Lung adenocarcinoma prognosis prediction model based on heme metabolism related genes, construction method and application

By integrating TCGA database and machine learning methods to screen heme metabolism-related genes, construct a risk score model for lung adenocarcinoma prognosis, solving the problem of population specificity and insufficient risk stratification of heme metabolism research in lung adenocarcinoma, and realizing data support and personalized treatment of precision medicine.

CN120472987APending Publication Date: 2025-08-12BEIJING CANCER HOSPITAL PEKING UNIV CANCER HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510482607.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The research on heme metabolism in lung adenocarcinoma in the prior art has insufficient population specificity, lack of effective risk stratification tools, and systematic lack of heme metabolism characteristic gene screening methods, which limits its application in clinical diagnosis and treatment.

Method used

By integrating lung adenocarcinoma data from TCGA database, using machine learning methods to screen out molecular markers and genes related to heme metabolism, construct a lung adenocarcinoma prognostic risk score model based on heme metabolism-related genes, combine it with deep neural network model to perform risk stratified prediction, identify high-risk patients and provide personalized treatment strategies.

Benefits of technology

The clinical significance of systematic analysis of heme metabolism in lung adenocarcinoma is achieved from the population level, and accurate prognosis prediction tools and personalized treatment plans are provided, breaking through the limitations of traditional laboratory models and improving prediction accuracy and targeted treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472987A_ABST
    Figure CN120472987A_ABST
Patent Text Reader

Abstract

The invention discloses a lung adenocarcinoma prognosis prediction model based on heme metabolism related genes and a construction method and application thereof. According to the invention, a heme metabolism related gene set is screened through a molecular characteristic database, and a gene significantly related to prognosis is screened by using transcriptome and clinical data of lung adenocarcinoma in a TCGA database and combining single-factor Cox regression, LASSO regression and multivariable Cox analysis. And through multi-time iterative modeling, selecting high-frequency stable genes and corresponding regression coefficient mean values to construct a heme metabolism risk score. According to the method, a random survival forest model is adopted to carry out iteration evaluation on gene importance for multiple times, the first 50% of genes which have the maximum influence on the model are screened, and finally the five important characteristic genes related to the heme metabolism risk are identified by taking an intersection with an LASSO result. The application provides new theoretical basis and technical support for precise treatment of lung adenocarcinoma, and has important clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedicine, and in particular to a lung adenocarcinoma prognosis prediction model based on heme metabolism-related genes using machine learning technology, as well as a construction method and application thereof. Background Art

[0002] Lung adenocarcinoma (LUAD) is one of the most common and lethal malignant tumors worldwide, posing a major challenge to public health. In recent years, the key role of tumor cell metabolism in disease progression and treatment response has gradually attracted attention. Among them, heme metabolism plays an important role in the occurrence, proliferation, metastasis, energy metabolism and regulation of treatment sensitivity of tumors. Heme metabolism is not only involved in basic metabolic processes such as oxidative phosphorylation, but also affects the survival and drug resistance of tumor cells by regulating programmed cell death pathways such as ferroptosis. Although existing studies have preliminarily revealed the importance of heme metabolism in tumors, the following problems and limitations still exist:

[0003] 1. Insufficient population-specific research: Current research on heme metabolism in LUAD is largely limited to laboratory models, such as cell lines and animal models. While these models help reveal molecular mechanisms, their results are often difficult to directly translate to clinical practice. Due to tumor heterogeneity and individual patient differences, laboratory models cannot fully reflect the diversity of the real-world population. Furthermore, systematic analyses based on large-scale clinical data are still lacking, resulting in the clinical significance of heme metabolism in LUAD remaining underexplored.

[0004] 2. Lack of effective risk stratification tools: Currently, methods for assessing the role of heme metabolism in the prognosis and treatment response of LUAD patients are relatively limited. Traditional biomarker screening methods typically rely on the expression levels of a single gene or protein, which makes it difficult to fully reflect the complex regulatory network of heme metabolism. Furthermore, existing methods lack effective tools for patient risk stratification, making it impossible to accurately predict patient prognosis or guide personalized treatment. This limits the application of research results related to heme metabolism in clinical practice.

[0005] 3. Systematic deficiencies in existing methods for screening heme metabolism characteristic genes: There are three key limitations in the current research on screening heme metabolism-related characteristic genes: methodological simplicity (relying on simple transcriptome differential expression analysis and univariate statistical tests at the cellular level, without considering synergistic effects or nonlinear relationships between genes), lack of functional annotation (ignoring the biological function weight of genes in the pathway), and insufficient clinical interpretability (screening results are statistically significant but the mechanism association is unclear, making it difficult to convert into clinical markers).

[0006] In summary, existing technologies in the field of heme metabolism research face challenges such as insufficient population-specific studies, a lack of risk stratification tools, and a lack of systematic approaches to screening for heme metabolism signature genes. These limitations restrict the application of heme metabolism in the clinical diagnosis and treatment of LUAD, and new technologies and methods are urgently needed to address these issues. Summary of the Invention

[0007] To address the aforementioned technical issues, the present invention provides a lung adenocarcinoma prognosis prediction model based on heme metabolism-related genes, as well as a method for constructing and applying such a model. This application not only addresses the existing challenges of insufficient population-specific research on heme metabolism and the lack of risk stratification tools, but also provides new theoretical basis and technical support for the precision treatment of lung adenocarcinoma, demonstrating its significant clinical application value.

[0008] In a first aspect, the present invention provides a use of a heme metabolism-related molecular marker, which is achieved through the following technical solutions.

[0009] The invention relates to an application of a heme metabolism-related molecular marker in the preparation of a product for predicting the prognosis of lung adenocarcinoma, wherein the heme metabolism-related molecular marker is composed of the following genes: ABCC2, SLCO1B3, SLCO2B1, JCHAIN, AQP3, DMTN, EIF2AK1, FBXO9, HTATIP2, LRP10, MAP2K3, NFE2, NNT, SLC2A1, SMOX, and TENT5C.

[0010] Preferably, the heme metabolism-related molecular markers are lung adenocarcinoma core heme metabolism risk characteristic genes consisting of SLC2A1, SMOX, ABCC2, SLCO1B3, and FBXO9.

[0011] In a second aspect, the present invention provides a lung adenocarcinoma prognostic risk scoring model based on heme metabolism-related genes, which is achieved through the following technical solutions.

[0012] A lung adenocarcinoma prognostic risk scoring model based on heme metabolism-related genes was constructed based on the above markers, and the calculation formula was: Risk Score=ABCC2×0.0738+SLCO1B3×0.0307+SLCO2B1×(-0.0524)+JCHAIN×(-0.0341)+AQP3×(-0.0289)+DMTN×(-0.0999)+EIF2AK1×0.0880+FBXO9×(-0.0941)+HTATIP2×0.0345+LRP10×0.1718+MAP2K3×0.0365+NFE2×(-0.0848)+NNT×(-0.0567)+SLC2A1×0.0764+SMOX×0.0530+TENT5C×(-0.0631). The patients were divided into two prognostic subgroups: high-risk group and low-risk group using the median of the risk score as the critical value.

[0013] In a third aspect, the present invention provides a method for constructing a lung adenocarcinoma prognostic risk scoring model based on heme metabolism-related genes, which is achieved through the following technical solutions.

[0014] A method for constructing the above model comprises the following steps:

[0015] Gene sets related to heme metabolism were collected through a molecular information database system, and the gene sets were sorted and duplicates were removed to obtain genes related to heme metabolism;

[0016] Obtain transcriptome and clinical data of lung adenocarcinoma from the TCGA database;

[0017] Using overall survival time and survival status as outcome variables and heme metabolism-related gene expression values as predictor variables, univariate Cox proportional hazards regression analysis was performed to screen heme metabolism-related genes that were significantly associated with the prognosis of lung adenocarcinoma.

[0018] Using behavioral genes, gene expression matrices with columns as samples, and patient survival data including survival time and survival status as input data, multiple iterations of LASSO regularized Cox proportional hazard model regression analysis were performed;

[0019] The frequency of each gene selected in the iterative process of the LASSO regularized Cox proportional hazard model regression iteration results was counted, and genes with a frequency threshold >50% were identified as high-frequency stable genes;

[0020] The expression levels of high-frequency stable genes and the average values of their regression coefficients in multiple iterations were weighted and summed to obtain a prognostic risk scoring model for lung adenocarcinoma based on heme metabolism-related genes.

[0021] Furthermore, the specific method for obtaining the transcriptome and clinical data of lung adenocarcinoma in the TCGA database is as follows: obtaining the mRNA TPM expression profile and supporting clinical information data, performing log2 normalization conversion on the original TPM expression matrix, and retaining only samples with sample numbers ending in "01A".

[0022] Furthermore, LASSO regularized Cox proportional hazards model regression analysis was performed for 100 iterations with the following parameters: alpha = 1, family = "cox", type.measure = "deviance", nfolds = 10, and standardize = TRUE.

[0023] In a fourth aspect, the present invention provides a computer-readable storage medium, which is implemented through the following technical solutions.

[0024] A computer-readable storage medium stores a computer program, which controls the device where the computer-readable storage medium is located to execute the above-mentioned model when the computer program is running.

[0025] In a fifth aspect, the present invention provides a method for screening core heme metabolism risk characteristic genes for lung adenocarcinoma, which is achieved through the following technical solutions.

[0026] A method for screening core heme metabolism risk signature genes for lung adenocarcinoma, comprising the following steps:

[0027] Gene sets related to heme metabolism were collected through a molecular information database system, and the gene sets were sorted and duplicates were removed to obtain genes related to heme metabolism;

[0028] Obtain transcriptome and clinical data of lung adenocarcinoma from the TCGA database;

[0029] Using overall survival time and survival status as outcome variables and heme metabolism-related gene expression values as predictor variables, univariate Cox proportional hazards regression analysis was performed to screen heme metabolism-related genes that were significantly associated with the prognosis of lung adenocarcinoma.

[0030] Using behavioral genes, gene expression matrices with columns as samples, and patient survival data including survival time and survival status as input data, multiple iterations of LASSO regularized Cox proportional hazard model regression analysis were performed;

[0031] The frequency of each gene selected in the iterative process of the LASSO regularized Cox proportional hazard model regression iteration results was counted, and genes with a frequency threshold >50% were identified as high-frequency stable genes;

[0032] A random survival forest model was constructed with the following input data: survival time and survival status, and the expression matrix of genes significantly associated with prognosis screened by univariate Cox regression. Multiple iterative modeling was used to ultimately screen out the top 50% of genes with the greatest impact on the model. These genes were then intersected with the high-frequency stable genes obtained by the LASSO regression model to obtain the core heme metabolism risk signature genes for lung adenocarcinoma.

[0033] Furthermore, the random survival forest model was run for 100 iterations with the following parameters: ntree = 500, nodesize = 5, samptype = "swr", and mtry = 7.

[0034] In a sixth aspect, the present invention provides a method for constructing a deep heme metabolism risk classification prediction model that integrates genes and clinical features, which is achieved through the following technical solutions.

[0035] A method for constructing a deep heme metabolism risk classification prediction model that integrates gene and clinical features adopts a deep neural network model. The model input data includes the core heme metabolism risk characteristic genes of lung adenocarcinoma and clinical information. The clinical information includes survival status, risk score calculated by the above model, and risk grouping based on the above model. The deep neural network model architecture adopts a three-layer fully connected network: input layer: 64 neurons, ReLU activation function, input dimension input_shape is the gene feature number ncol(train_x), train_x represents the number of genes in the important characteristic genes related to heme metabolism risk; hidden layer: 32 neurons, ReLU activation; output layer: 1 neuron, Sigmoid activation; binary_crossentropy is used as the loss function, Adam optimizer is used, and accuracy is used as the evaluation indicator.

[0036] This application has the following beneficial effects.

[0037] (1) This application integrates large-scale transcriptome data from lung adenocarcinoma (LUAD) patients in the TCGA database and uses machine learning methods to stratify the expression patterns of genes related to heme metabolism in patients. This method breaks through the limitations of traditional laboratory models and can systematically analyze the clinical significance of heme metabolism in LUAD at the population level, providing data support for precision medicine;

[0038] (2) The HMRS system constructed in this application can accurately assess the role of heme metabolism in the prognosis and treatment response of LUAD patients. By risk stratifying patients, the HMRS system provides clinicians with a reliable prognostic prediction tool, helping to identify high-risk patients and develop personalized treatment strategies;

[0039] (3) This application combines the transcriptome data of tumor patients and uses machine learning algorithms (LASSO regression, random forest) and deep neural network (DNN) learning to identify gene combinations with synergistic effects, breaking through the limitations of univariate analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0041] Figure 1 This is a graph showing the results of the univariate Cox proportional hazards regression analysis of the present invention to screen 49 heme metabolism-related genes that are significantly associated with prognosis;

[0042] Figure 2 This is the pathway diagram of 49 heme metabolism-related genes in the 100th iteration of the LASSO regularized Cox proportional hazards model regression analysis of the present invention (16 high-frequency stable genes and their corresponding line colors are marked in the upper right corner);

[0043] Figure 3 This is a diagram for determining the cutoff value for HMRS risk stratification in the TCGA-LUAD cohort of the present invention;

[0044] Figure 4 This is a Kaplan-Meier survival analysis result diagram based on HMRS risk stratification in the TCGA-LUAD cohort of the present invention;

[0045] Figure 5 3 is a graph showing the time-dependent ROC analysis results of HMRS-based risk stratification in the TCGA-LUAD cohort of the present invention;

[0046] Figure 6 This is a graph showing the evaluation results of the optimal number of decision trees for the random survival forest (RSF) model of the present invention (the arrow indicates the lowest point in the error rate, which means the error rate of the model reaches its lowest value, after which the error rate tends to stabilize);

[0047] Figure 7 This is an importance score graph for evaluating gene features based on the random survival forest (RSF) model of the present invention;

[0048] Figure 8 This is an important characteristic gene map related to heme metabolism obtained by screening the random survival forest (RSF) model and the LASSO regression model in the present invention;

[0049] Figure 9This is a graph showing the loss / accuracy of the deep learning DNN model of the present invention;

[0050] Figure 10 This is the confusion matrix heat map of the deep learning DNN model of the present invention (showing the correspondence between true and predicted labels. The table below shows that the model has an accuracy of 0.79, a sensitivity of 0.82, and a specificity of 0.77);

[0051] Figure 11 This is a stacked bar chart of the label distribution of the deep learning DNN model of the present invention (comparing the ratio of real and predicted labels);

[0052] Figure 12 This is the ROC curve of the deep learning DNN model of the present invention (showing the AUC value of the expression matrix of 5 important characteristic genes predicting risk classification);

[0053] Figure 13 This is a diagram showing the cutoff value determination for the validation cohort GSE31210 based on HMRS risk stratification;

[0054] Figure 14 This is a diagram showing the cutoff value determination for the validation cohort GSE68465 based on HMRS risk stratification;

[0055] Figure 15 This is a Kaplan-Meier survival analysis result diagram of the validation cohort GSE31210 based on HMRS risk stratification;

[0056] Figure 16 This is a graph showing the time-dependent ROC analysis results of the validation cohort GSE31210 based on HMRS risk stratification;

[0057] Figure 17 This is a Kaplan-Meier survival analysis result diagram of the validation cohort GSE68465 based on HMRS risk stratification;

[0058] Figure 18 This is a graph showing the time-dependent ROC analysis results of the validation cohort GSE68465 based on HMRS risk stratification;

[0059] Figure 19 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0060] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Prior to the description, it should be understood that the terms used in the specification and the appended claims are not to be construed as limited to their general and dictionary meanings, but rather should be interpreted based on the meanings and concepts corresponding to the technical aspects of the present invention, based on the principle that allows the inventor to appropriately define the terms for the best interpretation. Therefore, the description herein is merely a preferred example for illustrative purposes and is not intended to limit the scope of the present invention. It should be understood that other equivalent implementations and modifications may be made without departing from the spirit and scope of the present invention.

[0061] like Figure 19 As shown, the present invention systematically constructed and validated a prognostic prediction model for lung adenocarcinoma (LUAD) based on heme metabolism gene signatures by integrating multi-omics data with machine learning methods. First, a heme metabolism-related gene set was screened using molecular signature databases (e.g., MsigDB). Using transcriptome and clinical data from the TCGA-LUAD cohort, univariate Cox regression (p < 0.05), LASSO regression, and multivariate Cox analysis were combined to identify genes significantly associated with prognosis. Through multiple iterative modeling, a heme metabolism risk score (HMRS) was constructed by selecting high-frequency stable genes and their corresponding regression coefficient means. Patients were stratified into high- and low-risk groups based on the median HMRS score. Its prognostic predictive efficacy was validated by survival analysis and time-dependent receiver operating characteristic (ROC) curves. To further identify key signature genes, a random survival forest (RSF) model was used to iteratively evaluate gene importance, selecting the top 50% of genes with the greatest impact on the model. Finally, the intersection of the results with the LASSO results identified five key signature genes associated with heme metabolism risk. Based on this, a deep neural network (DNN) model with a 64-32-1 layer architecture (ReLU / Sigmoid activation) was used to validate the TCGA-LUAD dataset. Five key eigengenes associated with heme metabolism risk demonstrated significant discriminant performance in the HMRS-based risk stratification classification model in terms of accuracy, AUC, and other indicators. Furthermore, the HMRS model maintained stable prognostic prediction capabilities in an independent validation dataset. This present invention provides new heme metabolism-related molecular markers and calculation methods for LUAD prognosis prediction.

[0062] The invention is further described below with reference to the accompanying drawings and examples. Unless otherwise specified, the experimental methods used in this invention are conventional, and all experimental equipment, materials, and reagents used can be purchased from relevant material sales companies. R version 4.3.3 used in this application was downloaded from https: / / cran.r-project.org / bin / windows / base / old / 4.3.3 / . The software packages easyTCGA, randomForestSRC, and reticulate were downloaded from https: / / github.com / , and the software packages survival, glmnet, survmine, timeROC, ggplot2, keras, TensorFlow, and caret were downloaded from https: / / cran.r-project.org / .

[0063] 1. The screening process for heme metabolism-related genes (HMGs) involved in the present invention is as follows: First, the molecular information database MsigDB (Molecular Signatures Database, https: / / www.gsea-msigdb.org / gsea / msigdb) systematically collects gene sets related to heme metabolism, including five gene sets: "REACTOME HEME BIOSYNTHESIS", "REACTOME HEME DEGRADATION", "WIKIPATHOWS HEME BIOSYNTHESIS", "REACTOME SCAVENGING OF HEME FROM PLASMA", and "HALLMARK HEME METABOLISM". By integrating the genes in these gene sets and removing duplicates, 282 heme metabolism-related genes were finally obtained (as shown in Table 1).

[0064] Table 1. 282 heme metabolism-related genes

[0065]

[0066] 2. This study was based on the TCGA-LUAD (lung adenocarcinoma) dataset (https: / / portal.gdc.cancer.gov / , dbGaP Study Accession = phs000178). mRNA TPM expression profiles and accompanying clinical information data were obtained using the R language package easyTCGA (version 0.0.4.2000) using R version 4.3.3 (released on February 29, 2024) (download date: February 14, 2025). Through a bioinformatics preprocessing process, the original TPM expression matrix was first log2 normalized (log2(TPM+1)) to eliminate data skewness and improve analytical reliability. To ensure the homogeneity of the study subjects, only samples with sample numbers ending in "01A" were retained. These samples were all primary tumor samples, effectively eliminating interference from metastatic tumors, blood samples, and normal tissue samples. Samples with a follow-up time of 2000 days or less were further screened based on clinical data. Survival status was defined as status (1 = Dead, 0 = Alive). For samples with status = 0, only the record of the longest follow-up time within 2000 days was retained. For samples with status = 1, only the record of the last event was retained. Finally, 405 samples that met the requirements were obtained for the next analysis.

[0067] 3. The prognostic analysis method of heme metabolism-related genes of the present invention specifically includes the following steps: First, based on the TCGA-LUAD dataset, with 2000 days as the observation endpoint, the patient's overall survival time (Overall Survival, OS) and survival status (1=Dead, 0=Alive) are used as outcome variables, and the expression values of 282 heme metabolism-related genes are used as predictor variables. The 282 heme metabolism-related genes were subjected to univariate Cox proportional hazard regression analysis using the R language survival software package (version 3.5-8) to calculate the hazard ratio (Hazard Ratio, HR), regression coefficient (Coef), statistical significance level (p value) and 95% confidence interval (95% CI) of each gene. Using p < 0.05 as the significance threshold, 49 heme metabolism-related genes that were significantly correlated with the prognosis of lung adenocarcinoma were finally screened out, see Tables 2 and Figure 1 .

[0068] Table 2. 49 heme metabolism genes significantly associated with the prognosis of lung adenocarcinoma

[0069]

[0070]

[0071] 4. The present invention is based on the LASSO regularized Cox proportional hazard model regression analysis of heme metabolism-related gene prognostic model construction method, the specific implementation steps are as follows: using the cv.glmnet function in the R language glmnet software package (version 4.1-8), with the gene expression matrix (behavioral genes, columns as samples) and patient survival data (including survival time and survival status) as input data, setting the random number seed to set.seed (123) to ensure the reproducibility of the results. Through 100 repeated LASSO regularized Cox proportional hazard model regression analysis, each analysis uses the 10-fold cross-validation technique, and the key parameters are set as: alpha = 1 (pure LASSO penalty), family = "cox" (Cox proportional hazard model), type.measure = "deviance" (based on partial likelihood deviation evaluation criteria), nfolds = 10 (10-fold cross-validation), standardize = TRUE (expression value standardization). In each iteration, the optimal penalty parameter lambda.min was determined through cross-validation, and the regression coefficient of each gene was recorded (the coefficient of the unselected gene was recorded as 0), as shown in Table 3.

[0072]

[0073] Table 3

[0074]

[0075] Table 3

[0076]

[0077] Table 3

[0078]

[0079] 5. The results of 100 LASSO regularized Cox proportional hazard model regression iterations were integrated and analyzed, and the frequency of each gene selected during the iteration was counted to form a gene selection frequency table (Table 4). Using strict screening criteria, genes with a frequency threshold of >50% (threshold = 0.5) were set as high-frequency stable genes. 16 high-frequency stable genes were screened out from 49 candidate genes. The pathway diagram of the 16 high-frequency stable genes in heme metabolism can be found in Figure 2 The average regression coefficients of the 16 high-frequency stable genes across 100 iterations (accurate to 4 decimal places) were further calculated to form a gene coefficient table (Table 5). The screening method of the present invention ensures the stability of the results through multiple iterations. The 16 genes obtained will be used to construct a prognostic risk score model for lung adenocarcinoma.

[0080] Table 4. Selection frequency of 49 heme metabolism-related genes in 100 iterations based on LASSO regularized Cox proportional hazards model regression analysis

[0081]

[0082]

[0083] Table 5. Average regression coefficients of 16 high-frequency stable genes related to heme metabolism in 100 iterations based on LASSO regularized Cox proportional hazards model regression analysis

[0084]

[0085] 6. The expression of high-frequency stable genes and their average coefficients are weighted and summed to obtain the risk score named HMRS (Heme Metabolism Risk Score), HMRS=ABCC2×0.0738+SLCO1B3×0.0307+SLCO2B1×(-0.0524)+JCHAIN×(-0.0341)+AQP3×(-0.0289)+DMTN×(-0.0999)+EIF2AK1×0.0880+FBXO9 ×(-0.0941)+HTATIP2×0.0345+LRP10×0.1718+MAP2K3×0.0365+NFE2×(-0.0848)+NNT×(-0.0567)+SLC2A1×0.0764+SMOX×0.0530+TENT5C×(-0.0631).

[0086] 7. The HMRS formula was used to quantify the risk score of each patient in the TCGA-LUAD cohort, with the median risk score of all patients as the critical value (cutoff = 0.55) (see Figure 3 ), patients were divided into two prognostic subgroups: a high-risk group ("High") and a low-risk group ("Low"). The Kaplan-Meier survival model was constructed using the survfit function in the survival package (version 3.5-8) in R. The input parameters included: 1) survival time variable time; 2) survival status variable status (1 = Dead, 0 = Alive); and 3) risk group variable (High / Low). Survival curves with statistically significant markers were generated using the ggsurvplot function in the survminer package (version 0.5.0) (see Figure 4 ), visually showing the survival differences between high-risk and low-risk groups.

[0087] 8. The predictive efficacy of the HMRS risk score was evaluated based on time-dependent ROC analysis. Clinical follow-up time was first converted to standard days (1 year = 365 days, 3 years = 1095 days, and 5 years = 1825 days). Modeling and analysis were performed using the timeROC function in the R language timeROC package (version 0.4). Input parameters included: 1) survival time variable T (number of days); 2) survival status variable delta (1 = Dead, 0 = Alive); 3) risk score marker = HMRS; 4) event type parameter cause = 1 (Dead event); 5) weighting method weighting = "marginal" (marginal weight); and 6) iid = TRUE (calculation of confidence interval). Statistical power was ensured by systematically checking the number of events before each time point. The false positive rate (FPR), true positive rate (TPR), and AUC values (including 95% confidence intervals) were extracted to construct the evaluation data framework. Visualization was finally achieved using the ggplot2 package (version 3.5.1) (see ). Figure 5 ), where: 1) use geom_smooth(method="loess") to draw a smooth ROC curve; 2) color the predictions differently by time point (1 year = red, 3 years = green, 5 years = blue); 3) add a gray dashed reference line (AUC = 0.5); 4) accurately label the AUC value for each time point in the lower right corner of the chart (keep two decimal places).

[0088] 9. The random survival forest model was constructed using the rfsrc() function of the randomForestSRC software package (version 3.3.3). The input data included: 1) the survival time and survival status (1 = Dead, 0 = Alive) of TCGA-LUAD patients; 2) the expression matrix of 49 prognostic-significant genes screened by univariate Cox regression. The key parameters of the model were set as follows: ntree = 500 (number of decision trees), nodesize = 5 (minimum sample size of terminal nodes), samptype = "swr" (self-service sampling with replacement), and mtry = 7 (7 candidate genes were randomly selected at each tree split, calculated by sqrt(51), where 51 = 49 genes + 2 columns of survival data). The optimal tree size was determined by evaluating the out-of-bag error rate (OOB error rate) of 1-500 trees (see Figure 6 To enhance the reliability of the results, 100 iterations were used for modeling (num_iterations = 100), and different random seeds were set each time through set.seed(2024+i). The feature importance scores of each gene were calculated and the average values were summed up and sorted in descending order (see Figure 7By calculating the proportion of each gene's cumulative feature importance to the total importance, genes with a proportion not exceeding 50% were selected to determine the top 50% of genes with the greatest impact on the model. Finally, 9 genes were screened as the risk genes that contributed most to the RSF model (SLC2A1, SMOX, RAD23A, ABCC2, FABP1, SLCO1B3, GMPS, KAT2B, FBXO9). The intersection of these genes with the 16 risk genes obtained by the LASSO regression model identified 5 core heme metabolism risk signature genes: SLC2A1, SMOX, ABCC2, SLCO1B3, FBXO9 (see Figure 8 ).

[0089] 10. A deep neural network (DNN) survival risk prediction model was constructed based on the R language package keras (version 2.15.0) and TensorFlow (version 2.16.0) backend. The entire process was ensured to be reproducible by using set.seed (123). The data was converted to NumPy arrays using the reticulate package (version 1.37.0) to be compatible with the Keras backend. Finally, an end-to-end deep heme metabolism risk classification prediction model based on gene-clinical feature fusion was constructed. The model input data included the expression matrix of five important characteristic genes and clinical information (including survival status, risk score HMRS, and risk group based on HMRS). First, the data (TCGA-LUAD cohort) was partitioned into a test set and a training set according to the survival status ratio of 8:2 using the createDataPartition function of the caret package (version 7.0-1). The feature matrix was the gene expression data of the five important characteristic genes related to heme metabolism risk, and the label vector was the risk group (high risk and low risk) converted based on HMRS. The DNN model architecture uses a three-layer fully connected network: input layer (64 neurons, ReLU activation function, input dimension input_shape is the number of gene features ncol(train_x), train_x represents the number of genes in the important characteristic genes related to hemoglobin metabolism risk), hidden layer (32 neurons, ReLU activation) and output layer (1 neuron, Sigmoid activation), using binary_crossentropy as the loss function, Adam optimizer, and accuracy as the evaluation indicator. The training process sets 100 epochs and a batch_size of 32, and the test set performance is monitored by validation_data. The test set accuracy (accuracy), confusion matrix (confusionMatrix) and AUC value (pROC::roc) are calculated in the model evaluation phase, and the prediction results are binarized by a threshold of 0.5. The visualization part contains the training process curve (loss and accuracy changes) (see Figure 9 ), confusion matrix heat map (showing the correspondence between true and predicted labels) (see Figure 10 ), label distribution stacked bar chart (comparing the true and predicted label ratios) (see Figure 11 ) and ROC curve (showing that the AUC value of the expression matrix of 5 important characteristic genes in predicting risk classification is 0.9) (see Figure 12 ).

[0090] 11. Validate the model's applicability in other datasets. Download GSE31210 (https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE31210, a total of 163 lung adenocarcinoma patient tumor tissue data) and GSE68465 (https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE68465, a total of 273 lung adenocarcinoma patient tumor tissue data) from GEO (Gene Expression Omnibus, https: / / www.ncbi.nlm.nih.gov / geo) for independent validation. The gene expression matrix and clinical information data were extracted from the data set, and the original expression data were uniformly converted into the log2(TPM+1) standardized format; based on the above HMRS formula (HMRS=ABCC2×0.0738+SLCO1B3×0.0307+SLCO2B1×(-0.0524)+JCHAIN×(-0.0341)+AQP3×(-0.0289)+DMTN×(-0.0999)+EIF2AK1×0.0880+FBXO9×(-0.0941)+HTATIP2×0.0345+LRP10×0.1718+MAP2K3 The risk score of each patient was calculated using the survfit function of the survival package (input parameters: survival time, survival status, and high / low risk groups based on the median), and the ggsurvplot function of the survminer package was used to generate the survival curve. Following the method in step 8, time-dependent ROC analysis was performed using the timeROC function of the timeROC package (version 0.4) (parameter settings: survival time T, survival status delta, marker = HMRS, event type cause = 1, weighting method weighting = "marginal", iid = TRUE). The predictive efficacy (AUC value and 95% confidence interval) of the model at the 1-year (365-day), 3-year (1095-day) and 5-year (1825-day) time points was evaluated. Finally, it was confirmed that the HMRS model still maintained stable prognostic prediction ability in the independent validation data set. The experimental results are shown in [ 1 ]. Figure 13-18 .

[0091] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. Use of a heme metabolism-related molecular marker in the preparation of a product for predicting the prognosis of lung adenocarcinoma, characterized in that: Heme metabolism-related molecular markers are composed of the following genes: ABCC2, SLCO1B3, SLCO2B1, JCHAIN, AQP3, DMTN, EIF2AK1, FBXO9, HTATIP2, LRP10, MAP2K3, NFE2, NNT, SLC2A1, SMOX, and TENT5C.

2. The use according to claim 1, characterized in that: Heme metabolism-related molecular markers are the core heme metabolism risk characteristic genes of lung adenocarcinoma composed of SLC2A1, SMOX, ABCC2, SLCO1B3, and FBXO9.

3. A lung adenocarcinoma prognostic risk scoring model based on heme metabolism-related genes, characterized by: The model is constructed based on the markers described in claim 1, and the calculation formula is: Risk Score = ABCC2 × 0.0738 + SLCO1B3×0.0307+SLCO2B1×(-0.0524)+JCHAIN×(-0.0341)+AQP3×(-0.0289)+DMTN×(-0.0999)+EIF2AK1×0.0880+FBXO9×(-0.0941)+HTATIP2×0.0345+LRP10×0.1718+ MAP2K3×0.0365+NFE2×(-0.0848)+NNT×(-0.0567)+SLC2A1×0.0764+SMOX×0.0530+TENT5C×(-0.0631), with the median of the risk score as the critical value, divided into two prognostic subgroups: high-risk group and low-risk group.

4. A method for constructing the model according to claim 3, characterized in that: The following steps are involved: Gene sets related to heme metabolism were collected through a molecular information database system, and the gene sets were sorted and duplicates were removed to obtain genes related to heme metabolism; Obtain transcriptome and clinical data of lung adenocarcinoma from the TCGA database; Using overall survival time and survival status as outcome variables and heme metabolism-related gene expression values as predictor variables, univariate Cox proportional hazards regression analysis was performed to screen heme metabolism-related genes that were significantly associated with the prognosis of lung adenocarcinoma. Using behavioral genes, gene expression matrices with columns as samples, and patient survival data including survival time and survival status as input data, multiple iterations of LASSO regularized Cox proportional hazard model regression analysis were performed; The frequency of each gene selected in the iterative process of the LASSO regularized Cox proportional hazard model regression iteration results was counted, and genes with a frequency threshold > 50% were identified as high-frequency stable genes; The expression levels of high-frequency stable genes were weighted and summed with the average values of their regression coefficients in multiple iterations to obtain a prognostic risk scoring model for lung adenocarcinoma based on heme metabolism-related genes.

5. The construction method according to claim 4, characterized in that: The specific method for obtaining the transcriptome and clinical data of lung adenocarcinoma in the TCGA database is as follows: obtain the mRNA TPM expression profile and supporting clinical information data, perform log2 normalization transformation on the original TPM expression matrix, and retain only samples with sample numbers ending in "01A".

6. The construction method according to claim 4, characterized in that: The LASSO regularized Cox proportional hazards model regression analysis was performed for 100 iterations with the following parameters: alpha = 1, family = "cox", type.measure = "deviance", nfolds = 10, and standardize = TRUE.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the model described in claim 3.

8. A method for screening core heme metabolism risk signature genes for lung adenocarcinoma, characterized by: The following steps are involved: Gene sets related to heme metabolism were collected through a molecular information database system, and the gene sets were sorted and duplicates were removed to obtain genes related to heme metabolism; Obtain transcriptome and clinical data of lung adenocarcinoma from the TCGA database; Using overall survival time and survival status as outcome variables and heme metabolism-related gene expression values as predictor variables, univariate Cox proportional hazards regression analysis was performed to screen heme metabolism-related genes that were significantly associated with the prognosis of lung adenocarcinoma. Using behavioral genes, gene expression matrices with columns as samples, and patient survival data including survival time and survival status as input data, multiple iterations of LASSO regularized Cox proportional hazard model regression analysis were performed; The frequency of each gene selected in the iterative process of the LASSO regularized Cox proportional hazard model regression iteration results was counted, and genes with a frequency threshold > 50% were identified as high-frequency stable genes; A random survival forest model was constructed with input data including survival time and survival status, and the expression matrix of genes significantly associated with prognosis screened by univariate Cox regression. Using multiple iterative modeling, we finally screened out the top 50% of genes with the greatest impact on the model, and took the intersection with the high-frequency stable genes obtained by the LASSO regression model to obtain the core hemoglobin metabolism risk characteristic genes of lung adenocarcinoma.

9. The screening method according to claim 8, characterized in that: The random survival forest model was run for 100 iterations with the following parameters: ntree = 500, nodesize = 5, samptype = "swr", and mtry = 7.

10. A method for constructing a deep heme metabolism risk classification prediction model integrating gene and clinical features, characterized by: A deep neural network model is used, and the model input data includes the core heme metabolism risk characteristic genes of lung adenocarcinoma as described in claim 2 and clinical information. The clinical information includes survival status, risk score calculated using the model described in claim 3, and risk grouping based on the model described in claim 3; the deep neural network model architecture adopts a three-layer fully connected network: input layer: 64 neurons, ReLU activation function, input dimension input_shape is the gene feature number ncol(train_x), train_x represents the number of genes in the important characteristic genes related to heme metabolism risk; hidden layer: 32 neurons, ReLU activation; output layer: 1 neuron, Sigmoid activation; binary_crossentropy is used as the loss function, Adam optimizer, and accuracy is used as the evaluation indicator.

Citation Information

Cited By

  • Multi-omics malignant pleural effusion immunometabolism reprogramming space-time heterogeneity analysis device

    CN121601145A