Liver cancer prognosis and immune response prediction method based on multi-omics machine learning
By integrating multi-omics data and combining Cox regression with the random survival forest algorithm, the IMLIRI model was constructed, which addresses the shortcomings of existing technologies in predicting the prognosis and treatment response of hepatocellular carcinoma. It achieves multi-center, stable prognostic assessment and integrated assessment of immunotherapy response, providing a convenient clinical tool to support precision treatment decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-05
AI Technical Summary
Existing prognostic and treatment response prediction technologies for hepatocellular carcinoma suffer from limitations such as single data dimensions, insufficient algorithm robustness, and incomplete functional coverage. They cannot achieve integrated assessment of prognostic evaluation and immunotherapy response, and lack intuitive clinical tools, making it difficult to meet the needs of multi-center clinical applications.
By constructing a prognostic and immune response prediction method for hepatocellular carcinoma based on multi-omics machine learning, integrating multi-omics data, and using Cox regression combined with random survival forest algorithm, 11 core immune genes were identified, an immunotherapy response index (IMLIRI) was constructed, and multi-dimensional validation was performed to ensure the stability and accuracy of the model.
It achieves stable consistency index in multi-center external cohorts, reduces performance fluctuations, provides convenient and accurate risk characterization, identifies cancer subtypes and immunotherapy response potential, supports precise immunotherapy decisions, and reduces testing costs and complexity.
Smart Images

Figure CN121545597B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of biomedicine and artificial intelligence. More specifically, this invention relates to a method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning, used in scenarios such as clinical oncology diagnosis and treatment decision support, tumor molecular mechanism research, and precision medicine solution development for prognostic assessment and immunotherapy response prediction of hepatocellular carcinoma. Background Technology
[0002] Hepatocellular carcinoma is the sixth most common malignant tumor and the third most deadly malignant tumor worldwide. As the most common pathological subtype of liver cancer (accounting for 75% to 85%), its treatment methods include surgery, interventional therapy, radiotherapy, chemotherapy, immunotherapy and targeted therapy. However, most patients are diagnosed at an advanced stage and miss the opportunity for radical surgery. Therefore, immune checkpoint inhibitors have become a key treatment option for these patients.
[0003] In practical use, immune checkpoint inhibitors, by blocking immune checkpoint pathways such as programmed death receptor 1 (PDR1) / PDR-ligand 1 (PDL-1), reverse T cell depletion and restore anti-tumor immune responses, extending the median overall survival (OS) of patients from 13.4 months to 19.2 months. For example, the combination of atezolizumab and bevacizumab, compared to sorafenib, extended OS from 13.4 months to 19.2 months, and increased the objective response rate from 5% to 30%. The combination of durvalumab and tesimumab has also become the standard treatment for newly diagnosed unresectable hepatocellular carcinoma, with OS superior to sorafenib. However, clinical data shows that only about 30% of hepatocellular carcinoma patients respond to immune checkpoint inhibitors, and only 20%-30% achieve long-term survival. The heterogeneity of treatment responses among patients with similar clinical stages is extremely high, highlighting the limitations of relying solely on clinical staging to guide immunotherapy, and necessitating the development of precise biomarkers and predictive tools.
[0004] Existing prognostic and treatment response prediction technologies for hepatocellular carcinoma suffer from several drawbacks, including limited data dimensions, insufficient algorithm robustness, incomplete functional coverage, and limited clinical applicability. Specifically, these include:
[0005] I. Most existing hepatocellular carcinoma prediction models rely on single-mathematical data (such as mRNA expression profiles only), ignoring key molecular features such as DNA methylation, copy number variation, and somatic mutations. These models fail to comprehensively reflect tumor heterogeneity and the regulatory mechanisms of the immune microenvironment, resulting in concordance indices falling below 0.5 in external validation cohorts. For example, some gene expression-based prediction models only focus on genes related to immune cell infiltration, without incorporating epigenetic or mutational information, making it difficult to explain the phenomenon of "differences in patient responses under the same gene expression pattern."
[0006] Second, existing hepatocellular carcinoma prediction models mostly employ single machine learning algorithms with weak generalization ability (such as logistic regression and single random forest). They do not consider the complexity of high-dimensional multi-omics data, are prone to overfitting, and exhibit poor validation performance in multi-center cohorts, easily affected by data bias and overfitting. For example, some models can achieve a consistency index of 0.6 in the training cohort, but it drops sharply to below 0.5 in the external validation cohort, indicating weak generalization ability and failing to meet the needs of multi-center clinical applications.
[0007] Third, existing hepatocellular carcinoma prediction models are functionally fragmented, failing to simultaneously assess prognosis and predict immunotherapy response. They focus on a single indicator of prognosis or treatment response, lacking integrated prognosis-treatment response assessment capabilities. For example, some models can predict patient survival but cannot determine suitability for immunotherapy; while some immune response prediction models are not linked to long-term survival outcomes, requiring clinicians to use multiple tools, resulting in inefficiency. Some models rely on gene testing or tissue biopsy, increasing medical costs and patient burden; furthermore, they lack intuitive clinical tools (such as nomograms), requiring physicians to possess specialized bioinformatics knowledge to interpret the results, hindering widespread application.
[0008] Furthermore, compared to predictive models for other digestive system tumors such as colorectal cancer, hepatocellular carcinoma (HCC) has more complex molecular mechanisms (e.g., differences in hepatitis virus infection and cirrhosis background), requiring higher standards for multi-dimensional data integration and algorithm stability. Existing prognostic models for colorectal cancer have improved predictive performance through large-scale clinical data from public databases, multi-algorithm comparisons (e.g., extreme gradient boosting, logistic regression), and grid search hyperparameter optimization. However, the field of HCC lacks a complete technical system similar to this "multi-data source - multi-algorithm screening - clinical translation tools," urgently requiring the filling of this gap. Summary of the Invention
[0009] One object of the present invention is to solve at least the above-mentioned problems and / or defects, and to provide at least the advantages described below.
[0010] To achieve these objectives and other advantages of the present invention, a method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning is provided, comprising:
[0011] S1. Collect multi-omics data related to hepatocellular carcinoma from the Cancer Genome Atlas Database (TCGA) to form the TCGA-LIHC cohort, collect external validation cohorts from the Gene Expression Comprehensive Database (GEO), and combine them with immunotherapy-specific cohorts to construct datasets for model training and validation.
[0012] S2. Perform feature integration and cluster analysis on multi-omics data to obtain the corresponding molecular subtype distribution of hepatocellular carcinoma, and determine whether there is data bias in the TCGA-LIHC cohort based on the molecular subtype distribution of hepatocellular carcinoma.
[0013] S3. Construct a prognostic model for hepatocellular carcinoma based on Cox regression combined with random survival forest, and identify 11 core immune genes and their corresponding weight coefficients through model training, constructing the following immunotherapy response index, IMLIRI score:
[0014] IMLIRI = (-0.099 × EEF1B2 expression level) + (-0.308 × FHL3 expression level) + (-0.087 × MMP1 expression level) + (-0.128 × MAPK7 expression level) + (-0.259 × MPHOSPH6 expression level) + (-0.248 × NROB1 expression level) + (0.190 × RLF expression level) + (0.203 × RGS7 expression level) + (0.434 × GSTM1 expression level) + (0.341 × PIK3IP1 expression level) + (0.273 × PPP1R1A expression level);
[0015] S4. The IMLIRI score given by the hepatocellular carcinoma prognostic model was validated in multiple dimensions using an external validation cohort and an immunotherapy-specific cohort, so that the IMLIRI score can be applied clinically after validation.
[0016] Preferably, in S1, the multi-omics data includes: mRNA abundance quantified by TPM, normalized long non-coding RNA levels, log2-converted microRNA data, Illumina HumanMethylation450 chip DNA methylation data, somatic mutation profile, survival time, death status, tumor-lymph node-metastasis staging, and treatment regimen.
[0017] The external validation queue includes GSE76427, GSE15654, GSE10143, GSE14520, and GSE116174 downloaded from GEO. The relevant data information in the external validation queue is consistent with the multi-omics data content, and the external validation queue is integrated into the META queue after batch effect correction.
[0018] The immunotherapy-specific cohorts include: the IMvigor210 cohort and GSE78220, GSE135222, and GSE91061 downloaded from GEO. The IMvigor210 cohort consists of hepatocellular carcinoma patients receiving programmed death-ligand 1 inhibitor therapy, while the GSE78220, GSE135222, and GSE91061 cohorts contain data from patients receiving immune checkpoint inhibitor therapy.
[0019] The inclusion criteria for the TCGA-LIHC cohort and external validation cohort include: pathological diagnosis of hepatocellular carcinoma, complete survival time and death status information, at least mRNA abundance expression data, and exclusion of patients with other malignant tumors or those diagnosed solely through autopsy or death certificates.
[0020] Preferably, in S1, the batch effect correction is performed on the mRNA abundance expression data of the external validation cohort, with the cohort source as the batch variable and the survival status as the covariate to eliminate batch effects;
[0021] The correction content for the mRNA abundance expression data includes:
[0022] For numerical variables similar to mRNA abundance, missing values are filled using the median of the variable in the corresponding cohort;
[0023] For categorical variables similar to tumor staging, the most frequent category is used for imputation;
[0024] Variables with a missing rate >20% are removed directly.
[0025] Preferably, in S2, the method for obtaining the distribution of hepatocellular carcinoma molecular subtypes is as follows:
[0026] S20. Feature screening is performed on the data information in the TCGA-LIHC cohort to screen for immune-related gene features from mRNA expression data, to screen for the top-m prognostic features I before expression mutation from miRNA, lncRNA and methylation data, and to screen for the top-n prognostic features II before mutation frequency from somatic mutation data. The gene features, prognostic features I and prognostic features II are then combined to obtain a multi-omics integrated feature set.
[0027] S21. Multiple unsupervised clustering algorithms are used to perform integrated clustering analysis on the multi-omics integrated feature set. The clustering prediction index and gap statistics are combined to determine cancer subtype I and cancer subtype II. The majority voting method is used to determine the final cancer subtype for the clustering results of multiple unsupervised clustering algorithms.
[0028] S22. Evaluate the data offset of the TCGA-LIHC cohort through subtype verification and characterization analysis;
[0029] The classification reliability in the subtype verification is achieved through the nearest template prediction and the centroid-based partitioning algorithm. When the Kappa coefficient is greater than 0.6, the classification reliability is proven.
[0030] The subtype molecular distribution differences in the subtype verification are achieved by using principal component analysis, t-distribution random neighborhood embedding, unified manifold approximation and projection dimensionality reduction visualization.
[0031] The differences in pathway enrichment among subtypes in the characterization analysis were achieved through gene set variation analysis, in which cancer subtype I was enriched in metabolic pathways, and cancer subtype II was enriched in proliferation-related pathways.
[0032] The characterization analysis also includes: analyzing the tumor microenvironment by calculating the immune cell infiltration score and immune characteristic score, and identifying cancer subtype I as immune-activated and cancer subtype II as immunosuppressive.
[0033] Preferably, in S3, the hepatocellular carcinoma prognostic model is selected from 303 algorithm combinations generated by the basic algorithm and the ensemble strategy, and the optimal algorithm combination of Cox regression and random survival forest is selected.
[0034] The integrated strategy includes: a primary evaluation index based on the average consistency index of the training queue and the validation queue, and auxiliary indicators covering the area under the curve and the p-value of the log-rank test.
[0035] Preferably, in S3, the 11 core immune genes are obtained in the following manner:
[0036] Univariate Cox regression was performed in the S30, TCGA-LIHC cohort and external validation cohort to screen for candidate immune-related gene sets that are associated with prognosis and have consistent hazard ratios across all cohorts.
[0037] S31. Input the candidate immune-related gene set into the joint training framework of Cox regression and random survival forest, and use alternating iteration combined with adaptive adjustment of model structure hyperparameters to construct a closed-loop optimization of feature selection-hyperparameter optimization-survival prediction to promote the convergence of the candidate immune-related gene set, so as to obtain the corresponding 11 core immune genes after convergence.
[0038] In the training of the hepatocellular carcinoma prognostic model, the screening results after each iteration were evaluated using a three-level assessment criterion:
[0039] The average consistency index C-indexⅠ of the 5-fold cross-validation of the TCGA-LIHC queue is used as the immediate optimization target to perform a first-level evaluation on the screening results after each iteration. When the average C-indexⅠ is not lower than the preset performance threshold C1, and when the change of C-indexⅠ compared with the previous iteration result does not exceed the preset minimum performance decline threshold δ1, or when its improvement reaches or exceeds the preset minimum performance improvement threshold δ2, the first-level evaluation is deemed to have passed. Otherwise, the model structure hyperparameters are adaptively adjusted, and the next iteration is entered.
[0040] The generalization ability of the average consistency index C-index II in the five independent GEO validation queues is evaluated in a secondary evaluation. If C-index II is not lower than the preset generalization threshold C2 and the performance fluctuation between the validation queues does not exceed the preset allowable range, the secondary evaluation is deemed to pass; otherwise, the model is retrained.
[0041] The cross-model consistency index after calibration is checked to achieve a three-level evaluation. When the consistency index reaches or exceeds the preset consistency threshold C3, the three-level evaluation is deemed to have passed, the iteration is terminated and the final core immune gene set is output. Otherwise, joint optimization is achieved through alternating iteration.
[0042] Preferably, the adaptive adjustment of the model structure hyperparameters is performed as follows:
[0043] For the key hyperparameters of the two types of models, a wide initial interval is set. A first round of global coarse search is performed, and the performance is recorded with the 5-fold cross-validation consistency index C-indexⅠ as the objective function.
[0044] The initial results were sorted according to C-index I, and the high-potential clusters in the top few percentiles were extracted. The performance response surface was fitted by local weighted regression and combined with density clustering to locate the performance concentration area as the potential area.
[0045] The hyperparameters are adaptively shrunk and optimized by automatically shrinking the parameter boundaries of the potential region and conducting a high-resolution local dense search. In each round, the center and range of the potential region are dynamically updated based on the latest performance until the convergence condition is met.
[0046] Preferably, the alternating iteration involves extracting the importance index of features from the Cox regression combined with the random survival forest in each round to calculate the comprehensive contribution. When the marginal contribution of any gene to the consistency index C-index is positive during the cross-validation process and exceeds the preset minimum contribution threshold, the sampling weight or retention probability of the corresponding gene pair is increased in the next round; otherwise, the gene pair is gradually deweighted or eliminated.
[0047] The importance index refers to the stability and selection frequency of regression coefficients in the Cox regression algorithm, and the split gain or variable importance in the random survival forest algorithm.
[0048] Preferably, in S3, the optimal cutoff value for IMLIRI is determined based on the overall survival of the TCGA-LIHC cohort, and patients are divided into a low IMLIRI group and a high IMLIRI group.
[0049] Preferably, in S4, the multi-dimensional verification includes:
[0050] Dimension I: Compare and validate the IMLIRI score given by the hepatocellular carcinoma prognostic model with clinical parameters and existing models using the consistency index.
[0051] Dimension II: The IMLIRI scores given by the hepatocellular carcinoma prognostic model were divided into low-risk and high-risk groups. The tumor immune phenotype tracking algorithm and the tumor immune dysfunction and rejection framework were used to analyze and validate the low-risk and high-risk groups in terms of immunotherapy response probability, immune cell rejection, and myeloid-derived suppressor cell infiltration.
[0052] The present invention has at least the following beneficial effects:
[0053] Firstly, this invention employs a multi-omics + multi-cohort external validation dataset to lay the foundation for prediction. Specifically, by deeply fusing data from five omics dimensions, the system systematically covers tumor molecular characteristics and the immune microenvironment, achieving information complementarity and noise cancellation. Under the same evaluation metrics, its consistency index is higher than that of models built based on a single omics. The IMLIRI model achieves an average consistency index of 0.638, with 1 / 3 / 5-year AUCs exceeding 0.65, ranking among the top in comparison with 96 published models. Based on the TCGA-LIHC training cohort and optimization using five GEO external cohorts covering different regions, pathological types, and population backgrounds, combined with validation in an immunotherapy-specific cohort, the consistency index remains stable in multi-center external validation cohorts, reducing performance fluctuations between different cohorts. It maintains stable performance across different etiologies and hepatocellular carcinoma subtypes in multi-center cohorts, effectively addressing the key pain point in clinical research: "good training results, poor validation results."
[0054] Secondly, the 11 immune prognostic genes output by the model of this invention enable convenient and accurate risk characterization. Specifically, this invention uses cross-algorithm iterative screening of multidimensional immune-related candidate genes, ultimately converging into a core gene set with the most prognostic information, achieving maximum predictive benefit with minimal features. Without sacrificing accuracy, the detection scale is significantly reduced, shifting risk scoring from "multi-omics heavy detection" to "few genes, strong signals," greatly reducing experimental and data processing complexity. This 11-gene set can be rapidly determined using conventional methods such as immunohistochemistry, avoiding reliance on high-throughput sequencing and complex bioinformatics platforms, thereby achieving a low-cost, low-barrier, and scalable clinical testing pathway, providing a lightweight foundation for the widespread implementation of the model.
[0055] Thirdly, this invention assists in decision-making for hepatocellular carcinoma (HCC) immunotherapy through subtyping. While achieving accurate prediction, the model also provides crucial support for breakthroughs in HCC immunotherapy at the mechanistic level. By identifying two cancer subtypes and 11 core immune-related genes, the core subtyping mechanism of "immune activation-immunosuppression" in HCC is clearly revealed. Specifically, the identification of the two cancer subtypes determines whether there is any stratification bias in the pre-training data preparation stage, ensuring comprehensive and systematic data. The core subtyping, through the division into low and high IMLIRI groups, determines suitability for immunotherapy. This discovery provides an important theoretical basis for accurately matching immunotherapy regimens. Among the 11 core immune-related genes, the NROB1 gene promotes tumor proliferation through the GSK3β–BAX pathway, and the MMP1 gene participates in shaping the immunosuppressive microenvironment, providing novel targets for developing new drugs that target and regulate the immune microenvironment. Most importantly, the model clearly shows that the low IMLIRI group has high MSI scores and high immune infiltration characteristics. This characteristic directly points to the potential for this group to respond well to immune checkpoint inhibitors, thus providing a clear biological basis for highly effective immunotherapy strategies such as "ICI + anti-angiogenesis" and assisting in the decision-making of immunotherapy regimens for hepatocellular carcinoma.
[0056] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description
[0057] Figure 1 This is a schematic diagram of the prognosis and immune response prediction method for hepatocellular carcinoma based on multi-omics machine learning of the present invention.
[0058] Figure 2 This is a schematic diagram of multi-omics clustering in the embodiment;
[0059] Figure 3 This is a diagram of the core genes of the IMLIRI model in the examples.
[0060] Figure 4 As an example, this is a comparison chart of the IMLIRI score and clinical parameters of the present invention based on the CGA-LIHC cohort;
[0061] Figure 5 As shown in the embodiment, this invention presents a comparison chart of IMLIRI scores and clinical parameters based on the META cohort.
[0062] Figure 6 Here is a diagram of the immunotherapy response in the examples:
[0063] Figure 7 This is a line diagram from an embodiment. Detailed Implementation
[0064] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.
[0065] This invention proposes a prognostic and immunotherapy response prediction system for hepatocellular carcinoma. It integrates multi-dimensional omics data and optimizes machine learning algorithms to construct a predictive model for clinical applications. The specific scheme is as follows:
[0066] (1) Acquisition and standardization preprocessing of multi-omics data
[0067] Raw multi-omics data on hepatocellular carcinoma were obtained from the Cancer Genome Atlas (TCGA) database, including: mRNA expression levels per million transcripts, normalized long non-coding RNA (lncRNA) levels, log2-converted microRNA (miRNA) data, DNA methylation microarray (450K / 850K) data, somatic mutation profiles, and corresponding clinical information (survival time, stage, treatment regimen, etc.).
[0068] Five external validation cohorts (GSE76427, GSE15654, GSE10143, GSE14520, and GSE116174) were collected from the Gene Expression Omnibus database (GEO), and immunotherapy-specific cohorts such as IMvigor210 and GSE78220 were also included.
[0069] The acquired data underwent standardized preprocessing, including: using the `sva` package in R to correct systematic batch effects among samples from different sources in the test set, achieving unified scale integration of cross-cohort data; for missing data, median imputation was used for numerical variables, and most frequent class imputation was used for categorical variables to ensure data integrity. Furthermore, principal component analysis was used to verify the batch correction and cross-cohort integration effects, and outliers deviating from the overall distribution characteristics were identified and removed. The final result is a standardized dataset used for model training, internal validation, and multi-cohort aggregation validation.
[0070] (2) Cancer subtype identification based on multi-omics integration
[0071] An immune-related gene map was extracted from the mRNA transcriptome; for miRNA, lncRNA, and methylation data, the top 2000 expression variants were retained; prognostic-related genes / sites were screened using univariate Cox regression (p<0.05); for somatic mutation data, the top 5% of mutation frequencies were retained to reduce data redundancy. Ten unsupervised clustering algorithms (including K-means clustering, hierarchical clustering, spectral clustering, etc.) were used for ensemble clustering analysis, and the optimal number of clusters was determined to be two by combining the clustering prediction index and gap statistics; consensus clustering was performed to eliminate the bias of single algorithms, and finally, hepatocellular carcinoma molecular subtypes (cancer subtype I and cancer subtype II) with differences in survival outcome, pathway enrichment characteristics, and immune microenvironment characteristics were identified. Subtype stability was verified using recent template prediction and centroid-based partitioning algorithms, with Kappa coefficients all >0.6, demonstrating reliable subtyping. Principal component analysis, t-distribution random neighborhood embedding, unified manifold approximation, and projection dimensionality reduction were used to visualize the differences in molecular distribution among subtypes. Gene set variation analysis revealed pathway enrichment characteristics (cancer subtype I is enriched in metabolic pathways such as drug metabolism and bile acid synthesis, while cancer subtype II is enriched in cell cycle and DNA repair pathways). Tumor microenvironment analysis clarified that cancer subtype I is immune-activated (high immune cell infiltration, low tumor-associated fibroblasts), while cancer subtype II is immunosuppressive (high immune rejection / suppression characteristics, high levels of myeloid-derived suppressor cell infiltration).
[0072] 3) Prognostic model construction and algorithm optimization
[0073] Univariate Cox regression was performed in the training cohort (p<0.001) and validation cohort (p<0.05) to screen for prognostic immune-related genes. The risk ratio of the genes was required to be consistent across all cohorts, resulting in 22 candidate immune-related genes.
[0074] We constructed a framework of 10 basic algorithms (Cox regression based on boosting algorithm, random survival forest, logistic regression, support vector machine, etc.), and generated 303 algorithm combinations through "basic algorithms + ensemble strategy". Using "average consistency index of training queue + validation queue" as the core evaluation index, combined with auxiliary indicators such as area under the curve and p-value of log-rank test, we selected the optimal combination as Cox regression based on boosting algorithm combined with random survival forest (average consistency index of 0.638, significantly higher than other combinations). Eleven core immune-related genes (NROB1, MMP1, MPHOSPH6, FHL3, MAPK7, EEF1B2, RLF, PIK3IP1, PPP1R1A, GSTM1, and RGS7) were identified by ranking the feature importance using a Cox regression algorithm based on boosting. Among them, 5 genes, including NROB1 and MMP1, were risk genes (hazard ratio > 1), while 6 genes, including FHL3 and PIK3IP1, were protective genes (hazard ratio < 1). An integrated machine learning-derived immunotherapy response index (IMLIRI) was constructed based on a random survival forest algorithm. The formula is: IMLIRI = Σ (core gene expression level × CoxBoost coefficient). Patients were then divided into a low IMLIRI group (< cutoff value) and a high IMLIRI group (≥ cutoff value) based on the optimal cutoff value.
[0075] (4) Multi-dimensional model validation and performance evaluation
[0076] In all cohorts, the concordance index between IMLIRI and traditional clinical parameters (age, sex, tumor stage, grade, etc.) and the area under the receiver operating characteristic (ROC) curves at 1, 3, and 5 years were calculated. The results showed that the concordance index of IMLIRI (0.638-0.682) was significantly higher than that of tumor stage (0.581-0.615) and age (0.529-0.593). The area under the ROC curves at 1, 3, and 5 years was >0.65. Furthermore, univariate / multivariate Cox regression analysis (hazard ratio = 5.610-5.703, p < 0.001) confirmed that IMLIRI is an independent prognostic factor.
[0077] A systematic search of PubMed databases since 2019 for studies on "hepatocellular carcinoma prognostic models" was conducted, including 96 published models (covering different molecular mechanisms such as immunity, metabolism, and cell cycle). The concordance index was compared across TCGA and five GEO cohorts. Results showed that IMLIRI ranked in the top 5% in all seven cohorts, with a concordance index of 0.638 in the TCGA-LIHC cohort, significantly higher than most models (mean 0.572). In the IMvigor210 cohort, the low-IMLIRI group showed significantly better 12-month and 24-month restricted survival and 6-month long-term survival than the high-IMLIRI group (p values were 0.051, <0.001, and 0.003, respectively). In the three immunotherapy cohorts GSE78220, GSE135222, and GSE91061, the treatment response rate in the low-IMLIRI group was 2.3-3.1 times that of the high-IMLIRI group (p<0.05).
[0078] Validated by the tumor immune phenotype tracking algorithm, the low IMLIRI group showed higher activity in the whole immunization process of "antigen release-antigen presentation-immune cell infiltration-T cell recognition", and the tumor immune dysfunction and rejection scores showed that its immunotherapy response probability was increased by more than 40%.
[0079] (5) Development of clinical translation tools
[0080] The IMLIRI score was integrated with key clinical variables (tumor stage, grade, and presence of cirrhosis) to construct a visualized prognostic nomogram. Calibration was validated through 1000 sampling runs, and the calibration curve showed a deviation of <5% between predicted and actual survival values, indicating high reliability. Patients were categorized into three strata based on the IMLIRI score: low risk (<0.35), intermediate risk (0.35-0.65), and high risk (>0.65). Low-risk patients were preferentially recommended for single or combination therapy with immune checkpoint inhibitors; intermediate-risk patients were advised to combine immune checkpoint inhibitors with targeted therapy; and high-risk patients required local treatment to improve tumor burden before considering immunotherapy. An online calculation tool was also developed, which automatically generates IMLIRI scores, prognostic probabilities, and treatment recommendations after clinicians input multi-omics data and clinical information, offering convenient operation.
[0081] Example:
[0082] This embodiment provides a complete technical solution encompassing "multi-omics integration - subtype classification - model construction - clinical translation," offering a reliable basis for the precision diagnosis and treatment of hepatocellular carcinoma and addressing the current lack of systematic predictive tools in the field. Figure 1 As shown, the specific operation steps are as follows:
[0083] Step 1: Data Preparation and Preprocessing (TCGA-LIHC and GEO Queues)
[0084] 1.1 Acquisition of Multi-omics Data
[0085] Hepatocellular carcinoma multi-omics data were downloaded from the Cancer Genome Atlas (TCGA) database's GDCDataPortal (https: / / portal.gdc.cancer.gov), including mRNA abundance quantified by TPM, normalized long non-coding RNA (lncRNA) levels, log2-converted microRNA (miRNA) data, Illumina Human Methylation 450 microarray DNA methylation data, somatic mutation profiles, and corresponding clinical information (survival time, death status, tumor-lymph node-metastasis stage, treatment regimen, etc.). The original text did not specify a fixed sample size; ultimately, samples meeting the screening criteria were included (excluding patients with multiple primary cancers or those diagnosed solely by autopsy / death certificate).
[0086] mRNA expression data and clinical information of five hepatocellular carcinoma cohorts (GSE76427, GSE15654, GSE10143, GSE14520, and GSE116174) were downloaded from the Gene Expression Integrated Database (GEO). All data were obtained from microarray analysis to ensure consistency with gene names in the TCGA data. The five cohorts were integrated into a META cohort after batch effect correction.
[0087] The IMvigor210 cohort (hepatocellular carcinoma patients treated with programmed death-ligand 1 inhibitors) was obtained from a public dataset, including treatment response assessments (RECIST 1.1 criteria: CR / PR / SD / PD) and survival data; the three immunotherapy cohorts GSE78220, GSE135222, and GSE91061 were downloaded from the GEO database, all of which contain data on patients treated with immune checkpoint inhibitors.
[0088] 1.2 Data Preprocessing Operations
[0089] Batch effect elimination was performed on mRNA expression data from five GEO cohorts using the `sva` package in R. "Cohort origin" was used as the batch variable, and "survival status" as a covariate, achieving unified scale integration of cross-cohort data in the test set (i.e., batch correction combined with test set integration). Principal component analysis (PCA) validated the batch effect elimination, showing that samples from each cohort were evenly distributed in the principal component space. For numerical variables (e.g., mRNA expression level), missing values were imputed using the median of the variable in the corresponding cohort (missing value <20%). For categorical variables (e.g., tumor stage), "most frequent category imputed" was used (e.g., if T2 stage had the highest proportion in a cohort, T2 was used to impute missing values). Variables with a missing value >20% (e.g., some low-expression miRNAs or methylation sites) were directly removed. The "scale" function from the R language's base package was used to convert the mRNA, lncRNA, and miRNA data from the TCGA and GEO cohorts into a standard normal distribution (mean = 0, standard deviation = 1). For DNA methylation data, "β-normalization" (β = methylation signal intensity / (methylation signal intensity + non-methylation signal intensity + 100)) was used to ensure that the data range was between 0 and 1.
[0090] The inclusion criteria were as follows: ① Pathologically confirmed hepatocellular carcinoma; ② Complete survival time and death status information; ③ At least mRNA expression data; ④ Exclusion of patients with other malignant tumors or those diagnosed solely through autopsy or death certificates. The final sample size for the TCGA-LIHC cohort and the five GEO cohorts was determined based on actual screening, and the five GEO cohorts were integrated into a META cohort for subsequent validation.
[0091] Step 2: Cancer Subtype Identification
[0092] 2.1 Feature Filtering
[0093] Known immune-related genes were obtained from the ImmPort (https: / / www.immport.org) database and matched with mRNA expression data from the TCGA-LIHC cohort. Overlapping immune-related genes were identified as core features at the mRNA level (the exact number was not specified in the original text; the actual matching results prevail). For miRNA, lncRNA, and methylation data, features with the most significant expression variations were screened, retaining the top 2000 variable features for each. These screened features were further refined using univariate Cox regression analysis (with overall survival as the outcome, p<0.05) to identify prognostic features, resulting in a prognostic feature set for mRNA-immune-related genes, miRNA, lncRNA, and methylation sites. For somatic mutation data, the mutation frequency of each gene was calculated, and the top 5% of genes with the highest mutation frequencies were retained as mutation-level features.
[0094] By combining the above-mentioned mRNA-immune-related genes, miRNAs, lncRNAs, methylation sites, and prognostic features of somatic mutations, a multi-omics integrated feature set is formed for subsequent cluster analysis.
[0095] 2.2 Integration Clustering and Subtype Determination
[0096] Ten clustering algorithms were selected: K-means clustering, hierarchical clustering, spectral clustering, density clustering, fuzzy C-means clustering, mean-shift clustering, nearest-neighbor propagation clustering, hierarchical equilibrium iterative reduction clustering algorithm, agglomerative hierarchical clustering, and density-based cluster structure identification and ranking algorithm. Each algorithm was repeated 10 times to avoid random errors.
[0097] The cluster prediction index and gap statistic for each cluster size were calculated, and the results are as follows: Figure 2 As shown (it should be noted that, Figure 2 The blue dotted lines represent clustering prediction indices, and the red dotted lines represent gap statistics. When k=2, the clustering prediction index is the highest (0.82), and the gap statistics are the largest (0.32), indicating that the optimal number of clusters is 2. The clustering results of 10 algorithms were integrated, and the final subtype was determined by the "majority voting method": if a sample is classified into the same class in ≥6 algorithms, it is assigned to that class, resulting in cancer subtype I (212 cases) and cancer subtype II (153 cases).
[0098] 2.3 Subtype Validation and Characterization Analysis
[0099] Subtype stability was verified by recent template prediction and centroid-based partitioning clustering. Kappa coefficients were calculated (e.g., in the TCGA-LIHC cohort, the Kappa for cancer subtypes compared to recent template prediction was 0.613, and the Kappa for centroid-based partitioning clustering was 0.610, both p < 0.001). The verification was repeated in the META cohort to ensure subtype consistency.
[0100] Gene set variation analysis was used to assess the differences in pathway enrichment among subtypes. Metabolic pathways were enriched in cancer subtype I, and proliferation-related pathways were enriched in cancer subtype II (p<0.001). A transcriptional regulatory network was constructed to identify differentially activated transcription factors (such as AR and ESR1 activation in cancer subtype I, and FOXM1 and HIF1A activation in cancer subtype II) and chromatin remodeling factors (such as high expression of KAT2A and EHMT2 in cancer subtype II).
[0101] Step 3: IMLIRI Model Construction and Optimization
[0102] 3.1 Screening of immune-related genes
[0103] In the TCGA-LIHC training cohort, univariate Cox regression analysis was performed on immune-related genes at the mRNA level, with p < 0.001 used to screen for prognostic immune-related genes. In the five GEO validation cohorts, univariate Cox regression (p < 0.05) was also performed to screen for prognostic immune-related genes. The intersection of prognostic immune-related genes between the training and validation cohorts was taken, and all cohorts were required to have the same hazard ratio (hazard ratio > 1 or hazard ratio < 1) to determine the final set of candidate immune-related genes.
[0104] 3.2 Algorithm Combination Screening and Optimization
[0105] A machine learning framework comprising 10 basic algorithms (Cox regression based on boosting, random survival forest, logistic regression, support vector machine, K-nearest neighbors, neural network, extreme gradient boosting, gradient boosting decision tree, random forest, and adaptive boosting) was constructed. A combination of feature selection and survival prediction algorithms was used to generate 303 candidate models. Subsequently, all candidate combinations were systematically evaluated based on the 5-fold cross-validation consistency index on the training set, and the best-performing algorithm combination was ultimately selected as Cox regression combined with random survival forest. This combination demonstrated high accuracy and stable generalization ability on both the training and validation sets.
[0106] 3.3 Model Optimization
[0107] After selecting the optimal algorithm combination of Cox regression and random survival forest, this invention systematically optimizes this combination to further improve its prediction performance and generalization ability. The optimization process includes the following key steps, which are implemented simultaneously on the training set and five independent GEO validation sets to ensure robustness.
[0108] ① Select the optimal algorithm combination.
[0109] After comparing multiple algorithms, the optimal combination was determined to be "Cox regression + random survival forest", which was used as the unified object for subsequent system optimization and was implemented simultaneously on the training set and five independent GEO validation sets to ensure robustness.
[0110] ② Implement adaptive hyperparameter shrinkage optimization.
[0111] For the key hyperparameters of the two types of models, a wide initial interval is set. A first round of global coarse search is performed, using the 5-fold cross-validation consensus index (C-index) as the objective function to record performance. The initial results are sorted by C-index, and the top percentile "high-potential clusters" are extracted. Local weighted regression is used to fit the parameter-performance response surface, combined with density clustering to locate performance concentration areas. Subsequently, the parameter boundaries of these potential areas are automatically shrunk, and a high-resolution local dense search is conducted. In each round, the center and range of the potential areas are dynamically updated based on the latest performance until the convergence condition is met (e.g., C-index improvement < ε for two consecutive rounds or reaching the maximum number of iterations). This round-by-round shrinkage mechanism balances search accuracy and computational efficiency, improving the near-global optimal parameter capture rate.
[0112] ③ Introduce a two-way feedback mechanism between feature selection and prediction performance.
[0113] The candidate immune-related gene set is input into the joint training framework, employing alternating iterations: in each round, feature importance is extracted from both enhanced Cox regression and random survival forest (the former representing regression coefficient stability and selection frequency, the latter representing splitting gain / variable importance), and the overall contribution is calculated. If a gene has a significantly positive marginal contribution to the cross-validation C-index, its sampling weight or retention probability is increased in the next round; if the contribution is unstable or negative, its weight is gradually reduced or it is eliminated. Simultaneously, the model structure parameters are adaptively adjusted according to the current feature subset size and properties, forming a closed-loop optimization of "features—hyperparameters—performance," ultimately converging the feature set into a subset that highly matches the model and has the highest predictive value.
[0114] ④ Implement cross-algorithm performance consistency calibration.
[0115] Considering the differences between the two models in high-dimensional processing, sample sensitivity, and complexity, the performance score of the fusion output is uniformly scaled: first, the C-index of training and validation is normalized by sample size (performance estimate under equal sample distribution is obtained by bootstrapping), then penalty correction is applied based on the effective degrees of freedom and feature dimensions of the model, and finally the correction index is standardized by Z-score, so that the performance of different model stages and different algorithms can be fairly compared and merged on the same scale.
[0116] ⑤ Use multi-level evaluation criteria to screen the results of each round.
[0117] After each iteration, a three-level evaluation is performed: ① The average C-index of the 5-fold cross-validation of the training queue is used as the immediate optimization target; ② The average C-index of the five independent GEO validation sets is monitored to evaluate the generalization ability; ③ The calibrated cross-model consistency index is checked to ensure the fairness of the comparison and the rationality of the fusion.
[0118] In the training process of the hepatocellular carcinoma prognostic model, for the set of immune-related gene features obtained after each iteration, a prognostic model is constructed based on the TCGA-LIHC training cohort, and the corresponding average consistency index C-indexⅠ is calculated using five-fold cross-validation. When C-indexⅠ is not lower than the first performance threshold C1 (C1=0.60), and the change in C-indexⅠ compared to the previous iteration does not exceed the minimum performance decline threshold δ1 (δ1=0.03), or its improvement compared to the previous iteration reaches or exceeds the minimum performance improvement threshold δ2 (δ2=0.05), the prognostic model is further applied to five independent GEO external validation cohorts, and the corresponding average consistency index C-indexⅡ is calculated. When C-indexⅡ is not lower than the second performance threshold C2 (C2=0.80), and the five GEO... When the consistency index fluctuation between external validation queues is within a preset allowable range, the prediction results obtained by the model under different random initialization conditions or different algorithm combinations are calibrated, and the corresponding cross-model consistency index is calculated. When the cross-model consistency index reaches or exceeds the consistency threshold C3 (C3=0.85), the iteration is terminated and the final core immune gene set is output. Otherwise, according to the judgment result of the corresponding evaluation stage, the model structure hyperparameters and immune-related gene feature set are adjusted, and the next iteration is entered.
[0119] ⑥ Set the optimization stopping conditions and output the final model.
[0120] The optimization terminates when any of the following stopping conditions are met: no significant improvement on the independent validation set for k consecutive rounds (e.g., k=3) (below the preset threshold), or the C-index between training and validation tends to stabilize (fluctuation below the threshold), or the maximum number of iterations is reached; then the final feature subset and hyperparameter combination are locked to form the optimized joint model.
[0121] After final integration and robustness testing, the optimization results show that the hyperparameters and feature subsets based on the Cox regression + random survival forest combination have converged. The final model achieved a 5-fold cross-validation consistency index of 0.638 after average calibration during the training and validation phases, and exhibited small performance fluctuations in each independent GEO queue.
[0122] 3.4 Identification of Core Genes and Model Calculation
[0123] Candidate immune-related genes are input into a Cox regression algorithm, sorted based on feature importance scores (the frequency of gene selection in each iteration), and then selected as follows: Figure 3The 11 core immune-related genes shown are NROB1, MMP1, MPHOSPH6, FHL3, MAPK7, EEF1B2, RLF, PIK3IP1, PPP1R1A, GSTM1, and RGS7. The weight coefficients of these 11 core immune-related genes obtained using the Cox regression algorithm are shown in Table 1.
[0124] Table 1: Weight coefficients of 11 core immune-related genes
[0125]
[0126] The IMLRI scoring formula is as follows:
[0127] IMLIRI = (-0.099 × EEF1B2 expression level) + (-0.308 × FHL3 expression level) + (-0.087 × MMP1 expression level) + (-0.128 × MAPK7 expression level) + (-0.259 × MPHOSPH6 expression level) + (-0.248 × NROB1 expression level) + (0.190 × RLF expression level) + (0.203 × RGS7 expression level) + (0.434 × GSTM1 expression level) + (0.341 × PIK3IP1 expression level) + (0.273 × PPP1R1A expression level).
[0128] Among them, NROB1, MMP1, MPHOSPH6, MAPK7, and EEF1B2 are risk genes with positive coefficients, while FHL3, RLF, PIK3IP1, PPP1R1A, GSTM1, and RGS7 are protective genes with negative coefficients. The optimal cutoff value for IMLIRI was determined based on the overall survival of the TCGA-LIHC cohort, and patients were divided into low IMLIRI and high IMLIRI groups.
[0129] Step 4: Model Performance Evaluation and Clinical Application
[0130] 4.1 Comparison with clinical parameters and existing models
[0131] In the TCGA-LIHC and META cohorts, the concordance index between IMLIRI and traditional clinical parameters (age, sex, tumor stage, grade, etc.) and the area under the receiver operating characteristic curves (ROCs) for 1 / 3 / 5 years were calculated. Results are as follows: Figures 4-5As shown, the concordance index of IMLIRI (0.638) and the area under the receiver operating characteristic curve (both >0.65) were significantly higher than the clinical parameters. Univariate and multivariate Cox regression analyses confirmed that IMLIRI was an independent prognostic factor (hazard ratios of 5.610 and 5.703, respectively, both p<0.001).
[0132] The system searched the PubMed database and included 96 hepatocellular carcinoma prognostic models published since 2019. The consistency index of IMLIRI with these models was compared in TCGA-LIHC and 5 GEO cohorts. The results showed that IMLIRI's performance was among the best.
[0133] 4.2 Validation of Immunotherapy Response Prediction
[0134] like Figure 6 As shown, patients were divided into low-risk and high-risk groups according to their IMLIRI scores. The low-risk group had significantly better 12-month and 24-month limited survival, as well as significantly better 6-month long-term survival than the high-risk group (p = 0.051, <0.001, and 0.003, respectively). The treatment response rate in the low-risk group was significantly higher than that in the high-risk group (42.6% vs. 17.6%, p <0.001). In the GSE78220, GSE135222, and GSE91061 cohorts, the median survival and treatment response rate of patients in the low IMLIRI group were significantly higher. The low IMLIRI scores were significantly higher than those in the high-risk group (e.g., GSE78220 cohort p=0.013, GSE135222 cohort p<0.0001). Using the tumor immune phenotype tracking algorithm and the tumor immune dysfunction and rejection framework analysis, the low IMLIRI group showed higher activity in the immune processes of "antigen release-antigen presentation-immune cell infiltration-T cell recognition" and a higher probability of immunotherapy response (62.3% vs 21.8%, p<0.001); moreover, the low-risk group had higher microsatellite instability scores, while the high-risk group had higher scores for immune cell rejection and myeloid-derived suppressor cell infiltration.
[0135] Step 5: Nonograph Construction and Clinical Application
[0136] Integrating IMLIRI scores with clinical variables (tumor stage, grade, and presence of cirrhosis) to construct, as follows: Figure 7 The prognostic nomogram is shown; the 1-year, 3-year, and 5-year survival probabilities can be queried based on the total score of the nomogram. Internal validation was performed using 1000 samples, with mean absolute errors of <0.05 for the 1-year, 3-year, and 5-year calibration curves. External validation was performed in the GSE76427 cohort, with mean absolute errors of 0.038, 0.045, and 0.052, respectively, demonstrating the reliability of the predictions.
[0137] Clinical application example: A 65-year-old male patient with stage III (T3N1M0), grade G2, and cirrhosis was diagnosed with high IMLIRI risk based on core gene expression levels. Combined with the nomogram, the predicted 1-year survival probability was 68%. It was recommended to first perform transarterial chemoembolization to reduce the tumor burden, followed by combined treatment with a programmed death receptor-1 inhibitor to improve the response rate.
[0138] The above solution is merely an illustration of a preferred example and is not limited thereto. When implementing this invention, appropriate substitutions and / or modifications can be made according to the user's needs.
[0139] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. It can be applied to various fields suitable for the present invention. Other modifications can be readily made by those skilled in the art. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and examples shown and described herein.
Claims
1. A method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning, characterized in that, include: S1. Collect multi-omics data related to hepatocellular carcinoma from the Cancer Genome Atlas Database (TCGA) to form the TCGA-LIHC cohort, collect external validation cohorts from the Gene Expression Comprehensive Database (GEO), and combine them with immunotherapy-specific cohorts to construct datasets for model training and validation. S2. Perform feature integration and cluster analysis on multi-omics data to obtain the corresponding molecular subtype distribution of hepatocellular carcinoma, and determine whether there is data bias in the TCGA-LIHC cohort based on the molecular subtype distribution of hepatocellular carcinoma. S3. Construct a prognostic model for hepatocellular carcinoma based on Cox regression combined with random survival forest, and identify 11 core immune genes and their corresponding weight coefficients through model training, constructing the following immunotherapy response index, IMLIRI score: IMLIRI = (-0.099 × EEF1B2 expression level) + (-0.308 × FHL3 expression level) + (-0.087 × MMP1 expression level) + (-0.128 × MAPK7 expression level) + (-0.259 × MPHOSPH6 expression level) + (-0.248 × NROB1 expression level) + (0.190 × RLF expression level) + (0.203 × RGS7 expression level) + (0.434 × GSTM1 expression level) + (0.341 × PIK3IP1 expression level) + (0.273 × PPP1R1A expression level); S4. The IMLIRI score given by the hepatocellular carcinoma prognostic model was validated in multiple dimensions using an external validation cohort and an immunotherapy-specific cohort. In S1, the multi-omics data includes: mRNA abundance quantified by TPM, normalized long non-coding RNA levels, log2-converted microRNA data, Illumina Human Methylation 450 chip DNA methylation data, somatic mutation profile, survival time, death status, tumor-lymph node-metastasis staging, and treatment regimen. During the training of the hepatocellular carcinoma prognostic model, the screening results after each iteration were evaluated using a three-level assessment criterion: The average consistency index C-indexⅠ of the 5-fold cross-validation of the TCGA-LIHC queue is used as the immediate optimization target to perform a first-level evaluation on the screening results after each iteration. When the average C-indexⅠ is not lower than the preset performance threshold C1, and when the change of C-indexⅠ compared with the previous iteration result does not exceed the preset minimum performance decline threshold δ1, or when its improvement reaches or exceeds the preset minimum performance improvement threshold δ2, the first-level evaluation is deemed to have passed. Otherwise, the model structure hyperparameters are adaptively adjusted, and the next iteration is entered. The generalization ability of the average consistency index C-index II in the five independent GEO validation queues is evaluated in a secondary evaluation. If C-index II is not lower than the preset generalization threshold C2 and the performance fluctuation between the validation queues does not exceed the preset allowable range, the secondary evaluation is deemed to pass; otherwise, the model is retrained. The cross-model consistency index after calibration is checked to achieve a three-level evaluation. When the consistency index reaches or exceeds the preset consistency threshold C3, the three-level evaluation is deemed to have passed, the iteration is terminated and the final core immune gene set is output. Otherwise, joint optimization is achieved through alternating iteration.
2. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 1, characterized in that, The external validation queue includes GSE76427, GSE15654, GSE10143, GSE14520, and GSE116174 downloaded from GEO. The relevant data information in the external validation queue is consistent with the multi-omics data content, and the external validation queue is integrated into the META queue after batch effect correction. The immunotherapy-specific cohorts include: the IMvigor210 cohort and GSE78220, GSE135222, and GSE91061 downloaded from GEO. The IMvigor210 cohort consists of hepatocellular carcinoma patients receiving programmed death-ligand 1 inhibitor therapy, while the GSE78220, GSE135222, and GSE91061 cohorts contain data from patients receiving immune checkpoint inhibitor therapy. The inclusion criteria for the TCGA-LIHC cohort and external validation cohort include: pathological diagnosis of hepatocellular carcinoma, complete survival time and death status information, at least mRNA abundance expression data, and exclusion of patients with other malignant tumors or those diagnosed solely through autopsy or death certificates.
3. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 2, characterized in that, In S1, the batch effect correction is performed on the mRNA abundance expression data of the external validation cohort, with the cohort source as the batch variable and the survival status as the covariate to eliminate batch effects. The correction content for the mRNA abundance expression data includes: For numerical variables similar to mRNA abundance, missing values are filled using the median of the variable in the corresponding cohort; For categorical variables similar to tumor staging, the most frequent category is used for imputation; Variables with a missing rate >20% are removed directly.
4. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 1, characterized in that, In S2, the distribution of hepatocellular carcinoma molecular subtypes is obtained as follows: S20. Feature screening is performed on the data information in the TCGA-LIHC cohort to screen for immune-related gene features from mRNA expression data, to screen for the top-m prognostic features I before expression mutation from miRNA, lncRNA and methylation data, and to screen for the top-n prognostic features II before mutation frequency from somatic mutation data. The gene features, prognostic features I and prognostic features II are then combined to obtain a multi-omics integrated feature set. S21. Multiple unsupervised clustering algorithms are used to perform integrated clustering analysis on the multi-omics integrated feature set. The clustering prediction index and gap statistics are combined to determine cancer subtype I and cancer subtype II. The majority voting method is used to determine the final cancer subtype for the clustering results of multiple unsupervised clustering algorithms. S22. Evaluate the data offset of the TCGA-LIHC cohort through subtype verification and characterization analysis; The classification reliability in the subtype verification is achieved through the nearest template prediction and the centroid-based partitioning algorithm. When the Kappa coefficient is greater than 0.6, the classification reliability is proven. The subtype molecular distribution differences in the subtype verification are achieved by using principal component analysis, t-distribution random neighborhood embedding, unified manifold approximation and projection dimensionality reduction visualization. The differences in pathway enrichment among subtypes in the characterization analysis were achieved through gene set variation analysis, in which cancer subtype I was enriched with metabolic pathways and cancer subtype II was enriched with proliferation-related pathways. The characterization analysis also includes: analyzing the tumor microenvironment by calculating the immune cell infiltration score and immune characteristic score, and identifying cancer subtype I as immune-activated and cancer subtype II as immunosuppressive.
5. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 1, characterized in that, In S3, the hepatocellular carcinoma prognostic model is selected from 303 algorithm combinations generated by the basic algorithm and the ensemble strategy, and the optimal algorithm combination of Cox regression and random survival forest is selected. The integrated strategy includes: a primary evaluation index based on the average consistency index of the training queue and the validation queue, and auxiliary indicators covering the area under the curve and the p-value of the log-rank test.
6. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 1, characterized in that, In S3, the 11 core immune genes are obtained as follows: Univariate Cox regression was performed in the S30, TCGA-LIHC cohort and external validation cohort to screen for candidate immune-related gene sets that are associated with prognosis and have consistent hazard ratios across all cohorts. S31. Input the candidate immune-related gene set into the joint training framework of Cox regression and random survival forest, and use alternating iteration combined with adaptive adjustment of model structure hyperparameters to construct a closed-loop optimization of feature selection-hyperparameter optimization-survival prediction to promote the convergence of the candidate immune-related gene set, so as to obtain the corresponding 11 core immune genes after convergence.
7. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 6, characterized in that, The method for adaptively adjusting the hyperparameters of the model structure is as follows: For the key hyperparameters of the two types of models, a wide initial interval is set. A first round of global coarse search is performed, and the performance is recorded with the 5-fold cross-validation consistency index C-indexⅠ as the objective function. The initial results were sorted according to C-index I, and the high-potential clusters in the top few percentiles were extracted. The performance response surface was fitted by local weighted regression and combined with density clustering to locate the performance concentration area as the potential area. The hyperparameters are adaptively shrunk and optimized by automatically shrinking the parameter boundaries of the potential region and conducting a high-resolution local dense search. In each round, the center and range of the potential region are dynamically updated based on the latest performance until the convergence condition is met.
8. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 6, characterized in that, The alternating iteration involves extracting the importance index of features from the Cox regression combined with the random survival forest in each round to calculate the overall contribution. When the marginal contribution of any gene to the consistency index C-index is positive during the cross-validation process and exceeds the preset minimum contribution threshold, the sampling weight or retention probability of the corresponding gene pair is increased in the next round; otherwise, the gene pair is gradually deweighted or eliminated. The importance index refers to the stability and selection frequency of regression coefficients in the Cox regression algorithm, and the split gain or variable importance in the random survival forest algorithm.
9. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 1, characterized in that, In S3, the optimal cutoff value for IMLIRI was determined based on the overall survival of the TCGA-LIHC cohort, and patients were divided into low IMLIRI and high IMLIRI groups.
10. The method for predicting the prognosis and immune response of hepatocellular carcinoma based on multi-omics machine learning as described in claim 1, characterized in that, In S4, the multi-dimensional verification includes: Dimension I: Compare and validate the IMLIRI score given by the hepatocellular carcinoma prognostic model with clinical parameters and existing models using the consistency index. Dimension II: The IMLIRI scores given by the hepatocellular carcinoma prognostic model were divided into low-risk and high-risk groups. The tumor immune phenotype tracking algorithm and the tumor immune dysfunction and rejection framework were used to analyze and validate the low-risk and high-risk groups in terms of immunotherapy response probability, immune cell rejection, and myeloid-derived suppressor cell infiltration.
Citation Information
Patent Citations
Model for predicting curative effect of hepatocellular carcinoma immunotherapy and construction method thereof
CN112614546A
Method for constructing hepatocellular carcinoma typing system based on ferroptosis process
CN113192560A
Hepatocellular carcinoma prognosis prediction system based on multiple omics characteristics and prediction method thereof
CN114678062A
Hepatocellular carcinoma prognosis scoring model construction method and device, equipment and storage medium
CN117936111A