A Method for Constructing a Predictive Model for Immunotherapy Response in Hepatocellular Carcinoma

CN122575750APending Publication Date: 2026-08-14NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]因此,本发明提供了一种肝细胞癌免疫治疗响应预测模型构建方法解决现有技术因忽略肿瘤微环境空间异质性而难以精准捕捉决定免疫治疗响应的关键空间生态位特征,从而导致预测效能受限的问题

Benefits of technology

基于 Niche_9 的细胞组成、空间分布及 MEI1 的分子特征所筛选的核心预测基因集,构成了对 T+A 方案疗效预测的关键生物学基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575750A_ABST
    Figure CN122575750A_ABST
Patent Text Reader

Abstract

This invention discloses a method for constructing a predictive model for immunotherapy response in hepatocellular carcinoma, relating to the field of bioinformatics technology. The method includes: multi-omics data acquisition and integration: acquiring BulkRNA-seq data, single-cell RNA-seq (scRNA-seq) data, and spatial transcriptome (ST) data from hepatocellular carcinoma patients; performing quality control and batch effect removal on the scRNA-seq data, and performing slice alignment and pathological region annotation on the ST data; spatial niche identification: using the cell2location algorithm to perform cell deconvolution on the ST data based on the scRNA-seq reference map to obtain a spatial cell abundance matrix; and performing cluster analysis based on the abundance matrix to identify spatial niches related to immunotherapy response. This invention utilizes the XGBoost algorithm to construct and optimize the predictive model, and also discloses its application in the preparation of companion diagnostic kits and guiding clinical medication decisions for T+A regimens. The core also involves the characteristics and functions of the key biomarker MEI1.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics technology, and in particular to a method for constructing a predictive model for the response to immunotherapy in hepatocellular carcinoma. Background Technology

[0002] Hepatocellular carcinoma (HCC), a malignant tumor with a high mortality rate worldwide, is undergoing a paradigm shift in its treatment from traditional chemotherapy to immuno-oncology. In recent years, immune checkpoint inhibitors, represented by anti-PD-1 / PD-L1 inhibitors, have demonstrated sustained clinical benefits in the treatment of advanced HCC. However, the limitations of objective response rates and the existence of immune-related adverse events highlight the clinical urgency of constructing accurate predictive models. Current technological approaches primarily rely on bulk transcriptome data obtained through high-throughput sequencing, screening for immune-related gene sets through differential expression analysis, and constructing predictive models using machine learning algorithms (such as LASSO regression and support vector machines). While these methods achieve preliminary risk stratification at the population level, their fundamental assumption is that tissue samples are treated as homogeneous cell collections, neglecting the spatial heterogeneity and functional synergistic mechanisms of different cell subpopulations within the HCC microenvironment. This lack of spatial information may obscure key biological events determining the response to immunotherapy.

[0003] Existing technologies struggle to capture the spatial niches formed by immune cells, tumor cells, and stromal cells at specific anatomical sites. The composition and state of these niches directly determine the activation efficiency of immune checkpoint inhibitors and the persistence of anti-tumor immune responses. To address this technological gap, this invention proposes a method for constructing a predictive model for hepatocellular carcinoma (HCC) immunotherapy response that integrates multi-omics data. By integrating single-cell resolution transcriptome mapping with spatial transcriptomics techniques, the cell2location algorithm is used to perform cell deconvolution on ST data, accurately resolving the cell abundance distribution at different spatial sites. Cluster analysis is then used to identify key niches significantly correlated with immunotherapy response, and their specific hypervariable genes are extracted as candidate feature sets. Finally, the XGBoost algorithm and SHAP value interpretability analysis are combined to construct a high-precision predictive model in bulk data. This approach not only overcomes the limitations of traditional methods in spatial dimension resolution but also provides a combination of biomarkers with clear pathological significance for the accurate prediction of HCC immunotherapy. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a method for constructing a predictive model for hepatocellular carcinoma immunotherapy response to solve the problem that existing technologies, by ignoring the spatial heterogeneity of the tumor microenvironment, are unable to accurately capture the key spatial niche characteristics that determine the immunotherapy response, thus resulting in limited predictive efficacy.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for constructing a predictive model for immunotherapy response in hepatocellular carcinoma, comprising: acquiring and integrating multi-omics data: acquiring BulkRNA-seq data, single-cell RNA-seq (scRNA-seq) data, and spatial transcriptome (ST) data from hepatocellular carcinoma patients; performing quality control and batch effect removal on the scRNA-seq data; and performing slice alignment and pathological region annotation on the ST data. Spatial niche identification: The cell2location algorithm was used to perform cell deconvolution on ST data based on scRNA-seq reference maps to obtain a spatial cell abundance matrix; based on the abundance matrix, cluster analysis was performed to identify spatial niches related to immunotherapy response; Feature gene screening: Highly upregulated genes in the key spatial niches are extracted as a candidate feature gene set; Model construction: Based on the candidate feature gene set, an immunotherapy response prediction model was constructed in BulkRNA-seq data using the XGBoost algorithm, and core prediction features were screened using SHAP value analysis.

[0007] As a preferred embodiment of the method for constructing a predictive model for hepatocellular carcinoma immunotherapy response according to the present invention, the spatial niche identification step specifically includes: Using scRNA-seq data as a priori reference, the cell2location Bayesian hierarchical model was trained for 30,000 epochs to infer the cellular composition of each spot in the spatial transcriptome. Leiden clustering (Resolution=1.0) and hierarchical clustering were performed on the cell abundance matrix obtained by deconvolution to divide the tumor microenvironment into 9 consensus niches (Niche_1 to Niche_9). Niche_9 was identified as a niche significantly associated with the response to atezolizumab + bevacizumab (T+A) treatment using GSVA scores and chi-square tests.

[0008] Furthermore, in the spatial niche identification process, single-cell transcriptome (scRNA-seq) data is first used as a reference map and input into the cell2location Bayesian hierarchical model. The model is trained for 30,000 epochs to accurately infer the proportion of different cell types in each micro-spot of the spatial transcriptome. After obtaining the cell abundance matrix, both Leiden clustering (resolution set to 1.0) and hierarchical clustering are used to comprehensively divide the tumor microenvironment into nine common spatial niches, numbered from Niche_1 to Niche_9. Then, statistical association analysis is performed using GSVA enrichment scores and chi-square tests to select the only niche significantly associated with the treatment response to atezolizumab combined with bevacizumab (T+A regimen), namely Niche_9.

[0009] As a preferred embodiment of the method for constructing a predictive model for hepatocellular carcinoma immunotherapy response according to the present invention, wherein: Niche_9 is a TLS-like (tertiary lymphoid structure-like) spatial niche, characterized by including the following cellular composition and molecular features: Cellular composition: enriched with B cells, plasma cells, NK cells, Naïve T cells, Treg cells, and IL-4I1. + Macrophages; Spatial distribution: This niche is spatially characterized by close co-localization of B cells and T cells, and is located around or distributed in the tumor area; Functional characteristics: accompanied by enhanced TLS-related signaling, lymphocyte co-stimulation and adaptive immune activation pathways, as well as unique fatty acid metabolism activity.

[0010] Furthermore, the identified Niche_9 is essentially a tertiary lymphoid structure (TLS-like) spatial niche. In terms of cellular composition, it is significantly enriched with B cells, plasma cells, NK cells, naive T cells, regulatory T cells (Tregs), and IL4I1-positive macrophages. Spatially, the most typical characteristic of this niche is the highly co-localization of B cells and T cells, which are either surrounding the tumor region or directly distributed within the tumor. Functionally, it not only exhibits overall upregulation of TLS-related signaling pathways, lymphocyte co-stimulation pathways, and adaptive immune activation pathways, but also possesses unique and active fatty acid metabolism characteristics, which is a key marker distinguishing it from other niches.

[0011] As a preferred embodiment of the method for constructing a hepatocellular carcinoma immunotherapy response prediction model according to the present invention, the feature gene screening step specifically includes: The top 100 highly upregulated genes in the Niche_9 niche were extracted as candidate features; Genes playing key communication roles within their ecological niches were screened using ligand-receptor interaction analysis (CellPhoneDB) and spatial colocalization analysis (MistyR). The feature importance of the XGBoost model was analyzed using SHAP (Shapley Additive ex Planations) values, and MEI1, CD8B, CXCR4, GZMB, CD8A, CD19, BLK, IRF4 and CTLA4 were finally selected as the core prediction gene set.

[0012] Furthermore, in the feature gene screening stage, we first extracted the top 100 genes with the highest expression variation and most significant upregulation from the Niche_9 niche as initial candidate features. To further identify genes that truly play a key role within the niche, we combined ligand-receptor interaction analysis (CellPhoneDB) and spatial colocalization analysis (MistyR) to screen out genes that are only highly expressed individually but do not participate in intercellular communication. Finally, the remaining genes were put into the XGBoost model, and the feature importance of each gene was calculated using the SHAP additive interpretation value. After sorting and screening, the following nine genes were finally identified as the core prediction gene set: MEI1, CD8B, CXCR4, GZMB, CD8A, CD19, BLK, IRF4, and CTLA4.

[0013] As a preferred embodiment of the method for constructing a predictive model for hepatocellular carcinoma immunotherapy response according to the present invention, wherein: the MEI1 gene in the core gene set serves as a key biomarker, characterized in that: Specific expression: The MEI1 gene is specifically highly expressed in a plasma cell subset (i.e., MEI1). + Plasmacell); Functional correlation: MEI1 expression was significantly positively correlated with TLS maturity score, HEV characteristics, and better overall survival (OS); Molecular mechanism: MEI1 + Plasma cells exhibit enhanced endoplasmic reticulum protein processing and immunoglobulin synthesis, and are associated with IL4I1. + Macrophages exhibit active APOE–LDLR metabolic communication.

[0014] Furthermore, MEI1 is the most specific key biomarker. Firstly, in terms of expression specificity, MEI1 is highly expressed almost exclusively in plasma cell subsets, forming a unique MEI1... +Plasma cell populations; from a clinical functional perspective, MEI1 expression levels were significantly positively correlated with TLS maturity scores and high endothelial venule (HEV) characteristics, and were also closely associated with longer overall survival (OS); at the molecular mechanism level, MEI1... + Plasma cells exhibit a phenotype characterized by enhanced endoplasmic reticulum protein processing capacity and significantly upregulated immunoglobulin synthesis. Furthermore, these plasma cells also interact with IL-4I1. + Macrophages form active metabolic communications through the APOE–LDLR axis, working together to maintain the function of the immune microenvironment.

[0015] As a preferred embodiment of the method for constructing a predictive model for hepatocellular carcinoma immunotherapy response according to the present invention, the model construction step specifically includes: The BulkRNA-seq data were divided into training and validation sets at a ratio of 70% / 30%. Set the hyperparameters of the XGBoost classifier as follows: objective function is binary:logistic, maximum depth is 4, and learning rate is 0.1. A manual early stopping mechanism is used to prevent overfitting, and the optimal prediction model is obtained by using ROC-AUC as the evaluation index.

[0016] Furthermore, the bulk RNA-seq data was randomly divided into a 70% training set and a 30% validation set to ensure consistent data distribution between the two sets. Then, an XGBoost classifier was built, with key hyperparameters explicitly set: the objective function was binary:logistic for binary classification prediction, the maximum decision tree depth (max_depth) was set to 4, and the learning rate (learning_rate) was set to 0.1. During training, a manual early stopping strategy was employed, monitoring changes in the ROC-AUC metric on the validation set. Training was stopped when the metric ceased to improve, thus preventing overfitting and ultimately obtaining a stable and well-generalized optimal prediction model.

[0017] As a preferred embodiment of the method for constructing a prediction model for hepatocellular carcinoma immunotherapy response according to the present invention, it further includes: an application method based on the prediction model, specifically including: Obtain gene expression profile data from the patients to be tested; The GSVA algorithm was used to calculate the patient's core gene score; Patients were divided into high-score and low-score groups based on the median score. If a patient belongs to the high-score group, they are predicted to respond well to T+A combination therapy and have a low TIDE immune escape score and a high level of immune cell infiltration.

[0018] Further, it is necessary to obtain tissue or blood gene expression profile data of patients with hepatocellular carcinoma to be tested; then, the GSVA algorithm is used to calculate the core gene enrichment score for each patient based on the core gene set selected earlier; then, the patients are divided into high-score and low-score groups using the median score as the cutoff; if a patient is classified into the high-score group, it can be predicted that they are more likely to have a good response to T+A combined immunotherapy. At the same time, these patients usually also have lower TIDE immune escape scores and higher levels of immune cell infiltration, indicating that the overall immune microenvironment is more active.

[0019] As a preferred embodiment of the method for constructing a predictive model for hepatocellular carcinoma immunotherapy response described in this invention, the application method is specifically used to prepare a companion diagnostic kit for hepatocellular carcinoma immunotherapy, or to guide hepatocellular carcinoma patients in making clinical medication decisions for the T+A regimen (artezizumab combined with bevacizumab).

[0020] Furthermore, based on the expression detection of core genes, companion diagnostic kits for hepatocellular carcinoma immunotherapy can be developed and prepared to enable rapid screening of patients who will benefit from treatment. On the other hand, they can be directly used for clinical decision support, predicting the probability of treatment response in advance before hepatocellular carcinoma patients receive atezolizumab combined with bevacizumab, helping clinicians to more accurately select patients suitable for the T+A regimen, avoid ineffective treatment, and achieve personalized immunotherapy.

[0021] As a preferred embodiment of the method for constructing a predictive model for hepatocellular carcinoma immunotherapy response according to the present invention, the predictive model identifies a specific spatial niche, Niche_9, that is significantly associated with the treatment response to atezolizumab combined with bevacizumab (T+A), and the MEI1 gene, which is specifically highly expressed in this niche. Niche_9 is a tertiary lymphoid structure-like (TLS-like) niche, composed of B cells, plasma cells, NK cells, Naïve T cells, Treg cells, and IL4I1. + Macrophages are a specific component of cells and exhibit close co-localization of B cells and T cells in space. The MEI1 gene is specifically expressed in a plasma cell subset (MEI1) within Niche_9. + Plasma cells), whose expression levels were significantly positively correlated with TLS maturation scores, HEV characteristics, and overall patient survival, and were also associated with IL4I1. + Macrophages exhibit active APOE–LDLR metabolic communication; The core predictive gene set screened based on the cellular composition and spatial distribution of Niche_9 and the molecular characteristics of MEI1 constitutes the key biological basis for predicting the efficacy of the T+A regimen.

[0022] The beneficial effects of this invention are as follows: it acquires and integrates multidimensional omics data of hepatocellular carcinoma patients and performs preprocessing, then uses relevant algorithms to perform cell deconvolution and cluster analysis on spatial transcriptome data based on single-cell RNA-seq reference maps, identifies TLS-like spatial niches related to T+A treatment response, extracts hypervariable upregulated genes in these niches as candidate features, combines multiple analyses to screen out a core predictive gene set, uses the XGBoost algorithm to construct and optimize a prediction model, and discloses the application of this model in the preparation of companion diagnostic kits and in guiding clinical drug decisions for T+A regimens. The core also involves the characteristics and functions of the key biomarker MEI1. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating the method for constructing a predictive model for the response to immunotherapy in hepatocellular carcinoma. Detailed Implementation

[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0026] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0027] Secondly, the term "one embodiment" or "example" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the invention. The appearance of an embodiment in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that mutually excludes other embodiments.

[0028] Reference Figure 1 This is one embodiment of the present invention, which provides a method for constructing a prediction model for hepatocellular carcinoma immunotherapy response, comprising the following steps: Example 1 1. Acquisition and integration of multidimensional omics data Bulk RNA-seq data, single-cell RNA-seq (scRNA-seq) data, and spatial transcriptome (ST) data were collected from 50 patients with hepatocellular carcinoma. The scRNA-seq data underwent quality control to remove cells with excessively high mitochondrial gene content and insufficient sequencing depth, and batch effect removal was performed using Seurat software. The ST data were sliced ​​and aligned, and pathological region annotation was completed based on HE staining results to distinguish tumor regions from normal liver tissue regions.

[0029] 2. Spatial niche identification Using preprocessed scRNA-seq data as a priori reference, the cell2location Bayesian hierarchical model was trained for a specified number of rounds to infer the cellular composition of each spot in the spatial transcriptome, thus obtaining a spatial cell abundance matrix. Leiden clustering and hierarchical clustering were performed on this abundance matrix to divide the tumor microenvironment into 9 consensus niches (Niche_1 to Niche_9). Niche_9, which was significantly associated with the response to atezolizumab + bevacizumab (T+A) treatment, was selected through GSVA scoring and chi-square test. Niche_9 is a TLS-like spatial niche, enriched with B cells, plasma cells, NK cells, etc. B cells and T cells are closely co-localized, accompanied by enhanced TLS-related signals and adaptive immune activation pathways.

[0030] 3. Feature gene screening The top 100 highly upregulated genes in the Niche_9 niche were extracted as candidate features. Genes playing key communication roles within the niche were screened using ligand-receptor interaction analysis (CellPhoneDB) and spatial colocalization analysis (MistyR). The feature importance of the XGBoost model was analyzed using SHAP values, and MEI1, CD8B, CXCR4, GZMB, CD8A, CD19, BLK, IRF4, and CTLA4 were finally selected as the core predictive gene set. Among them, MEI1 was specifically highly expressed in plasma cell subsets, and its expression was significantly positively correlated with TLS maturation score and overall patient survival.

[0031] 4. Model Building BulkRNA-seq data were divided into training and validation sets proportionally. The hyperparameters of the XGBoost classifier were set as follows: objective function: binary:logistic, maximum depth: 4, and learning rate: 0.1. A manual early stopping mechanism was used to prevent overfitting. The optimal prediction model was obtained by using ROC-AUC as the evaluation metric.

[0032] 5. Model Application Gene expression profiles of 10 patients with hepatocellular carcinoma were obtained; core gene scores were calculated using the GSVA algorithm; patients were divided into high-score and low-score groups based on the median score; patients in the high-score group were predicted to respond well to T+A combination therapy and had lower TIDE immune escape scores and higher levels of immune cell infiltration.

[0033] Example 2 1. Acquisition and integration of multidimensional omics data BulkRNA-seq, scRNA-seq, and ST data were collected from 80 patients with hepatocellular carcinoma, including 40 patients in each of the T+A treatment response and non-response groups. Strict quality control was performed on the scRNA-seq data; after removing abnormal cells, the Harmony algorithm was used for batch effect removal to improve cell clustering accuracy. After slice alignment of the ST data, a deep learning algorithm was used to assist in pathological region annotation, reducing errors from manual annotation.

[0034] 2. Spatial niche identification Following the same spatial niche identification steps as in Example 1, Niche_9, which was significantly associated with T+A treatment response, was screened out. In addition, niche function verification was performed by immunohistochemical experiments to verify the distribution characteristics of B cells and plasma cells in Niche_9, confirming its TLS-like structural properties and immune activation function.

[0035] 3. Feature gene screening Following the same characteristic gene screening steps as in Example 1, a core predicted gene set was obtained; a core gene verification step was added, using qPCR technology to detect the expression level of the core gene in patient tissue samples, verifying the specific expression of the MEI1 gene in plasma cells and its association with treatment response, removing genes with poor expression stability, and further optimizing the core gene set.

[0036] 4. Model Building BulkRNA-seq data were divided into training and validation sets proportionally, and 5-fold cross-validation was used to optimize the hyperparameters of the XGBoost classifier. Based on the manual early stopping mechanism, an L1 regularization term was added to further reduce the risk of model overfitting. ROC-AUC, accuracy, and recall were used as comprehensive evaluation indicators to obtain the optimal prediction model with stronger generalization ability.

[0037] 5. Model Application The model application steps are the same as in Example 1; additionally, clinical validation of the model is added, the prediction results are compared with the actual treatment effects of patients, the model prediction accuracy is calculated, and the core gene score is correlated with the patient's progression-free survival to further verify the clinical applicability of the model and provide a more reliable basis for clinical medication decisions.

[0038] Comparative Example 1 The difference between this comparative example and Example 1 is that the spatial niche identification step is omitted, and the model is constructed directly based on scRNA-seq data to screen for highly variable genes. The specific steps are as follows: 1. Acquisition and integration of multidimensional omics data The steps for acquiring and integrating multidimensional omics data are the same as in Example 1.

[0039] 2. Feature gene screening The top 100 highly upregulated genes from tumor cells in scRNA-seq data were directly extracted as a candidate feature gene set. Spatial niche identification and niche gene screening were not performed. The feature importance of the XGBoost model was directly analyzed using SHAP values ​​to screen the core prediction gene set.

[0040] 3. Model Construction and Application The model construction and application steps are the same as in Example 1.

[0041] Because this comparative example did not identify the spatial niche associated with treatment response and the core gene set did not incorporate the spatial distribution characteristics of the tumor microenvironment, it could not accurately capture cell interactions and functional characteristics related to immunotherapy response. As a result, the model prediction accuracy and ROC-AUC value were significantly lower than those in Example 1, and it could not effectively distinguish between T+A treatment-responsive and non-responsive patients.

[0042] Comparative Example 2 The difference between this comparative example and Example 1 is that the characteristic gene screening step does not use ligand-receptor interaction analysis and SHAP value analysis, but only uses conventional differential expression analysis to screen core genes. The specific steps are as follows: 1. Acquisition and integration of multidimensional omics data, and identification of spatial ecological niches. Following the same steps as in Example 1, the Niche_9 niche was selected.

[0043] 2. Feature gene screening The top 100 highly upregulated genes in the Niche_9 niche were extracted as a candidate feature gene set. Conventional differential expression analysis (t-test) was used to screen genes associated with T+A treatment response. Ligand-receptor interaction analysis and SHAP value analysis were not performed. The screened differentially expressed genes were directly used as the core predictive gene set.

[0044] 3. Model Construction and Application The model construction and application steps are the same as in Example 1.

[0045] This comparative model failed to screen key communication genes within the niche through ligand-receptor interaction analysis and failed to clarify the predictive importance of genes through SHAP value analysis. As a result, there were redundant genes in the core gene set, and some genes were not directly related to the treatment response. This increased the risk of model overfitting and made it impossible to accurately screen patients who benefited from T+A treatment.

[0046] This embodiment also provides a computer device applicable to the method for constructing a prediction model for hepatocellular carcinoma immunotherapy response, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for constructing a prediction model for hepatocellular carcinoma immunotherapy response as proposed in the above embodiment.

[0047] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0048] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0049] In summary, this invention acquires and integrates multi-omics data from hepatocellular carcinoma patients and preprocesses them. Then, it utilizes relevant algorithms to perform cellular deconvolution and clustering analysis on spatial transcriptome data based on single-cell RNA-seq reference maps to identify TLS-like spatial niches associated with T+A treatment response. Subsequently, it extracts hypervariable upregulated genes in these niches as candidate features, combines multiple analyses to screen out a core predictive gene set, and uses the XGBoost algorithm to construct and optimize a predictive model. The invention also discloses the application of this model in the preparation of companion diagnostic kits and in guiding clinical medication decisions for T+A regimens. The core of the invention also involves the characteristics and functions of the key biomarker MEI1.

[0050] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for constructing a predictive model for immunotherapy response in hepatocellular carcinoma, characterized in that: Includes the following steps: Multidimensional omics data acquisition and integration: Acquiring BulkRNA-seq data, single-cell RNA-seq (scRNA-seq) data, and spatial transcriptome (ST) data from hepatocellular carcinoma patients; The scRNA-seq data were subjected to quality control and batch effect removal, and the ST data were subjected to slice alignment and pathological region annotation. Spatial niche identification: The cell2location algorithm was used to perform cell deconvolution on ST data based on scRNA-seq reference maps to obtain a spatial cell abundance matrix; based on the abundance matrix, cluster analysis was performed to identify spatial niches related to immunotherapy response; Feature gene screening: Highly upregulated genes in the key spatial niches are extracted as a candidate feature gene set; Model construction: Based on the candidate feature gene set, an immunotherapy response prediction model was constructed in BulkRNA-seq data using the XGBoost algorithm, and core prediction features were screened using SHAP value analysis.

2. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 1, characterized in that: The spatial niche identification step specifically includes: Using scRNA-seq data as a priori reference, the cell2location Bayesian hierarchical model was trained for 30,000 epochs to infer the cellular composition of each spot in the spatial transcriptome. Leiden clustering (Resolution=1.0) and hierarchical clustering were performed on the cell abundance matrix obtained by deconvolution to divide the tumor microenvironment into 9 consensus niches (Niche_1 to Niche_9). Niche_9 was identified as a niche significantly associated with the response to atezolizumab + bevacizumab (T+A) treatment using GSVA scores and chi-square tests.

3. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 2, characterized in that: The Niche_9 is a TLS-like (tertiary lymphoid structure) spatial niche, characterized by the following cellular composition and molecular features: Cellular composition: enriched with B cells, plasma cells, NK cells, Naïve T cells, Treg cells, and IL-4I1. + Macrophages; Spatial distribution: This niche is spatially characterized by close co-localization of B cells and T cells, and is located around or distributed in the tumor area; Functional characteristics: accompanied by enhanced TLS-related signaling, lymphocyte co-stimulation and adaptive immune activation pathways, as well as unique fatty acid metabolism activity.

4. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 2, characterized in that: The feature gene screening step specifically includes: The top 100 highly upregulated genes in the Niche_9 niche were extracted as candidate features; Genes playing key communication roles within their ecological niches were screened using ligand-receptor interaction analysis (CellPhoneDB) and spatial colocalization analysis (MistyR). The feature importance of the XGBoost model was analyzed using SHAP (Shapley Additive ex Planations) values, and MEI1, CD8B, CXCR4, GZMB, CD8A, CD19, BLK, IRF4 and CTLA4 were finally selected as the core prediction gene set.

5. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 4, characterized in that: The MEI1 gene in the core gene set serves as a key biomarker, characterized by: Specific expression: The MEI1 gene is specifically highly expressed in a plasma cell subset (i.e., MEI1). + Plasmacell); Functional correlation: MEI1 expression was significantly positively correlated with TLS maturity score, HEV characteristics, and better overall survival (OS); Molecular mechanism: MEI1 + Plasma cells exhibit enhanced endoplasmic reticulum protein processing and immunoglobulin synthesis, and are associated with IL4I1. + Macrophages exhibit active APOE–LDLR metabolic communication.

6. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 1, characterized in that: The model construction step specifically includes: The BulkRNA-seq data were divided into training and validation sets at a ratio of 70% / 30%. Set the hyperparameters of the XGBoost classifier as follows: objective function is binary:logistic, maximum depth is 4, and learning rate is 0.

1. A manual early stopping mechanism is used to prevent overfitting, and the optimal prediction model is obtained by using ROC-AUC as the evaluation index.

7. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 1, characterized in that: It also includes application methods based on the prediction model, specifically including: Obtain gene expression profile data from the patients to be tested; The GSVA algorithm was used to calculate the patient's core gene score; Patients were divided into high-score and low-score groups based on the median score. If a patient belongs to the high-score group, they are predicted to respond well to T+A combination therapy and have a low TIDE immune escape score and a high level of immune cell infiltration.

8. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 7, characterized in that: The application method is specifically used to prepare companion diagnostic kits for hepatocellular carcinoma immunotherapy, or to guide clinical medication decisions for hepatocellular carcinoma patients using the T+A regimen (artezizumab combined with bevacizumab).

9. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 1, characterized in that: The predictive model identified a specific spatial niche, Niche_9, that was significantly associated with the response to atezolizumab combined with bevacizumab (T+A) treatment, and the MEI1 gene, which is specifically highly expressed in this niche. Niche_9 is a tertiary lymphoid structure-like (TLS-like) niche, composed of B cells, plasma cells, NK cells, Naïve T cells, Treg cells, and IL4I1. + Macrophages are a specific component and exhibit close co-localization of B cells and T cells in space; The MEI1 gene is specifically expressed in a plasma cell subset (MEI1) within Niche_9. + Plasma cells), whose expression levels were significantly positively correlated with TLS maturation scores, HEV characteristics, and overall patient survival, and were also associated with IL4I1. + Macrophages exhibit active APOE–LDLR metabolic communication; The core predictive gene set screened based on the cellular composition and spatial distribution of Niche_9 and the molecular characteristics of MEI1 constitutes the key biological basis for predicting the efficacy of the T+A regimen.

10. The method for constructing a predictive model for hepatocellular carcinoma immunotherapy response as described in claim 1, characterized in that: The MEI1 gene in the core predictive feature gene set serves as a non-traditional use of a companion diagnostic target for hepatocellular carcinoma immunotherapy: The application of the MEI1 gene in the prediction model reveals MEI1 + Plasma cells and IL4I1 + There is an active APOE–LDLR metabolic communication mechanism among macrophages. By combining MEI1 with core genes such as CD8B and CXCR4, a predictive model for the combination of atezolizumab and bevacizumab in hepatocellular carcinoma can be constructed. This model can be used to prepare companion diagnostic kits for immunotherapy of hepatocellular carcinoma or to guide clinical medication decisions.