Method for evaluating the prognostic value of an immune infiltrate model

CN116206681BActive Publication Date: 2026-08-11XIANGYA HOSPITAL CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,尽管已经开发了诸如xCell、CIBERSORT和TIMER等算法来量化基于批量/单细胞测序数据集的TIICs的表达水平,以促进TIICs的研究,但这些方法受到可能导致TIICs的不同参考基因组的限制

Benefits of technology

[0029]本发明全面收集了肿瘤微环境中的免疫细胞类型,并引入了细胞对算法以在GBM中开发强大的免疫特征,免疫特征可以帮助识别具有更好免疫治疗反应的GBM患者,此外,巨噬细胞/周细胞和CD163/MCAM被证实主要影响GBM患者的生存,基于已识别免疫细胞类型的相对丰度构建细胞对ICP评分,因此,高ICP评分预测GBM患者的总生存期较差,此外,ICP评分与各种致瘤和免疫原性因素密切相关,可以灵敏地预测抗PD-1免疫治疗的反应,ICP评分有望加深对GBM TME中TIICs的理解,改善GBM患者的临床管理,同时,基因对CD163/MCAM有望成为GBM的潜在预后标志物和治疗靶点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206681B_ABST
    Figure CN116206681B_ABST
Patent Text Reader

Abstract

This invention provides a method for evaluating the prognostic value of gene pairs in an immune-infiltrating cell model, belonging to the field of gene technology. The method includes the following steps: collecting immune cell gene sets from tumor immunology research; constructing an ICP score in GBM samples based on Gaussian and cell pair algorithms; determining mutational characteristics of the ICP score; defining immunogenicity characteristics of the ICP score; identifying the optimal prognostic cell pair of endothelial cells and macrophages based on the constructed ICP score; identifying the optimal prognostic gene pair of CD163 / MCAM by combining cell surface molecules; determining the role of CD163 / MCAM in cell interactions; and validating the prognostic value of the CD163 / MCAM gene pair in sequencing data and immunohistochemical samples from the Xiangya cohort. This method comprehensively collects immune cell types in the tumor microenvironment and introduces cell pair algorithms to develop robust immune signatures in GBM. These immune signatures can help identify GBM patients with better immunotherapy responses. Macrophage / peripheral cells and CD163 / MCAM have been shown to primarily influence the survival of GBM patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene technology, and in particular to a method for evaluating the prognostic value of gene pairs in an immune infiltrating cell model. Background Technology

[0002] The World Health Organization (WHO) classification defines grade I and II gliomas as low-grade gliomas (LGG) and grade III and IV gliomas as high-grade gliomas (HGG). Glioblastoma (GBM) is widely recognized as the most destructive primary brain tumor, with an extremely high mortality rate. Typically, the 10-year survival rate for LGG patients is 47%, with a median survival of 11.6 years, while the median survival for GBM patients is less than 15 months. Despite surgical resection followed by adjuvant chemoradiotherapy, the prognosis for GBM patients remains poor. To date, biomarkers including isocitrate dehydrogenase (IDH), 1p19q, and O-6-methylguanine-DNA methyltransferase (MGMT), and molecular subtypes including protoneurotic, classical, and mesenchymal, have been used for the precise classification of GBM patients to facilitate clinical management and enable personalized treatment.

[0003] Tumor-infiltrating immune cells (TIICs), including T cells, mast cells, tumor-associated macrophages (TAMs), cancer-associated fibroblasts (CAFs), and natural killer (NK) cells, can trigger a strong immune response against tumors. TIICs play a central role in regulating immune surveillance of cancer and creating a relaxed microenvironment that accelerates tumor progression. Previous studies have explored the roles of several TIICs in various cancer types, including ovarian cancer, lung adenocarcinoma, pancreatic tumors, and melanoma. Notably, TIICs have been proposed as mediators or targets for immunotherapy. Furthermore, with the rapid development of bioinformatics, which provides new insights for cancer research based on large-scale analyses, many studies have established risk profiles based on immune-infiltrating cells in various cancer types. However, the comprehensive role of TIICs in the tumor microenvironment (TME) of GBM lacks in-depth understanding. Therefore, developing TIIC-based profiling could help determine the prognostic value of TIICs in GBM and improve the efficacy of immunotherapy. However, although algorithms such as xCell, CIBERSORT, and TIMER have been developed to quantify the expression levels of TIICs based on batch / single-cell sequencing datasets to facilitate TIIC research, these methods are limited by the different reference genomes that may lead to TIICs. Different research results from different studies also contribute to this limitation. Given that the proportion of each TIIC in the tumor microenvironment lies within a relatively stable range, exploring the proportions of different TIICs could potentially optimize TIIC quantification in TME studies. Summary of the Invention

[0004] The purpose of this invention is to provide a method for evaluating the prognostic gene pair value of an immune infiltrating cell model, thereby solving the technical problems mentioned in the background art.

[0005] TIICs modulate immune surveillance and immune escape in cancer cells. The prognostic value of TIICs in various cancer types has been reported. However, the overall survival benefit of multiple TIICs in GBM has not been fully explored, and a consensus-oriented risk profile for TIICs has not yet been reached. Furthermore, previous immune cell-derived prognostic models have been limited in cross-validation across different transcriptome datasets due to the heterogeneous reference genome and immune cell signatures. Frequent updates to immune cell and reference genome annotations may hinder their widespread application and impede their prospects in clinical practice (50). To address this issue, we collected and integrated 65 immune cells to construct a robust and comprehensive risk profile. In addition, we introduced the concept of cell pairs for constructing prognostic immune profiles. We explored the possibility of using the relative expression levels of immune cells to calculate CP scores, which significantly reduces the impact of updated reference genome annotations, eliminates the need for data standardization, and improves the accuracy of the designed models.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A method for evaluating the prognostic gene pair value of an immune-infiltrating cell model, the method comprising the following steps.

[0008] Step 1: Collect the immune cell gene set in tumor immunology research, and construct the ICP score in the GBM sample based on the Gaussian algorithm and cell pair algorithm;

[0009] Step 2: Determine the mutation characteristics of the ICP score;

[0010] Step 3: Define the immunogenicity characteristics of the ICP score;

[0011] Step 4: Based on the construction of ICP scores, the optimal prognostic cell pairs of endothelial cells and macrophages are identified, and further, by combining cell surface molecules, the optimal prognostic gene pair of CD163 / MCAM is identified.

[0012] Step 5: Determine the role of CD163 / MCAM in cell interactions at the single-cell level;

[0013] Step 6: Validate the prognostic value of the CD163 / MCAM gene pair in sequencing data and immunohistochemical samples from the Xiangya cohort.

[0014] Furthermore, the specific process of step 1 is as follows:

[0015] Step 1.1: Collection and preprocessing of immune cell gene sets. A total of 1127 GBM patient samples were collected from 6 cohorts and defined as the consolidation cohort. 523 GBM patient samples were from TCGA, and single-cell RNA sequencing data of 33 GBM patient samples were from the Single Cell Portal platform. Raw data from the microarray dataset generated by Agilent was downloaded from GEO. Gene expression profiles and corresponding clinical information generated by Illumina were downloaded from TCGA and CGGA. The raw data from the Agilent dataset were processed using the RMA algorithm in the limma package for background adjustment. The raw data from Illumina were processed using the lumi package. The fragment per kilobase per million value of RNA-seq data was converted to transcript per kilobase per million value. The batch effect was removed using the R package sva.

[0016] Step 1.2: Immune cell gene set. Immune cell features were integrated from publicly available resources. By integrating gene sets of immune cell types from different literature, 65 immune cell features were finally obtained, and a list of 65 immune cell types was provided in advance.

[0017] Step 1.3: Develop a reliable risk model in GBM and perform univariate Cox analysis to screen the GBM dataset TCGAGBM-RNAseq. The TCGAGBM-RNAseq dataset contains 523 samples with prognostic value for prognostic-related immune cell types. Then, the prognostic-related immune cell type Ci is paired with all 65 immune infiltrating cell types Cj. For cell pairs starting with immune cell type Ci and immune infiltrating cell type Cj, Score_ij = 1 (exp_Ci – exp_Cj > 0) and Score_ij = 0 (exp_Ci – exp_Cj < 0) are used. The 2-year area under the curve (AUC) was used to estimate the performance of each Score_ij, and cell pairs with statistically significant prognosis and the highest 2-year AUC were identified. For each immune cell type Ci, Score_ij was determined as the highest 2-year AUC. The identified cell pairs with the highest 2-year AUC were further ranked, with a hazard ratio (HR) > 1, and duplicate cell pairs were removed. Then, hierarchical agglomerative clustering based on a cell pair model using a Gaussian finite mixture model (GMM) was used for classification. Finally, the ICP score was calculated using the selected Score_ij, where ICP score = Score_ij.

[0018] Furthermore, the specific process of step 2 is as follows:

[0019] Step 2.1: Genomic alterations of ICP scores. Download somatic mutations and somatic copy number variations (CNVs) corresponding to GBM samples with RNA-seq data from TCGA. Visualize somatic mutations using the R package maftools. Use GISTIC 2.0 analysis to determine CNVs associated with the two ICP score groups and the threshold copy number of alteration peaks.

[0020] Step 2.2: Immune infiltration analysis of ICP score. Gene features of 115 metabolic signaling pathways and seven types of immune checkpoint molecules were obtained from existing technologies. Several immunomodulators were collected. The xCell algorithm, TIMER algorithm, EPIC algorithm, MCPcounter algorithm, quanTlseq algorithm and CIBERSORT algorithm were used to identify immune infiltrating cells in the GBM tumor microenvironment.

[0021] Further, the specific process of step 3 is as follows: prediction of ICP score in immunotherapy response. GBM samples receiving anti-PD1 immunotherapy in the PRJNA482620 dataset are collected to evaluate the predicted value of ICP score. The urothelial carcinoma cohort and melanoma dataset GSE78220 are further used to predict immunotherapy response. The original data from the two datasets are standardized using the DEseq2R package, and the expression values ​​of the original matrix are converted into TPM values. ICP scores are calculated in both cohorts respectively.

[0022] Further, step 4 involves identifying the cell pair macrophages / peripheral cells and the gene pair CD163 / MCAM. Based on 2y-AUC, the cell pairs most associated with prognosis are explored. Functional annotations are performed on the identified cell pairs macrophages / peripheral cells, including biological processes, metabolic pathways, inflammatory features, and immune infiltration. CD31, NG2, PDGFR beta, CD146, and Nestin are used as peripheral cell markers, while CD11b, CD68, CD163, CD14, and CD16 are used as macrophage markers. Then, markers from macrophages and markers from peripheral cells are paired. Furthermore, gene pairs most associated with prognosis are explored based on 2y-AUC. Functional annotations are performed on the identified gene pair CD163 / MCAM, including biological processes, metabolic pathways, inflammatory features, and immune infiltration.

[0023] Further, the specific process of step 5 is as follows: annotate genes and perform single-cell sequencing on CD163 / MCAM. Based on the R package infercnv, tumor cells are first identified. After performing principal component analysis (PCA) using the R package RunPCA, the R package FindNeighbors is used to define K-nearest neighbors. Based on the level of gene alteration, the R package FindClusters is used to combine cells with the highest gene alteration. The R packages UMAP and tSNE are used for dimensionality reduction. The R package scCATCH is used for annotation of non-malignant cell types. The R package FindMarkers is used to screen for genes with significant differential expression in cell types. The Scalop algorithm is used to define four types of GBM at the single-cell level. The R package CellChat is used to explore cell communication patterns and analyze and visualize different receptor-ligand signaling pathways.

[0024] Furthermore, the specific process of step 6 is as follows:

[0025] Step 6.1: Formalin-fixed paraffin-embedded tumor tissues from 73 GBM patients were collected for sequencing. 1 μg of RNA was used as input material for RNA sample preparation. DNA was cut, and sequencing libraries were prepared using the NEBNext Ultra RNA Library Prep Kit. PCR was then performed using Phusion high-fidelity DNA polymerase, universal PCR primers, and index X primers. After capturing the target region with biotin-labeled probes, the captured library was sequenced on the Illumina Hiseq platform to generate 125 / 150 bp paired-end reads. Internal perlscripts were used to process the raw data, including reads containing adapters and poly-N. Low-quality reads were removed to obtain clean reads. Reference genomes and gene model annotation files were obtained from a genome website. The reference genome index was constructed using Hisat2 v2.0.5. The paired-end clean reads were aligned with the reference genome, and then FeatureCounts were used. v1.5.0-p3 calculates the number of reads mapped to each gene. The TPM for each gene is calculated based on the gene length, and the read count is mapped to the corresponding gene.

[0026] Step 6.2: The tissue was surgically removed from the patient during GBM surgery at the hospital. The tissue was then fixed in formalin and embedded in paraffin for subsequent sectioning. Sections were 4 μm in size and then boiled for antigen retrieval. 3% H2O2 was used as an inhibitor of endogenous HPR activity, and 5% BSA was used for section blockade. Rabbit polyclonal anti-CD163 and anti-MCAM antibodies were used, while endogenous HRP-labeled goat anti-rabbit IgG was used as the secondary antibody. Sections with the primary antibody were incubated overnight at 4°C. The substrate was mixed with Solution 1 and Solution 2 at a 1 drop / 1000 ml / min ratio. The ratio of ml was used to examine the signal. The substrate was 3,3'-diaminobenzidine. DAB and hematoxylin were used for staining the sections. After staining, the sections were observed with an optical microscope. For intensity scoring, four intensity levels, negative, weak, medium and strong, were assigned as level 0, level 1, level 2 and level 3, respectively. As for degree scoring, it was the proportion of stained cells. 10%, 10-25%, 25-50%, 50-75% and >75% were assigned as 0, 1, 2, 3 and 4, respectively. The H score was calculated as range * intensity, with a range of 0-12.

[0027] Step 6.3: The log-rank test was used to determine survival differences, and survival curves were generated using the R package survminer. The clinical significance of prognostic factors was determined by univariate and multivariate Cox regression analysis. Correlation coefficients were calculated by Pearson correlation analysis. The R package pROC was used to visualize receiver operating characteristics ROC analysis. The R package maftools was used to depict the mutational landscape of TCGA via OncoPrint. All statistical analyses were performed on R item 3.6.3, and P < 0.05 was considered statistically significant.

[0028] The present invention, by adopting the above-described technical solution, has the following beneficial effects:

[0029] This invention comprehensively collects immune cell types in the tumor microenvironment and introduces a cell pair algorithm to develop robust immune signatures in GBM. These signatures can help identify GBM patients with better immunotherapy responses. Furthermore, macrophage / pericyllocyte and CD163 / MCAM have been shown to significantly influence GBM patient survival. Based on the relative abundance of identified immune cell types, a cell pair ICP score is constructed. Therefore, a high ICP score predicts poorer overall survival in GBM patients. In addition, the ICP score is closely related to various tumorigenic and immunogenic factors and can sensitively predict responses to anti-PD-1 immunotherapy. The ICP score is expected to deepen the understanding of TIICs in GBM TME and improve the clinical management of GBM patients. Meanwhile, the gene pair CD163 / MCAM is expected to become a potential prognostic biomarker and therapeutic target for GBM. Attached Figure Description

[0030] Figure 1 This is a flowchart of the method of the present invention;

[0031] Figure 2 This is a flowchart of the cell pairing algorithm of the present invention and a summary diagram of related sample data;

[0032] Figure 3 This is a graph showing the immunogenicity and tumorigenicity characteristics of ICP scores in TCGA and their correlation with ICP scores in this invention.

[0033] Figure 4 This is a graph showing the immune infiltration characteristics of ICP score and the infiltration characteristics related to ICP score in the TCGA of this invention;

[0034] Figure 5 This is a graph showing the predictive value and related data of the ICP score in immunotherapy according to the present invention;

[0035] Figure 6 This is a graph showing the prognostic value and related data of the gene of this invention for CD163 / MCAM;

[0036] Figure 7 This is a graph showing the molecular characteristics and related data of the CD163 / MCAM gene pair at the single-cell sequencing level of this invention. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and preferred embodiments. However, it should be noted that many details listed in the specification are merely to provide the reader with a thorough understanding of one or more aspects of the present invention, and these aspects of the invention can be implemented even without these specific details.

[0038] like Figure 1 As shown, a method for evaluating the prognostic cell pair value of an immune-infiltrating cell model includes the following steps:

[0039] Step 1: Collect the immune cell gene set in tumor immunology research, and construct the ICP score in the GBM sample based on the Gaussian algorithm and cell pair algorithm. The ICP score is referred to as CP score below.

[0040] Step 1.1: Dataset Collection and Preprocessing:

[0041] The publicly available GBM cohorts were obtained from the Gene Expression Omnibus (GEO; https: / / www.ncbi.nlm.nih.gov / geo / ), the Cancer Genome Atlas (TCGA) (https: / / xenabrowser.net / ), and the Chinese Glioma Genome Atlas (CGGA; http: / / www.cgga.org.cn / ). A total of 523 GBM patient samples were obtained from TCGA. A total of 1127 GBM patient samples were collected from 6 cohorts and defined as the consolidation cohort. Single-cell RNA sequencing data from 33 GBM patient samples (accessions SCP50 and SCP393) were obtained from the Single Cell Portal platform (http: / / singlecell.broadinstitute.org). Information on the platform and patient samples is provided in Table S1. Raw data from the microarray dataset generated by Agilent were downloaded from GEO. Gene expression profiles and corresponding clinical information generated by Illumina were downloaded from TCGA and CGGA. Raw data from the Agilent dataset underwent background adjustment using the RMA algorithm in the limma software package. Raw data from Illumina were processed using the lumi package (18). FPKM values ​​of RNA-seq data were converted to TPM values, which were more comparable to RMA-normalized values ​​from microarray datasets (19). Computational batch effects were removed using the R package sva.

[0042] Step 1.2: Immune cell gene set:

[0043] Immune cell signatures were integrated from publicly available resources. By integrating gene sets of immune cell types from different literature, 65 immune cell signatures were ultimately obtained and deemed reliable. Our previous findings provided a list of 65 immune cell types.

[0044] Step 1.3: Develop a reliable risk model in GBM:

[0045] Univariate Cox analysis was performed to screen for prognostic-related immune cell types with prognostic value in the GBM dataset TCGAGBM-RNAseq (523 samples). The prognostic-related immune cell types (Ci) were then paired with all 65 immune-infiltrating cell types (Cj). For cell pairs starting with Ci, Ci, and Cj, Score_ij = 1 (exp_Ci – exp_Cj > 0) and Score_ij = 0 (exp_Ci – exp_Cj < 0). The area under the curve (AUC) over 2 years was used to estimate the performance of each Score_ij, and cell pairs with statistically significant prognosis and the highest 2-year AUC were identified (23). For each Ci, Score_ij was determined as the highest 2-year AUC. The identified cell pairs with the highest 2-year AUC were further ranked, with a hazard ratio (HR) > 1, and duplicate cell pairs were removed. Subsequently, hierarchical agglomerative clustering based on a cell pair model using a Gaussian finite mixture model (GMM) was used for classification. Then, use these selected Score_ij to calculate the CP score:

[0046] CP score = Score_ij

[0047] Step 2: Determine the mutation characteristics of the ICP score.

[0048] Step 2.1: Genomic alterations in CP scores:

[0049] Download somatic mutations and somatic copy number variations (CNVs) corresponding to GBM samples with RNA-seq data from TCGA. Visualize somatic mutations using the R package maftools. Use GISTIC 2.0 analysis (https: / / gatk.broadinstitute.org) to determine CNVs associated with the two CP score groups and threshold copy numbers of altered peaks.

[0050] Step 2.2: Immune infiltration analysis based on CP score:

[0051] Gene signatures for 115 metabolic signaling pathways were derived from previously published work. Seven types of immune checkpoint molecules were identified from a previous study. A variety of immunomodulators were collected. Immune infiltrating cells in the GBM tumor microenvironment were identified using the xCell, TIMER, EPIC, MCPcounter, quanTlseq, and CIBERSORT algorithms.

[0052] Step 3: Analyze the ICP score for immunotherapy purposes.

[0053] To predict CP scores in immunotherapy response, GBM samples receiving anti-PD1 immunotherapy in the PRJNA482620 dataset were collected to evaluate the predictive values ​​of CP scores. The IMvigor210 cohort (urothelial carcinoma cohort) and GSE78220 (melanoma dataset) were further used to predict immunotherapy response (37,38). The raw data from the two datasets were standardized using the DEseq2R package, and the expression values ​​of the raw matrices were converted into TPM values. CP scores were calculated in both cohorts.

[0054] Step 4: Based on the construction of ICP scores, the optimal prognostic cell pairs of endothelial cells and macrophages are identified, and further, by combining cell surface molecules, the optimal prognostic gene pair of CD163 / MCAM is identified.

[0055] Cell pairings of macrophages / peripheral cells and gene pairings of CD163 / MCAM were identified based on 2y-AUC to explore the cell pairs most associated with prognosis. Functional annotations of the identified macrophage / peripheral cell pairs were performed, including biological processes, metabolic pathways, inflammatory features, and immune infiltration.

[0056] CD31, NG2, PDGFR beta, CD146, and Nestin were used as pericyte markers, while CD11b, CD68, CD163, CD14, and CD16 were used as macrophage markers. Markers from macrophages were then paired with those from pericytes. Gene pairs most associated with prognosis were also explored based on 2y-AUC. Functional annotations of the identified gene pair CD163 / MCAM were performed, including biological processes, metabolic pathways, inflammatory features, and immune infiltration.

[0057] Step 5: Determine the role of CD163 / MCAM in cell interactions at the single-cell level.

[0058] Single-cell sequencing for gene pairing CD163 / MCAM, based on the R package "infercnv", initially identifies tumor cells. Principal component analysis (PCA) is performed using the R package "RunPCA", and K-nearest neighbors are defined using the R package "FindNeighbors". Cells with the highest gene alterations are grouped using the R package "FindClusters" based on their level of gene alteration. The R packages "UMAP" and "tSNE" are used for dimensionality reduction. The R package "scCATCH" is used for annotation of non-malignant cell types. The R package "FindMarkers" is used to screen for genes with significant differential expression in cell types. Four types of GBM are defined at the single-cell level using the "Scalop" algorithm. Cell communication patterns are explored using the R package "CellChat", analyzing and visualizing different receptor-ligand signaling pathways.

[0059] Step 6: Validate the prognostic value of the CD163 / MCAM gene pair in sequencing data and immunohistochemical samples from the Xiangya cohort.

[0060] Step 6.1: Transcriptome sequencing of the Xiangya cohort. Formalin-fixed paraffin-embedded tumor tissue from 73 GBM patients was collected for sequencing. Briefly, 1 μg of RNA was used as input material for RNA sample preparation. DNA was cut, and sequencing libraries were prepared using the NEBNext Ultra RNA Library Prep Kit. PCR was then performed using Phusion high-fidelity DNA polymerase, universal PCR primers, and index (X) primers. After capturing the target region with biotin-labeled probes, the captured library was sequenced on the Illumina Hiseq platform to generate 125 / 150 bp paired-end reads. Internal perlscripts were used to process the raw data (raw reads). Reads containing adapters and poly-N were then used to remove low-quality reads, resulting in clean reads. Reference genome and gene model annotation files were obtained from the genome website (http: / / genome.ucsc.edu). The reference genome index was constructed using Hisat2 v2.0.5, and the paired-end clean reads were aligned with the reference genome. Then, FeatureCounts v1.5.0-p3 is used to calculate the number of reads mapped to each gene. The TPM for each gene is calculated based on the gene length, and the read count is mapped to that gene.

[0061] Step 6.2: Tissue was sourced from patients (n=45) who underwent GBM surgery at Xiangya Hospital, Central South University. The tissue was then fixed in formalin and embedded in paraffin for subsequent sectioning (4 μm). Sections were then boiled for antigen retrieval using 3% H2O2 as an inhibitor of endogenous HPR activity. 5% BSA was used for section blocking. Rabbit polyclonal anti-CD163 and anti-MCAM antibodies (1:50; Proteintech; Wuhan, China) were primary antibodies, while HRP-labeled goat anti-rabbit IgG was the secondary antibody. Sections with primary antibodies were incubated overnight at 4°C. A substrate (3,3'-diaminobenzidine, DAB) was mixed with solutions 1 and 2 at a ratio of 1 drop / 1 ml for signal examination. Hematoxylin was used for section staining. The sections were finally observed using an optical microscope after staining. For intensity scoring, four intensity levels—negative, weak, medium, and strong—were assigned as 0, 1, 2, and 3, respectively. As for the severity rating (the percentage of stained cells), 10%, 10-25%, 25-50%, 50-75%, and >75% were assigned 0, 1, 2, 3, and 4, respectively. The H score was calculated as range * intensity, ranging from 0 to 12.

[0062] Step 6.3: The log-rank test was used to determine survival differences, and survival curves were generated using the R package survminer. The clinical significance of prognostic factors was determined by univariate and multivariate Cox regression analysis. Correlation coefficients were calculated by Pearson correlation analysis. Receiver operating characteristics (ROC) analysis was visualized using the R package pROC. The R package maftools was used to depict the mutational landscape of TCGA using OncoPrint(40). All statistical analyses were performed on R item 3.6.3. P < 0.05 was considered statistically significant.

[0063] Implementation process:

[0064] The CP score was constructed and its prognostic value was determined by identifying 26 immune cell types with prognostic value in the TCGA GBM sample using univariate Cox regression analysis, and pairing them with 65 integrated immune cell types collected from previously published studies. Each cell pair was assigned a score of 1 or 0 based on its relative expression level. After calculating the 2-year AUC for all cell pairs, the 13 immune cell pairs with the highest 2-year AUC were identified. After performing GMM, the CP score based on 6 immune cell pairs was ultimately selected based on the highest AUC. Figure 2 A). The HR of the 13 cell pairs with the highest 2-year AUC values ​​is shown. The AUC of the CP scoring model, ranked by the GMM classifier across all 8191 formulas, is as follows: Figure 2 As shown in C. CP score predicted poor survival rates for LGG, GBM, and panglioma samples from TCGA (log-rank test, p < 0.001; respectively). Figure 2 D, 2E, and 2F). Furthermore, CP score was a risk factor in both the GBM and panglioma samples from the Xiangya cohort (log-rank test, p < 0.001; respectively). Figure 2 G and 2H). The 2-, 3-, 4-, and 5-year AUCs of ROC analysis were 0.703, 0.738, 0.767, and 0.797, respectively, confirming that the CP score can serve as a prognostic biomarker for predicting the survival of GBM patients with TCGA. Figure 2 I).

[0065] Immune escape mechanisms associated with CP scores were investigated, and high CP scores were found to be significantly associated with N-glycan biosynthesis, kynurenine metabolism, and prostaglandin biosynthesis in TCGA and the integration cohort (for...). Figure 2A). The cancer immune cycle has been proposed as a comprehensive reflection of the functions of several chemokines and immunomodulators (42,43). Notably, most steps in the cancer immune cycle are upregulated in high CP score groups, including the release of cellular antigens, tumor antigen presentation, and recruitment of immune cells (CD8 T cells, dendritic cells, macrophages, MDSCs, monocytes, neutrophils, NK cells, Th1 cells, Th17 cells, and Th22 cells), as well as the infiltration of immune cells into the tumor in TCGA and integration cohorts (for...). Figure 3 B).

[0066] First, a range of tumorigenic and immunogenic factors were assessed. The high CP score group exhibited a higher T-cell inflammatory gene expression profile (GEP), indicating a higher response rate to anti-PD-1 therapy. Figure 3 C). The high CP score group also showed lower homologous recombination deficiency (HRD), an indicator of cell death. Figure 3 D). Interestingly, high CP score arrays are associated with a higher number of segments ( Figure 3 E). Matrix characteristics, including TGF-β response, white blood cell count, matrix count, interferon-γ (IFNG), IFNG marker gene set (IFNG.GS), and interferon-stimulated gene resistance trait (ISG.RS), were all higher in the high CP score group. Figure 3 F-3K). Regarding antigen presentation capacity, the high CP score group exhibited higher levels of T cell receptor (TCR) Shannon index, TCR abundance, and higher antigen processing and presentation mechanism (APM) scores (respectively...). Figure 3 L-3N).

[0067] Immune infiltration characteristics were also assessed in the CP score groups. Therefore, a high CP score was associated with higher levels of ESTIMATE score, immune score, and matrix score. Figure 4 A). Based on six different algorithms, high CP scores were significantly associated with immunosuppressive cells, including regulatory T cells (Tregs), TAMs, CAFs, T helper 2 cells (Th2), and dendritic cells (DCs). Figure 4 B). GO results from GSVA confirmed that oncogenic pathways include regulation of the ERBB signaling pathway, Toll-like receptor signaling pathway, NF-κB transcription factor activity, and glial cell activation, as well as immunogenic pathways including regulation of macrophages, chemokine production, and mast cell activation. Activation was observed in high CP score groups in TCGA and the integration cohort. Figure 4 C). Furthermore, high CP scores were significantly associated with PD-1 treatment efficacy, T-cell signaling, hypoxia signaling, exosome signaling, and immunosuppressive cell signaling, including Tregs, myeloid-derived suppressor cells (MDSCs), TAMs, and CAFs in the TCGA and integration cohorts ( Figure 4 D).

[0068] Next, we explored the association between CP scores and seven types of immune checkpoint molecules. High CP scores were positively correlated with most immune checkpoint molecules, including ICOS, PDCD1, CTLA4, and CD40, and it is possible that these classic immune checkpoint molecules from TCGA can be used to evade immune responses. Figure 5 A). Notably, the differential expression of immune checkpoint molecules in the CP score groups was not present in somatic mutations and CNVs, but was closely associated with methylation. Figure 5 A).

[0069] CP scores predict immunotherapy response, and immunotherapy has revolutionized cancer treatment. Therefore, the predictive value of CP scores in immunotherapy response was also explored. In a cohort exploring the response to anti-PD-1 immunotherapy in GBM patients, patients with higher CP scores had lower responses to PD-1 immunotherapy. Figure 5 B). In the melanoma dataset GSE78220, high CP scores predict poor survival outcomes. Figure 5 C). Similarly, patients with high CP scores exhibit both stable and progressive disease (C). Figure 5 D). The CP score was also constructed in the IMvigor210 cohort (urothelial carcinoma dataset). As expected, a high CP score predicted poorer survival outcomes. Figure 5 E). Patients with high CP scores exhibited both stable disease and disease progression (E). Figure 5 F).

[0070] Based on six different algorithms, cell group M>P was significantly associated with immunosuppressive cells, including Treg, TAM, CAF, and DC ( Figure 6 A). Furthermore, M>P in cell groups predicts poor survival rates in TCGA (A). Figure 6 C).

[0071] Functional annotation of CD163 / MCAM genes, based on six different algorithms, showed that high levels were significantly associated with immunosuppressive cells, including Treg, TAM, CAF, Th2, and DC. Figure 6 B). Furthermore, high-group predictions showed poorer survival rates for TCGA (B). Figure 6 D).

[0072] Validation of the CD163 / MCAM gene pair in the Xiangya cohort: In sequencing data from 73 GBM samples in the Xiangya cohort, a high level of CD163 / MCAM was associated with decreased survival. Figure 6E). In addition, IHC staining was performed on 45 GBM samples from the Xiangya cohort. Based on the H-scores of CD163 and MCAM in the IHC staining results, the 45 GBM samples were divided into a high group (CD163>MCAM) and a low group (CD163>MCAM). <MCAM)( Figure 6 G). Notably, according to IHC staining data from the Xiangya cohort, the high group was also associated with decreased survival (G). Figure 6 F).

[0073] Characterization of CD163 / MCAM at the Single-Cell Level: To further elucidate the role of the gene in CD163 / MCAM within the GBM tumor microenvironment, we performed single-cell sequencing analysis on 33 GBM samples. Tumor cells were defined as aneuploid cells, and after t-SNE dimensionality reduction, a total of 11 cell types were identified (…). Figure 7 A). Figure 6 B shows differentially expressed genes (DEGs) among 11 cell types. CD163 was found to be more enriched in dendritic cells (DCs), microglia, and M0 / M1 / M2 macrophages, while MCAM was more enriched in OPCs, oligodendrocytes, and vascular cells. Figure 7 C). Based on the relative expression of CD163 and MCAM, after UMAP dimensionality reduction, cells were divided into a high-expression group (CD163>MCAM) and a low-expression group (CD163>MCAM). <MCAM)( Figure 7 D). The relative comparisons of the two groups in 11 cell types are as follows: Figure 7 As shown in E, the high-dose group was more dominated by M0 macrophages and microglia, while the low-dose group was more dominated by neurons, tumors, and OPCs. Figure 7 E). Malignant cells in GBM were classified into four main types at the single-cell level: neural progenitor-like (NPC-like), oligodendrocyte progenitor-like (OPC-like), astrocyte-like (AC-like), and mesenchymal-based GBM cell expression profile similarity (MES-like). AC-like and MES-like malignant cells were found to be more associated with the high-level group, while NPC-like malignant cells were enriched in the low-level group. Figure 6 F). GSEA results confirmed that the tumorigenic pathway was more activated in the low-dose group, while the immunogenic pathway was more activated in the high-dose group. Figure 7 G).

[0074] Figure 2A. Flowchart of the cell pair algorithm. B. Forest plot depicting the 13 cell pairs with the highest 2y-AUC values. C. Patterns of the logistic regression model correlated with AUC values ​​and identified by Gaussian mixture. Nine clusters with 8191 combinations. Kaplan-Meier curves for two CP score groups in the D.LGG, E.GBM, and F. glioma samples from the TCGA dataset. Log-rank test, P<0.001. Kaplan-Meier curves for two CP score groups in the G. Xiangya cohort GBM sample. Log-rank test, P<0.001. Kaplan-Meier curves for two CP score groups in the H. Xiangya cohort glioma sample. Log-rank test, P<0.001. I. ROC curves measuring the sensitivity of CP scores in predicting 2-, 3-, 4-, and 5-year survival rates of GBM patients in the TCGA dataset. The areas under the ROC curves were 0.703, 0.738, 0.767, and 0.797, respectively.

[0075] Figure 3 Immunogenicity and tumorigenicity characteristics of CP scores in TCGA. A. Heatmap illustrating the expression patterns of metabolic features in CP scores. B. Differences in various steps of the cancer immune cycle between high and low CP score groups. C. GEP scores in high and low CP score groups. D. HRD in high and low CP score groups. E. Segments in high and low CP score groups. F. TGF-β response in high and low CP score groups. G. White blood cell count in high and low CP score groups. H. Matrix count in high and low CP score groups. I. IFNG score in high and low CP score groups. J. IFNG.GS score in high and low CP score groups. K. ISG.RS score in high and low CP score groups. L. TCR Shannon index in high and low CP score groups. M. TCR abundance in high and low CP score groups. N. APM score in high and low CP score groups.

[0076] Figure 4 In TCGA, the immune infiltration characteristics of CP scores are analyzed. A. Differences in the expression of ESTIMATE, immune, and matrix scores between high and low CP score groups. B. Estimation of the correlation between immune cells and CP scores using different algorithms. C. A heatmap illustrating the expression patterns of immune function in CP scores. D. A heatmap illustrating the expression patterns of immune regulatory features in CP scores.

[0077] Figure 5The predictive value of CP scores in immunotherapy. A. Heatmap illustrating the expression patterns of seven immunomodulators in CP scores. From left to right: CP score; mutation frequency; amplification frequency; deletion frequency and methylation (correlation between gene expression and DNA methylation value) of immunomodulators in the two CP score groups. B. Difference in CP score expression between patients with and without PD-1 response. P = 0.012. C. Kaplan-Meier curves of two CP score groups in the GSE78220 dataset. Log-rank test, P < 0.1412. D. CP scores of groups with different PD-1 clinical response statuses (CR / PR and SD / PD). Differences between groups were compared using the Wilcoxon test (Wilcoxon, P = 0.036). E. Kaplan-Meier curves of CP score groups in the IMvigor210 dataset. Log-rank test, P = 0.00174. F. CP scores of groups with different PD-1 clinical response statuses (CR, PR, SD, PD). Differences between groups were compared using the Kruskal-Wallis test (Kruskal-Wallis, P = 0.013).

[0078] Figure 6 The prognostic value of genes for CD163 / MCAM. A. Estimation of the correlation between immune cells and cell pairs (macrophages / peripheral cells) under different algorithms. B. Estimation of the correlation between immune cells and genes for CD163 / MCAM under different algorithms. C. Kaplan-Meier curves of two cell pairs in TCGA. Log-rank test, P < 0.001. D. Kaplan-Meier curves of two genomes in TCGA. Log-rank test, P = 0.00115. E. Kaplan-Meier curves of two genomes based on sequencing data from the Xiangya cohort. Log-rank test, P = 0.04908. F. Kaplan-Meier curves of two genomes based on IHC staining from the Xiangya cohort. Log-rank test, P = 0.02014. G. IHC staining of CD163 and MCAM in four representative samples from the Xiangya cohort.

[0079] Figure 7 The molecular characterization of CD163 / MCAM genes at the single-cell sequencing level. A. t-SNE plot for visualization of aneuploid cells, diploid cells, and 11 identified cell types. B. Dot plot showing differentially expressed genes among the 11 cell types. C. Differences in CD163 and MCAM expression in the 11 cell types. D. UMAP plots visualizing high (CD163 expression > MCAM expression) and low (MCAM expression > CD163 expression) cell groups. E. Bar chart showing the proportional differences of the 11 cell types in the high and low groups. F. Relative proportions of the four cell types in the high and low groups.

[0080] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for evaluating the prognostic gene pair value of an immune-infiltrating cell model, characterized in that: The method includes the following steps. Step 1: Collect the immune cell gene set in tumor immunology research, and construct the ICP score in the GBM sample based on the Gaussian algorithm and cell pair algorithm; Step 2: Determine the mutation characteristics of the ICP score; Step 3: Predict immunotherapy response based on ICP scores; Step 4: Based on the construction of ICP scores, the optimal prognostic cell pair of macrophages / peripheral cells was identified, and further, by combining cell surface molecules, the optimal prognostic gene pair of CD163 / MCAM was identified. Step 5: Determine the role of CD163 / MCAM in cell interactions at the single-cell level; Step 6: Validate the prognostic value of the CD163 / MCAM gene pair in sequencing data and immunohistochemical samples from the Xiangya cohort; The specific process of step 1 is as follows: Step 1.1: Collection and preprocessing of immune cell gene sets. A total of 1127 GBM patient samples were collected from 6 cohorts and defined as the modeling cohort. 523 GBM patient samples were from TCGA, and single-cell RNA sequencing data of 33 GBM patient samples were from the Single Cell Portal platform. Raw data from the microarray dataset generated by Agilent was downloaded from GEO. Gene expression profiles and corresponding clinical information generated by Illumina were downloaded from TCGA and CGGA. The raw data from the Agilent dataset underwent background adjustment using the RMA algorithm in the limma software package. The raw data from Illumina... The lumi package was used to process the RNA-seq data into kilobase-million-fragment values, which were then converted to transcript-million-fragment values. The R package sva was used to remove batch processing effects. Step 1.2: Integrate immune cell features from public resources. By integrating gene sets of immune cell types from different literature, 65 immune cell features were finally obtained, and a list of 65 immune cell types was provided in advance. Step 1.3: Develop a reliable risk model in GBM and perform univariate Cox analysis to screen the GBM dataset TCGAGBM-RNAseq. The TCGAGBM-RNAseq dataset contains 523 samples with prognostic value for prognostic-related immune cell types. Then, the prognostic-related immune cell type Ci is paired with all 65 immune infiltrating cell types Cj. For cell pairs starting with immune cell type Ci and immune infiltrating cell type Cj, Score_ij = 1 (exp_Ci – exp_Cj > 0) and Score_ij = 0 (exp_Ci – exp_Cj < 0), the 2-year area under the curve (AUC) is used to estimate the performance of each Score_ij, and cell pairs with statistically significant prognosis and the highest 2-year AUC are identified. For each immune cell type Ci, Score_ij is determined as the highest 2-year AUC. The identified cell pairs with the highest 2-year AUC are further ranked, with a hazard ratio (HR) > 1. Duplicate cell pairs are removed, and then hierarchical agglomerative clustering based on a cell pair model based on a Gaussian finite mixture model (GMM) is used for classification. The ICP score is then calculated using the selected Score_ij, where ICP score = Score_ij.

2. The method for evaluating the prognostic gene pair value of an immune-infiltrating cell model according to claim 1, characterized in that: The specific process of step 2 is as follows: Step 2.1: Genomic alterations of ICP scores. Download somatic mutations and somatic copy number variations (CNVs) corresponding to GBM samples with RNA-seq data from TCGA. Visualize somatic mutations using the R package maftools. Use GISTIC2.0 analysis to determine CNVs associated with the two ICP score groups and the threshold copy number of alteration peaks. Step 2.2: Immune infiltration analysis of ICP score. Gene features of 115 metabolic signaling pathways and seven types of immune checkpoint molecules were obtained from existing technologies. Several immunomodulators were collected. The xCell algorithm, TIMER algorithm, EPIC algorithm, MCPcounter algorithm, quanTlseq algorithm or CIBERSORT algorithm were used to identify immune infiltrating cells in the GBM tumor microenvironment.

3. The method for evaluating the prognostic gene pair value of an immune-infiltrating cell model according to claim 1, characterized in that: Step 3 involves the prediction of ICP scores in immunotherapy response. GBM samples receiving anti-PD1 immunotherapy in the PRJNA482620 dataset were collected to evaluate the predicted values ​​of ICP scores. The urothelial carcinoma cohort and the melanoma dataset GSE78220 were further used to predict immunotherapy response. The raw data from the two datasets were standardized using the DEseq2 R package, and the expression values ​​of the original matrix were converted into TPM values. ICP scores were calculated in both cohorts.

4. The method for evaluating the prognostic gene pair value of an immune-infiltrating cell model according to claim 1, characterized in that: Step 4 involves identifying the cell pair of macrophages / peripheral cells and the gene pair CD163 / MCAM. Based on 2y-AUC, the cell pairs most associated with prognosis are explored. Functional annotations are performed on the identified cell pairs of macrophages / peripheral cells, including biological processes, metabolic pathways, inflammatory features, and immune infiltration. CD31, NG2, PDGFR beta, CD146, and Nestin are used as peripheral cell markers, while CD11b, CD68, CD163, CD14, and CD16 are used as macrophage markers. Then, markers from macrophages and markers from peripheral cells are paired. Furthermore, gene pairs most associated with prognosis are explored based on 2y-AUC. Functional annotations are performed on the identified gene pair CD163 / MCAM, including biological processes, metabolic pathways, inflammatory features, and immune infiltration.

5. The method for evaluating the prognostic gene pair value of an immune-infiltrating cell model according to claim 1, characterized in that: Step 5 involves the following steps: Gene annotation is performed on CD163 / MCAM single-cell sequencing. Based on the R package infercnv, tumor cells are first identified. After performing principal component analysis (PCA) using the R package RunPCA, the R package FindNeighbors is used to define K-nearest neighbors. Based on the level of gene alteration, the R package FindClusters is used to group cells with the highest gene alterations. The R packages UMAP and tSNE are used for dimensionality reduction. The R package scCATCH is used for annotation of non-malignant cell types. The R package FindMarkers is used to screen for genes with significant differential expression in cell types. The Scalop algorithm is used to define four types of GBM at the single-cell level. The R package CellChat is used to explore cell communication patterns and analyze and visualize different receptor-ligand signaling pathways.