Head and neck squamous cell carcinoma prognosis prediction model based on characteristic macrophage subpopulation and construction method thereof

The prognostic prediction model based on TREM2+TAMs cell population, constructed by integrating multi-omics data, solves the problem of insufficient accuracy in HNSCC prognostic prediction models, achieves accurate prediction of HNSCC after immunotherapy, and improves the accuracy of prognostic assessment and guidance for individualized treatment.

CN121812175APending Publication Date: 2026-04-07HOSPITAL OF STOMATOLOGY SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing prognostic prediction models for head and neck squamous cell carcinoma (HNSCC) have limited accuracy in assessing lymph node metastasis. Traditional imaging examinations have a low detection rate for micrometastases, and traditional treatments rely on anatomical parameters that cannot reflect the intrinsic degree of malignancy, resulting in inaccurate prognostic stratification and high recurrence and mortality rates.

Method used

A prognostic prediction model based on characteristic macrophage subsets was constructed. By integrating single-cell and microarray data through multi-omics analysis, TREM2+TAMs cell populations were identified. Combined with multiplex immunofluorescence and in vitro co-culture techniques, the role of TREM2+TAMs in promoting CD8+T cell exhaustion was determined. Genes were screened using univariate Cox regression, LASSO regression, and random forest algorithms to establish the prediction model.

Benefits of technology

Accurate prediction of prognosis after immunotherapy for head and neck squamous cell carcinoma improves the accuracy of prognostic prediction, provides a basis for personalized treatment, and reduces the recurrence rate and mortality of HNSCC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121812175A_ABST
    Figure CN121812175A_ABST
Patent Text Reader

Abstract

The invention provides a head and neck squamous cell carcinoma prognosis prediction model based on a characteristic macrophage subset and a construction method of the head and neck squamous cell carcinoma prognosis prediction model, and belongs to the technical field of biomedicine. According to the invention, an immune spectrum of LN metastatic HNSCC is firstly established, single cell sequencing identification is utilized, then a TREM < 2 + > TAMs cell population with significant prediction value for HNSCC prognosis is identified by combining microarray data, multiple immunofluorescence, an in-vitro co-culture technology and the like, and a mechanism that TREM < 2 + > TAMs induces CD8 < + > T cell depletion through an SPP1-CD44 signal axis to promote HNSCC LN metastasis is clarified. Then, single-factor COX regression and LASSO regression are jointly used, a random forest algorithm is introduced, modeling is carried out in microarray sample data based on a TREM2 + TAMs feature gene set, and the established model can accurately predict prognosis after head and neck squamous cell carcinoma immunotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, specifically to a prognostic prediction model for head and neck squamous cell carcinoma based on a characteristic macrophage subset and its construction method. Background Technology

[0002] Head and neck squamous cell carcinoma (HNSCC) is the sixth most common cancer worldwide, primarily including laryngeal, pharyngeal, or oral mucosal cancers, causing approximately 350,000 deaths annually. Despite advances in multimodal therapy, the overall 5-year survival rate for HNSCC patients remains around 50%. Among the various prognostic factors in HNSCC, lymph node (LN) metastasis is the primary mode of metastasis and an independent prognostic factor, with most HNSCC patients presenting with regional LN metastases at initial diagnosis. The 5-year survival rate for HNSCC patients decreases with increasing LN metastasis grade (N stage), particularly for patients in N2c stage (≥1 contralateral or bilateral LN ≤6 cm) or higher, where the survival rate is less than 40%. Therefore, LN is often used as one of the prognostic indicators for HNSCC. Currently, the American Joint Committee on Cancer (AJCC) TNM staging system is the primary clinical method for prognostic assessment of HNSCC patients. However, TNM staging, based solely on anatomical parameters (primary tumor size, number of lymph node metastases, and distant metastases), exhibits significant survival differences among patients at the same stage, resulting in limited prognostic accuracy. Furthermore, traditional imaging modalities (such as CT, MRI, and PET-CT) lack sufficient sensitivity and specificity for diagnosing lymph node metastases, particularly for micrometastases (<2mm in diameter) or isolated tumor cells (ITCs), leading to inaccurate staging. While postoperative pathological examination is the gold standard, it is limited by sampling errors and cannot provide molecular-level information, failing to reflect tumor heterogeneity and thus limiting its application.

[0003] The tumor immune microenvironment (TIME) plays a crucial role in tumor progression and response to immunotherapy. Suppressive TIME promotes lymphoma (LN) metastasis by facilitating the formation of metastatic niches and the colonization and spread of circulating tumor cells. The formation of suppressive TIME is characterized by the aberrant infiltration of specific cellular components, including exhausted T cells (Tex), tumor-associated macrophages (TAMs), myeloid-derived suppressor cells (MDSCs), and cancer-associated fibroblasts (CAFs). However, the dynamic characteristics of the immune landscape behind lymph node metastasis in HNSCC remain unclear, keeping HNSCC immunotherapy technology in a "broad-spectrum application" stage. Metastasis control still relies on traditional treatments, and prognostic stratification is limited by anatomical parameters such as the number, location, and size of metastatic lymph nodes, failing to reflect the "intrinsic malignancy" of metastasis. This results in limited accuracy of the constructed prognostic prediction models, thus failing to effectively reduce patient recurrence and mortality rates. Furthermore, in the construction of tumor prediction models, using Cox regression / LASSO regression alone faces the problem of insufficient training and validation sets.

[0004] Therefore, there is an urgent need to construct a prognostic prediction model for HNSCC that integrates multi-omics data to improve the accuracy of prognostic prediction and provide a basis for individualized treatment. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention establishes an immunochromatographic profile of metastatic HNSCC using multi-omics analysis, and obtains TREM2, a marker with significant predictive value for HNSCC prognosis, through microarray sequencing data, multiplex immunofluorescence, and in vitro co-culture techniques. + A prognostic prediction model for head and neck squamous cell carcinoma was constructed based on a population of TAMs, which can accurately predict the prognosis of head and neck squamous cell carcinoma after immunotherapy.

[0006] A prognostic prediction model for head and neck squamous cell carcinoma and its construction method are provided based on a characteristic macrophage subset.

[0007] To achieve the above objectives, the present invention provides the following technical solution to address the technical problem: In a first aspect, the present invention provides a method for constructing a prognostic prediction model for head and neck squamous cell carcinoma based on a characteristic macrophage subset, the method comprising the following steps: S1. Integrating single-cell and microarray data from head and neck squamous cell carcinoma to construct an immune invasion atlas, confirming CD8 enrichment in lymph node metastatic head and neck squamous cell carcinoma. + Tex was identified as TREM2 through bioinformatics analysis, multiplex immunofluorescence assay, and in vitro co-culture. + TAMs cell population; S2. Collect clinical samples of head and neck squamous cell carcinoma and identify TREM2 using multiplex immunohistochemical staining techniques. + TAMs against CD8 + Tex co-infiltration was elevated in metastatic head and neck squamous cell carcinoma of the lymph nodes, and TREM2 was determined by in vitro co-culture technique. + TAMs promote CD8 + The role of T cell exhaustion; S3. Infiltration correlation analysis, pseudo-time analysis, transcription factor analysis and cell-cell communication techniques were used to determine that TREM2⁺ TAMs were terminally differentiated phenotypes and that the key transcription factor was ETV5. S4. Validate ETV5 and SPP1 in TREM2 using Western blot and flow cytometry. + Specific expression of TREM2 in TAMs was determined using multiplex immunofluorescence, immunoprecipitation, and multiplex immunofluorescence staining techniques. + SPP1 binds to CD44 on the surface of CD8⁺ T cells in TAMs, driving exhaustion; S5. TREM2⁺TAMs infiltration was confirmed to be positively correlated with poor prognosis in a microarray cohort, and then TREM2 was integrated. + TAMs-related gene set yielded TREM2 + TAMs characteristic gene set; S6. Six genes were selected by using univariate Cox regression, LASSO regression, and random forest algorithm. The random forest algorithm was then used for training to obtain the trained model, which is the head and neck squamous cell carcinoma prognostic prediction model.

[0008] In some preferred embodiments, the single-cell transcripts are from 212 samples from the GEO and HCA databases; the microarray data are from 419 samples from the GSE65858, GSE41613, and GSE42743 datasets; and the GEO database includes GSE164690, GSE139324, GSE226620, GSE234933, GSE182227, and GSE181919.

[0009] In some preferred embodiments, the TREM2 + TAMs promote CD8 +Increased expression of exhaustion markers LAG-3, PD-1, and TIM-3 in T cells.

[0010] In some preferred embodiments, the TREM2 + TAMs induce CD8 through the SPP1–CD44 signaling pathway. + T cell depletion.

[0011] In some preferred embodiments, the TREM2 + Increased SPP1 expression and CD8 in TAMs + Increased Tex wettability is positively correlated.

[0012] In some preferred embodiments, the TREM2 + TAMs characteristic gene set is obtained through TREM2⁺ TAMs characteristic genes, ETV5 regulatory network and TREM2 + TAMs–CD8 + Tex interacting genes were obtained.

[0013] In some preferred embodiments, the six genes are THBS1, TNFAIP6, DDIT4, ADA, EXOSC4, and SERPINH1.

[0014] Secondly, the present invention utilizes the above-described construction method to construct a prognostic prediction model for head and neck squamous cell carcinoma based on a characteristic macrophage subset.

[0015] In some preferred embodiments, the head and neck squamous cell carcinoma prognostic prediction model uses TREM2 identified by single-cell transcriptome sequencing, microarray data, multiplex immunofluorescence, and in vitro co-culture techniques. + TAMs are used as risk prediction indicators.

[0016] In some preferred embodiments, the prognostic prediction model for head and neck squamous cell carcinoma includes a detection unit and an analysis unit; The detection unit detects the expression levels of THBS1, TNFAIP6, DDIT4, ADA, EXOSC4, and SERPINH1 in the sample. Analysis Unit: Loads a pre-trained prognostic prediction model for head and neck squamous cell carcinoma; inputs the expression levels of the above 6 genes, and outputs an individualized risk score. The higher the risk score, the worse the prognosis.

[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention integrates single-cell data from independent datasets using a multi-omics analysis approach, constructs an immunoassay profile of metastatic HNSCC (low-negative leukemia-associated non-carcinoma), identifies TREM2 as a gene with significant predictive value for HNSCC prognosis through single-cell sequencing, and further analyzes it using microarray data, multiplex immunofluorescence, and in vitro co-culture techniques. + TAMs cell population and elucidated TREM2 + TAMs induce CD8 via the SPP1–CD44 signal axis. + The mechanism by which T cell exhaustion promotes HNSCC LN metastasis was investigated by combining univariate Cox regression and LASSO regression with a random forest model, based on TREM2 in microarray sample data. + Modeling was performed using a set of TAMs characteristic genes, and the resulting model can accurately predict the prognosis of head and neck squamous cell carcinoma after immunotherapy.

[0018] This invention employs calibration curves and ROC curves, and uses the Youden index to define the optimal high / low risk thresholds in the ROC curve. The risk score calculated based on the model validates the model's good generalization ability in 171 microarray samples from two new datasets. Attached Figure Description

[0019] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This invention relates to single-cell identification of macrophage subsets and TREM2-based methods. + A schematic diagram illustrating the process of constructing a prognostic model for HNSCC using TAMs.

[0021] Figure 2 This is a single-cell transcriptome atlas according to an embodiment of the present invention. In the figure: A is a UMAP diagram of all single-cell transcripts; B is a UMAP diagram of the T cell subset, colored by the identified subsets; C is a heatmap showing the differentially expressed genes (DEGs) of the T cell subsets; D is CD8. + Differentiation trajectories of T cells obtained by Monocle2 analysis are colored using subsets (left) and pseudotime (right); E is a bar chart showing the proportion of T cell subsets in LN metastatic and negative tumor tissues, highlighting CD8. + Tex; F represents the infiltration analysis results of the scRNA-seq dataset, showing CD8 +Tex infiltration was elevated in LN metastatic tumor tissues, as confirmed by the Mann-Whitney U test (p < 0.01). G is a bar graph showing the composition of T cell subsets in each histological type, with CD8 highlighted in metastatic LN. + Tex.

[0022] Figure 3 This is a multiplex immunohistochemical (mIHC) analysis diagram from an embodiment of the present invention. In the diagram: H represents representative mIHC images of CD8+ Tex-infiltrated LN metastatic positive and negative tumor tissues, using CD8a, TIM-3, and LAG-3 as surface markers. Scale bar: 100 μm, insert scale bar: 50 μm; I represents mIHC quantitative CD8+ Tex-infiltrated LN metastatic tumor tissues. + The proportion of T cells in LN metastatic positive (n=7) and negative (n=7) tumor tissues. Error bars represent mean ± SD. The results were analyzed by two-way ANOVA, and p < 0.01.

[0023] Figure 4 This is a single-cell microenvironment component analysis diagram of tumor-associated myeloid cell (TAMs) subsets and lymph node metastases (LNs) according to an embodiment of the present invention. In the diagram: A is the correlation matrix of the infiltration ratio among different cell types, which is p<0.05 according to Spearman's test; B is the correlation matrix of the infiltration ratio between TAMs and T cell subsets, which is *p<0.05 and **p<0.01 according to Spearman's test; C is TREM2. + The proportion of TAMs in LN metastatic and negative tumor tissues, highlighting TREM2 + TAMs.

[0024] Figure 5 This is a multiplex immunohistochemical (mIHC) analysis image from an embodiment of the present invention. In the image: A is an mIHC image showing CD8. + Tex and TREM2 + TAMs co-infiltrated in both LN metastatic and LN-negative tumor tissues. Scale bar: 50 μm, insert: 20 μm; B represents TREM2 in both LN metastatic and LN-negative tumors. + The proportion of TAMs to total TAMs is shown in the error bars, which represent the mean ± SD. Two-way ANOVA was performed, and **p < 0.01; C represents the Spearman test TREM2. + TAMs and CD8 + Correlation of Tex infiltration levels in mIHC

[0025] Figure 6The CD8 stimulation of TAMs subsets in this embodiment of the invention + Flow cytometry analysis of T cell exhaustion markers. Two-way ANOVA. Error bars represent mean ± SD. ∗∗p<0.01.

[0026] Figure 7 This is a set of pseudo-temporal trajectory maps and transcription factor activity heatmaps of TAMs subsets in an embodiment of the present invention. In the figure: A is the TAMs subset trajectory analysis and TAMs subset distribution along pseudo-time based on scTour; B is the expression map of the marker genes of the main TAMs subsets along pseudo-time; C is a dot plot showing the top TF regulatory network with the highest rss in the TAMs subsets.

[0027] Figure 8 This is a graph showing the enrichment of TAMs for immunosuppressive function and the validation analysis of the key transcription factor ETV5 in an embodiment of the present invention. In the graph: D represents the TREM2 values ​​between the LN metastasis-positive group and the LN metastasis-negative group, as determined by the Mann-Whitney U test. + Regulon scores for TAMs-specific TFs showed no statistical significance (ns), *p<0.05, **p<0.01; E is the Kaplan-Meier survival curve of ETV5 modulator scores generated from the microarray dataset, with samples divided into high and low groups using the log-rank test (optimal cutoff point); F and G are comparisons of GSEA results with TREM2. + ETV5-regulated gene sets in TAMs and other TAMs subpopulations; H represents Western blot analysis of ETV5 expression in TAMs subpopulations.

[0028] Figure 9 This is embodiment TREM2 of the present invention. + TAMs communicate with CD8 via the SPP1–CD44 signal path. + The validation results of Tex interactions are shown in the figure. In the figure, A is a circular plot depicting the crosstalk between TAMs and T cell subsets, and the width represents the weight of the interaction strength. B is a circular plot depicting CD8. + Crosstalk between Tex and TAMs subsets, with width representing the weight of the interaction strength; C is a heatmap describing the interaction path between TREM2⁺ TAMs and CD8⁺ Tex, with the output signal (left) and input signal (right) colored by the interaction strength; D is a heatmap showing the interaction probability of the SPP1 signal path between subsets; E is a chord diagram showing the interaction strength of SPP1–CD44 between subsets.

[0029] Figure 10This is a verification diagram of the SPP1–CD44 signaling pathway expression in an embodiment of the present invention. In the diagram: F is a violin plot showing the expression levels of SPP1 and other ligands in the TAMs subset; G and H are flow cytometry analyses showing SPP1 expression in the TAMs subset (n=4). Two-way ANOVA was performed, and the error bars represent the mean ± SD, with p < 0.01.

[0030] Figure 11 This is embodiment TREM2 of the present invention. + TAMs promote CD8 via SPP1–CD44 + The image shows the validation results of T cell exhaustion. In the image, A is an mIHC image showing CD8. + Tex and SPP1 + TREM2⁺ TAMs in LN-positive and LN-negative tumor tissues, scale bar: 100 μm, insert: 50 μm; B represents TREM2 in mIHC LN-positive and LN-negative tumors. + SPP1 + The proportion of TAMs in total TAMs was determined by two-way ANOVA, *p<0.05; C represents the TREM2 concentration in mIHC detected using the Spearman method. + SPP1 + TAMs and CD8 + Correlation of Tex infiltration ratio.

[0031] Figure 12 This is a functional experimental diagram from an embodiment of the present invention verifying that TREM2⁺ TAMs induce CD8⁺ T cell exhaustion via SPP1–CD44. In the diagram: D and E represent CD8⁺ T cells. + T cells co-cultured with TAMs subsets, flow cytometry analysis of cell exhaustion markers with or without anti-SPP1 or anti-CD44 antibodies, analyzed by two-way ANOVA, error bars represent mean ± SD, ∗p<0.01; F represents Western blot analysis of SPP1 and CD44 binding in co-IP assay; G represents immunofluorescence showing SPP1 and CD44 in CD44-silenced or unsilenced CD8+. + Colocalization on T cells, scale bar: 50 μm.

[0032] Figure 13 This is a Kaplan-Meier survival curve from an embodiment of the present invention. In the figure: A represents the survival curve based on three microarray datasets TREM2. +Kaplan-Meier survival curves for TAMs infiltration scores. The optimal cutoff value using the log-rank test stratified the samples into high-expression and low-expression groups; B represents the Kaplan-Meier survival curves based on SPP1 expression levels in three microarray datasets. The optimal cutoff value was used to stratify the samples into high-expression and low-expression groups using the log-rank test.

[0033] Figure 14 This is a feature importance graph of the random forest screening of core genes and model construction in an embodiment of the present invention. In the graph: C is the LASSO coefficient distribution of 119 genes in the training set (GSE65858); D is the optimal value of the tenfold cross-validation identification penalty parameter (λ); E is the relationship between the RSF model error rate and the number of classification trees; F is the importance score of the first 6 genes selected by the RSF model.

[0034] Figure 15 The figure shows the time-dependent performance verification results of the prognostic model in the embodiment of the present invention. In the figure: A is the calibration curve of the 1-year OS prediction of the training set and the validation set; B is the ROC curve and AUC value of the predicted 1, 3 and 5-year OS over time using the Kaplan-Meier method in the training set and the validation set.

[0035] Figure 16 This diagram shows the expression heatmap of risk-related genes in high- and low-risk groups according to embodiments of the present invention, and the Kaplan-Meier survival curves based on risk scores. In the diagram: C is the heatmap of risk-related gene expression in the training and validation sets, using the Youden index for log-rank testing to divide the samples into high- and low-risk groups; D is the K-Kaplan-Meier survival curve based on the risk scores of the training and validation sets. The log-rank test, using the Youden index for truncation, divides the samples into high- and low-risk groups. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to specific examples. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Unless otherwise specified, the experimental methods used in this invention are conventional methods; unless otherwise specified, the materials and reagents used are commercially available.

[0038] The method used in this invention is as follows: Pseudotime analysis: To infer cell differentiation trajectories, we applied several trajectory inference methods, including Monocle2 and scTour. Monocle2... Monocle The R package scTour is used for trajectory inference and gene-level analysis. It utilizes Python (version 3.12) and includes... scTour The package captures the dynamic transcriptional program during differentiation.

[0039] Regulon analysis: To identify characteristic transcription factors (TFs) within TAMs subsets, we applied the SCENIC R package to infer gene regulatory networks. We first identified co-expressed modules and then linked them to candidate TFs to define potential regulators, each composed of a TF and its predicted target genes. Regulator-specific scores (RSS) were calculated to assess the specificity of regulator activity across different subsets.

[0040] Cell-cell communication analysis: use CellChat This R package is used to study intercellular communication within TIME. A CellChat object is constructed using a selected subset as input, and all preprocessing steps are performed with default parameters. The resulting communication network is visualized using pie charts, bar charts, heatmaps, and chord plots, providing a comprehensive view of the intercellular signaling landscape across subsets.

[0041] Immunofluorescence staining analysis (mIHC and immunofluorescence staining analysis): For mIHC staining of tumor tissues, a series of primary HNSCC tumor sections were used to assess the spatial distribution and relative localization of subpopulations. Sections were dewaxed, rehydrated, and subjected to antigen recovery in EDTA buffer at high temperature for 15 minutes. Sections were then blocked with goat serum for 30 minutes. Sections were incubated overnight with primary antibody at 4°C, and with secondary antibody incubation using Opal Tyramide Signal Amplification (TSA) dye. Weaker positive markers were labeled with stronger dyes. Primary antibodies used included CD8a (Proteintech, 66868-1-Ig), TIM-3 (CST, 45208S), LAG-3 (Abcam, ab209236), CD68 (Proteintech, 65515-1-RR), TREM2 (Abcam, ab318262), and SPP1 (Proteintech, 83341-1-RR). TSA dyes were PPD480, PPD570, PPD620, and PPD690, respectively. The above procedure was repeated three times. Cell nuclei were then stained with 4′,6-diamino-2-phenylindole (DAPI). Sections were imaged using a digital slide scanner (Pannoramic MIDI, 3DHISTECH).

[0042] For immunofluorescence staining of cells, cells cultured in confocal dishes were fixed with 4% paraformaldehyde and then blocked with goat serum at room temperature for 1 hour to prevent nonspecific binding. The samples were then incubated overnight at 4°C with the indicated primary antibody, followed by incubation with a fluorescently labeled secondary antibody at room temperature for 1 hour. The cell nuclei were then stained with DAPI at room temperature for 15 minutes. Cells were imaged using a laser scanning confocal microscope (Olympus FV3000).

[0043] Flow cytometry (FACS of TAMs and flow cytometry): Flow cytometry detection of CD8 + T cell exhaustion and SPP1 in TREM2 + Expression in TAMs. In co-culture experiments, CD8 was collected from the suspension fraction of the co-culture medium. +T cells. Cells were processed into single-cell suspensions, stained with live / dead dyes (BD Biosciences), and then incubated at room temperature for 20 minutes with antibodies targeting specific surface markers. The following antibodies were used for CD8 (biolgend, 344704), LAG-3 (BD ​​Biosciences, 565716), PD-1 (BD Biosciences, 563245), and TIM-3 (BD ​​Biosciences, 565562) for CD8. + T cells; TREM2 (R&D Systems, FAB17291P) and SPP1 (eBioscience, 3002684) preparation of TREM2 + TAMs. To detect intracellular proteins, cells were fixed and permeated with flow cytometry intracellular fixation and lysis buffer (eBioscience) before antibody staining. After staining, the expression levels of the specified markers were detected using a FACScan flow cytometer (BD FACSCelesta; Beckman Coulter).

[0044] Protein extraction and Western blotting analysis: To extract proteins, cells were collected, washed three times with PBS, and thawed on ice for 30 minutes in RIPA buffer containing protease and phosphatase inhibitors. Lysis buffer was centrifuged at 12,000 × g for 30 min at 4°C, and the supernatant was used for subsequent analysis. Isolated membrane proteins were detected using the Minut Plasma Membrane Protein Separation and Cell Separation Kit (Invent Biotechnologies) according to the manufacturer's instructions. Protein concentrations were determined using the BCA Protein Assay Kit (CWBio). For western blot analysis, 20 μg of total protein per sample was dissolved in 10% SDS-PAGE and then transferred to a PVDF membrane. After blocking with goat serum to prevent nonspecific binding, the membrane was incubated overnight at 4°C with the designated primary antibody, followed by incubation at room temperature for 1 hour with enzyme-labeled secondary antibody. The following antibodies were used: ETV5 (Proteintech, 13011-1-AP), SPP1 (Abcam, ab214050), and CD44 (Abcam, ab316123). Protein signals were visualized using an ECL detection kit and quantified using ImageJ software.

[0045] co-IP assays: For protein-protein interaction analysis between SPP1 and CD44, co-IP analysis was used. After 48 h of co-culture, CD84 cells bound to SPP1 were harvested from the lower chamber of the Transwell system. + T cells. Approximately 1 × 10⁻⁶ cells lysed. 7 Sorted CD8 + T cells were used to extract membrane proteins. Anti-SPP1 antibody (Abcam, ab214050) was then added and incubated overnight at 4°C. The immune complexes were then captured using Protein A / G agarose beads. After thorough washing, the bound proteins were eluted from the beads and subjected to Western blot analysis to verify the interaction between SPP1 and CD44.

[0046] Survival analysis: Overall survival data were derived from the microarray dataset. Enrichment scores for the gene set were calculated using GSVA, and the association between the target gene set and individual genes or risk scores calculated by the prognostic model and patient survival was assessed. survival and survminer The R software package was used to generate Kaplan-Meier survival curves to assess the prognostic significance of the selected variables. Patients were divided into high- and low-risk groups based on the optimal cutoff point or the Youden index cutoff point, and the difference was assessed using the log-rank test.

[0047] Predictive model construction: The publicly available microarray dataset (GSE65858) containing 248 samples was used as the training set. A primary gene set consisting of 1049 genes analyzed by scRNA-seq was used as the input feature set. The first step involved using... survival We performed univariate Cox proportional hazards regression using the R package to identify genes significantly associated with overall survival (OS) (p<0.05). To enhance the interpretability and robustness of the model, we then used... glmnet The R package (alpha = 1, nlambda = 100) applies LASSO regression with minimum absolute shrinkage and selection operator and performs 10-fold cross-validation. Regularization parameters are selected based on λ. A more concise prediction model is obtained by preserving the genes with the most information. Following this, we employ the Random Survival Forest (RSF) algorithm, using... randomForestSRC The R package (ntree=1000, nodesize=10) further improves the model. Only genes with positive importance values ​​are retained for downstream modeling. These genes are then used to build the final prognostic model using the R package. survcomp We used the consistency index (C-index) to conduct an initial evaluation to assess its predictive performance on the training dataset itself.

[0048] Prognostic model validation: The prognostic model was evaluated on the training set (GSE65858) and two independent validation sets (GSE41613 and GSE42743). First, using... rms The R package was used to generate calibration curves to assess the consistency between predicted and observed survival probabilities. Subsequently, using... survivvalroc The R package was used for time-dependent receiver operating characteristic (ROC) curve analysis to evaluate the model's discriminative ability by comparing predicted survival probabilities with actual outcomes at different time points. To further evaluate the model's predictive ability, we used risk scores calculated by the prognostic model and, based on the cutoff value determined by the Youden Index derived from the ROC analysis, divided patients into high-risk and low-risk groups. Kaplan-Meier survival analysis was applied to compare the overall survival rates of the two risk groups, and the log-rank test was used to assess statistical significance and verify the correlation between predicted risk scores and actual survival outcomes.

[0049] Statistical analysis: Statistical analyses were performed using R (version 4.41) and GraphPad Prism44. All experiments were conducted in three independent biological replicates; detailed methodologies are provided in the corresponding subsection of the Methods section. Two-way ANOVA and the Mann-Whitney U test were used to assess statistical significance between groups. Pearson correlation analysis was used to evaluate the correlation between two variables. The log-rank test was used to compare survival curves. A p-value less than 0.05 was considered statistically significant, *p<0.05, and **p<0.01.

[0050] Example 1

[0051] I. Source of Sample Data A total of 688,866 single-cell transcripts from 212 HNSCC samples from 7 publicly available datasets (GSE164690, GSE139324, GSE226620, GSE234933, GSE182227, GSE181919, HCA) and microarray data from 419 samples from 3 publicly available datasets (GSE65858, GSE41613, GSE42743) were used to construct a complete HNSCC immune infiltration atlas (see [link to dataset]). Figure 1Immune infiltration in LN metastatic lesions is mainly driven by the CD8⁺ Tex subset. Differential analysis of immune infiltration and mIHC staining revealed that LN metastatic HNSCC is characterized by high infiltration of CD8⁺ exhausted T cells (Tex) and a suppressive TIME (see [link to study]). Figure 2 ,3), and identified a subset of TREM2⁺ tumor-associated macrophages (TAMs) associated with LN metastasis (see Figure 4 ).

[0052] II. Collect clinical HNSCC cohorts and verify TREM2 using multiplex immunohistochemical (mIHC) staining. + TAMs and CD8 + Increased co-infiltration of Tex in LN-transferred HNSCC (see...) Figure 5 Furthermore, TREM2 was separated by in vitro flow cytometry. + TAMs and CD8 + T cells were then used to construct an in vitro co-culture model, and TREM2 was discovered. + TAMs promote CD8 + T cell exhaustion markers LAG-3, PD-1, and TIM-3 (detected by flow cytometry, validating TREM2) + TAMs promote CD8 + T cell exhaustion) expression is elevated (see... Figure 6 ).

[0053] III. Combining infiltration correlation analysis, Monocle2 / scTour pseudo-time-sequence cell trajectory inference, SCENIC transcription factor analysis, and Cellchat intercellular communication analysis, the terminal differentiation phenotype of TREM2⁺TAMs and the key transcription factor ETV5 were elucidated, and its biological potential to promote CD8+ T cell exhaustion through the SPP1–CD44 signaling pathway, thereby driving HNSCC LN metastasis, was revealed (see...). Figure 7 , 8 9).

[0054] IV. Detection of TREM2 by Western blot and flow cytometry analysis + The expression levels of ETV5 and SPP1 in TAMs were used to verify the role of these two molecules in TREM2. + Specific expression in TAMs (see) Figure 8 ,10); TREM2 was discovered using multiplex immunofluorescence staining. + Increased SPP1 expression and CD8 in TAMs + Increased Tex infiltration is positively correlated with (see...) Figure 11 And it was found that blocking TREM2...+ Both SPP1 secreted by TAMs and blocking CD44 can weaken TREM2. + TAMs induce CD8 + The role of T cell exhaustion (using neutralizing antibodies) was then verified using co-immunoprecipitation and multiplex immunofluorescence staining. + SPP1 derived from TAMs specifically binds to CD8. + CD44 on the surface of T cell membrane (see...) Figure 12 This confirms TREM2 + TAMs induce CD8 through the SPP1–CD44 signaling pathway. + The mechanism by which T cell depletion drives HNSCC LN metastasis.

[0055] V. Through survival analysis of microarray data, TREM2 was found. + High TAM infiltration is indeed significantly positively correlated with poor prognosis in HNSCC (see...). Figure 13 Based on the above results, the following three TREM2... + TAMs-related gene set: (1)TREM2 + TAMs' characteristic gene set (536 genes); (2) TREM2 + The core transcription factor ETV5 regulatory network of TAMs (578 genes); (3) TREM2 + TAMs and CD8 + Tex interacting ligand-receptor pairs (71 genes) were intersected to obtain 1049 genes, which were used to construct TREM2. + TAMs characteristic gene set (see Figure 1 ).

[0056] VI. A combined approach using multiple model building methods was employed to construct a prognostic prediction model for HNSCC. A microarray dataset containing 248 samples was used as the training set. First, univariate Cox regression was used to analyze the data from TREM2. + From the TAMs feature gene set (1049 genes), 119 genes significantly associated with patient survival were screened. LASSO regression analysis was then used to identify 8 genes that significantly contributed to prognostic prediction. Random forest analysis was then used to identify 6 genes that positively contributed to model construction (THBS1, TNFAIP6, DDIT4, ADA, EXOSC4, SERPINH1), and a random forest prognostic prediction model was built based on these 6 genes (see...). Figure 14The model was validated using an independent test set containing 171 microarray samples. Results showed that the model maintained high predictive consistency on new data, demonstrating excellent prognostic predictive ability and promising clinical translation prospects (see...). Figure 15 , 16 ).

[0057] In this embodiment, Using survival status and survival time as target variables, a random forest risk prediction model rf_model was constructed using 248 samples from the training set GSE65858 and six target genes (THBS1, TNFAIP6, DDIT4, ADA, EXOSC4, SERPINH1) selected from the screening.

[0058] Random Forest (RF) is a classification and regression algorithm model based on ensemble learning, consisting of multiple independent decision trees. When applied to a single new sample or a new sample queue, the input to the Random Forest model (rf_model) should be a matrix or data frame consisting of the expression levels of the six target genes mentioned above. These gene expression levels are obtained through microarray detection and logarithmic transformation. The final output of the Random Forest model (rf_model) is the predicted risk score for that sample. The risk score is obtained by weighted averaging of the predictions from all decision trees in the model, thus ensuring the stability of the prediction results.

[0059] The random forest model rf_model of this invention accepts input data in the following format: Line name: Sample name (e.g., sample_1, sample_2, etc.); Column names: c("THBS1", "TNFAIP6", "DDIT4", "ADA", "EXOSC4", "SERPINH1"); Expression levels: Log2 transformed expression levels of the six genes on the microarray platform.

[0060] Example formats are shown in Table 1 below: Table 1

[0061] Based on the model prediction results, this invention further establishes risk stratification criteria. Through analysis of a large number of real samples and statistical verification, this invention determines the risk score threshold to be 26.45221.

[0062] The specific hierarchical rules are as follows: When risk_score > 26.45221, the sample is defined as a high-risk group; When risk_score ≤ 26.45221, the sample is defined as the low-risk group.

[0063] The above thresholds can effectively distinguish individuals with different risk levels, which is helpful for HNSCC prognostic assessment and subsequent clinical decision-making.

[0064] In summary, this invention integrates single-cell transcript and microarray data from multiple publicly available datasets to construct a complete immune infiltration atlas of HNSCC. Infiltration abundance analysis of large-scale single-cell data revealed a repressive TIME characterized by high infiltration of CD8⁺ exhausted T cells (Tex) in metastatic LN HNSCC, and identified a TREM2⁺ tumor-associated macrophage (TAM) subset associated with LN metastasis. Furthermore, by combining infiltration correlation analysis, Monocle2 / scTour pseudo-time-sequence cell trajectory inference, SCENIC transcription factor analysis, and Cellchat intercellular communication analysis, the terminal differentiation phenotype of TREM2⁺ TAMs and the key transcription factor ETV5 were elucidated, revealing that it may induce CD8⁺ T cells through the SPP1–CD44 signaling pathway. + T cell depletion, thereby driving the biological potential of HNSCC LN metastasis.

[0065] This invention collects clinical HNSCC cohorts and verifies TREM2 through multiple immunohistochemical staining. + TAMs and CD8 + Increased co-infiltration of Tex in LN-transferred HNSCC. Furthermore, TREM2 was observed through the construction of an in vitro co-culture model. + TAMs promote CD8 + Elevated levels of exhaustion markers LAG-3, PD-1, and TIM-3 in T cells were observed to validate TREM2. + TAMs against CD8 + T cell induction.

[0066] This invention also detects the interaction between ETV5 and SPP1 in TREM2 using Western blot and flow cytometry analysis. + Specific expression of TAMs; and detection of TREM2 using multiplex immunofluorescence staining. + Increased SPP1 expression and CD8 in TAMs + A positive correlation was found between increased Tex infiltration and increased TREM2. + Both SPP1 secreted by TAMs and blocking CD44 can weaken TREM2. + TAMs induce CD8 + The role of T cell exhaustion (using neutralizing antibodies) was then verified using co-immunoprecipitation and multiplex immunofluorescence staining.+ SPP1 derived from TAMs specifically binds to CD8. + CD44 on the surface of T cell membrane; thus confirming TREM2 + TAMs induce CD8 through the SPP1–CD44 signaling pathway. + The mechanism by which T cell depletion drives HNSCC LN metastasis.

[0067] This invention uses a training set of 248 samples and combines univariate Cox regression, LASSO regression analysis, and random forest analysis to select six predictive genes (THBS1, TNFAIP6, DDIT4, ADA, EXOSC4, and SERPINH1) to construct a random forest prognostic prediction model. The model's high predictive accuracy and good generalization ability are validated using an independent test set containing 171 samples.

[0068] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for constructing a prognostic prediction model for head and neck squamous cell carcinoma based on characteristic macrophage subsets, characterized in that, The construction method includes the following steps: S1. Integrating single-cell and microarray data from head and neck squamous cell carcinoma to construct an immune invasion atlas, confirming CD8 enrichment in lymph node metastatic head and neck squamous cell carcinoma. + Tex was identified as TREM2 through bioinformatics analysis, multiplex immunofluorescence assay, and in vitro co-culture. + TAMs cell population; S2. Collect clinical samples of head and neck squamous cell carcinoma and verify TREM2 using multiplex immunohistochemical staining techniques. + TAMs and CD8 + Tex showed elevated infiltration in metastatic head and neck squamous cell carcinoma of the lymph nodes, and TREM2 was determined by in vitro co-culture technique. + TAMs promote CD8 + The role of T cell exhaustion; S3. Infiltration correlation analysis, pseudo-time analysis, transcription factor analysis and cell-cell communication technology were used to determine that TREM2⁺TAMs were terminally differentiated phenotypes and the key transcription factor was ETV5. S4. Verify ETV5 and SPP1 in TREM2 using Western blot and flow cytometry analysis. + Specific expression of TREM2 in TAMs was determined using immunoprecipitation and multiplex immunofluorescence staining techniques. + SPP1 binds to CD44 on the surface of CD8⁺ T cells in TAMs, driving exhaustion; S5. TREM2⁺ TAMs infiltration was confirmed to be positively correlated with poor prognosis in a microarray cohort, and then TREM2 was integrated. + TAMs-related gene set yielded TREM2 + TAMs characteristic gene set; S6. Six genes were selected by using univariate Cox regression, LASSO regression, and random forest algorithm. The random forest algorithm was then used for training to obtain the trained model, which is the head and neck squamous cell carcinoma prognostic prediction model.

2. The construction method according to claim 1, characterized in that, The single-cell transcripts were from 212 samples from the GEO and HCA databases; the microarray data were from 419 samples from the GSE65858, GSE41613, and GSE42743 datasets; the GEO database included GSE164690, GSE139324, GSE226620, GSE234933, GSE182227, and GSE181919.

3. The construction method according to claim 1, characterized in that, The TREM2 + TAMs promote CD8 + Increased expression of exhaustion markers LAG-3, PD-1, and TIM-3 in T cells.

4. The construction method according to claim 1, characterized in that, The TREM2 + TAMs induce CD8 through the SPP1–CD44 signaling pathway. + T cell depletion.

5. The construction method according to claim 1, characterized in that, The TREM2 + Increased SPP1 expression and CD8 in TAMs + Increased Tex wettability is positively correlated.

6. The construction method according to claim 1, characterized in that, The TREM2 + TAMs characteristic gene set is obtained through TREM2⁺ TAMs characteristic genes, ETV5 regulatory network and TREM2 + TAMs–CD8 + Tex-interacting genes were obtained.

7. The construction method according to claim 1, characterized in that, The six genes are THBS1, TNFAIP6, DDIT4, ADA, EXOSC4, and SERPINH1.

8. A prognostic prediction model for head and neck squamous cell carcinoma based on a characteristic macrophage subset, constructed using the construction method described in any one of claims 1-7.

9. The prognostic prediction model for head and neck squamous cell carcinoma according to claim 8, characterized in that, The prognostic prediction model for head and neck squamous cell carcinoma, TREM2, was identified using single-cell transcriptome sequencing, microarray data, multiplex immunofluorescence, and in vitro co-culture techniques. + TAMs are used as risk prediction indicators.

10. The prognostic prediction model for head and neck squamous cell carcinoma according to claim 8, characterized in that, The prognostic prediction model for head and neck squamous cell carcinoma includes a detection unit and an analysis unit; The detection unit detects the expression levels of THBS1, TNFAIP6, DDIT4, ADA, EXOSC4, and SERPINH1 in the sample. Analysis Unit: Loads a pre-trained prognostic prediction model for head and neck squamous cell carcinoma; inputs the expression levels of the above 6 genes, and outputs an individualized risk score. The higher the risk score, the worse the prognosis.