Head and neck squamous cell carcinoma prognosis and immunotherapy response applicability prediction marker and application thereof

By constructing mRNA marker models based on DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E and FOXD1, the accuracy of prognosis and immunotherapy response of squamous cell carcinoma in the head and neck was solved, and the treatment effect and survival rate were improved.

CN120272595APending Publication Date: 2025-07-08ANHUI PROVINCIAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510433924.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art In the treatment of head and neck squamous cell carcinoma, it is difficult to effectively optimize the treatment response rate and improve the overall survival rate, especially in the prediction of the applicability of immunotherapy.

Method used

A prediction model for prognosis and immunotherapy response suitability of squamous cell carcinoma in the head and neck was constructed. Through the gene expression data set, a random survival forest algorithm was used to construct a cumulative risk function, combined with mRNA markers such as DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E and FOXD1, risk scores were performed to predict patients' prognosis and immunotherapy response.

Benefits of technology

It has achieved accurate prediction of the prognosis of patients with squamous cell carcinoma of the head and neck and effective evaluation of immunotherapy response, providing a reliable basis for formulating treatment plans, and improving the treatment response rate and survival rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120272595A_ABST
    Figure CN120272595A_ABST
Patent Text Reader

Abstract

The invention discloses a head and neck squamous cell carcinoma prognosis and immunotherapy response applicability prediction marker and application thereof, the marker is an mRNA marker, and the mRNA marker comprises DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E and FOXD1. The invention further discloses a model for predicting the prognostic risk and the immunotherapy applicability of the head and neck squamous cell carcinoma, a system for predicting the prognostic risk and the immunotherapy applicability of the head and neck squamous cell carcinoma and a kit for predicting the prognostic risk and the immunotherapy applicability of the head and neck squamous cell carcinoma. According to the method, markers DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E and FOXD1 related to head and neck squamous cell carcinoma progression are found through a bioinformatics technology, then a model for predicting the head and neck squamous cell carcinoma prognosis risk and immunotherapy applicability is constructed according to the eight genes, and then a risk score is obtained. The prognosis of the head and neck squamous cell carcinoma and the applicability of immunotherapy can be predicted through risk scoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of gene technology and medical technology, and in particular to a prognostic and immunotherapy response applicability prediction marker for head and neck squamous cell carcinoma and its application. Background Art

[0002] Head and neck squamous cell carcinoma (HNSCC) is one of the common malignant tumors worldwide, including multiple subtypes such as oral cancer, laryngeal cancer, and pharyngeal cancer. According to the statistics of the World Health Organization (WHO), head and neck tumors account for 6% ± of all cancer cases and are highly prevalent in the Asian region. Despite significant progress in multidisciplinary treatment (including surgery, radiotherapy, and chemotherapy), for example, immune checkpoint inhibitors have become the first-line or second-line treatment for recurrent and metastatic HNSCC. However, optimizing the treatment response rate and improving the overall survival rate remain major challenges. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems in the related art to some extent. To this end, an object of the present invention is to propose a prognostic and immunotherapy response applicability prediction marker for head and neck squamous cell carcinoma. By constructing a prognostic and immunotherapy response applicability prediction model for head and neck squamous cell carcinoma, a risk score can be obtained, and the prognosis of head and neck squamous cell carcinoma and the applicability of immunotherapy can be predicted through the risk score.

[0004] In a first aspect, a prognostic and immunotherapy response applicability prediction marker for head and neck squamous cell carcinoma proposed by the present invention, the marker is an mRNA marker, and the mRNA marker includes DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1.

[0005] In a second aspect, an application of the prognostic and immunotherapy response applicability prediction marker for head and neck squamous cell carcinoma proposed by the present invention, the application includes evaluating or predicting prognostic risk, predicting the applicability of immunotherapy, predicting survival rate, formulating a treatment / medication plan, constructing a model for predicting the prognostic risk of head and neck squamous cell carcinoma, constructing a model for the applicability of immunotherapy, constructing a model for predicting the survival rate of head and neck squamous cell carcinoma, preparing a detection reagent or device for predicting the prognostic risk of head and neck squamous cell carcinoma, and preparing a detection reagent or device for predicting the survival rate of head and neck squamous cell carcinoma, any one or a combination of several of them.

[0006] In a third aspect, a prognostic and immunotherapy response applicability prediction model for head and neck squamous cell carcinoma proposed by the present invention, which contains the above-mentioned prognostic and immunotherapy response applicability prediction marker for head and neck squamous cell carcinoma. The method steps for constructing the prognostic and immunotherapy response applicability prediction model for head and neck squamous cell carcinoma are as follows:

[0007] S1: Obtain a gene expression dataset: Obtain head and neck squamous cell carcinoma samples and obtain the mRNA expression data of each sample;

[0008] S2: Define the survival data of patients as survival time and survival status, where death is defined as 1 and survival is defined as 0;

[0009] S3: Use the random survival forest algorithm, which consists of N tree decision trees, and each tree is constructed according to the following rules:

[0010] S31: Randomly select a target percentage of samples from the total samples collected in step S1 as bootstrap samples, and the remaining samples are used as out-of-bag samples to verify the model performance;

[0011] S32: Randomly select some gene expression variables at each node, use the Log-rank statistic as the criterion to automatically select the best split point, split the samples within the node to maximize the difference in survival curves between nodes, and recursively split the data until the number of samples in the node reaches the preset minimum value;

[0012] S4: The cumulative risk function constructed for each tree is expressed as follows: At each node during the decision tree splitting process, according to the Log-rank statistic, the best split point is selected, the samples are divided into subgroups with different risk levels, and the data is recursively split until the minimum node sample number requirement is met. The occurrence frequency of survival events within the node is used to estimate the node-specific risk, and finally, it is summarized into the overall cumulative risk function of the tree, specifically expressed as:

[0013] h b (DKK1 i ,STC2 i ,INHBA i ,AMIGO2 i ,VEGFC i ,SPOCK1 i ,MT1E i ,FOXD1 i )

[0014] S5: For the overall risk score RS of patient i, the calculation method is the average of the cumulative risk functions of all trees:

[0015]

[0016] where: RS i represents the predicted risk score of the i-th patient; N tree represents the total number of decision trees in the random survival forest; h b () represents the cumulative risk function calculated by the b-th tree, which is automatically constructed by the algorithm based on the gene expression level;

[0017] S6: Evaluate the prediction accuracy and generalization ability of the model through the consistency index C-index statistical indicator.

[0018] Preferably, N in step S3 tree is 1000, the target percentage of bootstrap samples in step S31 is 63%, the percentage of out-of-bag samples is 37%, and the preset minimum value of the number of node samples in step S32 is 5.

[0019] Fourthly, a prediction system for the prognosis and immune therapy response suitability of head and neck squamous cell carcinoma proposed by the present invention includes any one of the above-mentioned schemes of the prediction model for the prognosis and immune therapy response suitability of head and neck squamous cell carcinoma. The system further includes a data input module, a model calculation module, and a result output module. The data input module is used to input the expression value results of DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1 into the model calculation module. The model calculation module outputs the RS i value, and defines the patient prediction risk score LRS = RS i . The result output module is used to predict the prognosis risk and immune therapy suitability according to the value of the patient prediction risk score LRS.

[0020] Preferably, when the patient prediction risk score LRS > 53.89, it is determined as a high risk of prognosis of head and neck squamous cell carcinoma, indicating poor prognosis and poor immune therapy response of the patient. When the patient prediction risk score ≤ 53.89, it is determined as a low risk of prognosis of head and neck squamous cell carcinoma, indicating good prognosis and good immune therapy response of the patient.

[0021] Fifthly, a prediction kit for the prognosis and immune therapy response suitability of head and neck squamous cell carcinoma proposed by the present invention includes any one of the above-mentioned schemes of the prediction model for the prognosis and immune therapy response suitability of head and neck squamous cell carcinoma.

[0022] Sixthly, the method steps for predicting the immune therapy risk using the above kit are as follows:

[0023] A1: Detect the expression levels of DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1 in the head and neck squamous cell carcinoma patient samples, and perform standardization processing;

[0024] A2: Substitute the standardized gene expression levels obtained in step S1 into the model for predicting the prognosis risk and immune therapy suitability of head and neck squamous cell carcinoma, and define the patient prediction risk score LRS = RS i , and calculate the risk score LRS;

[0025] A3: Determine whether the risk score LRS is greater than 53.89. When the determination result is yes, execute step A4; when the determination result is no, execute step A5;

[0026] A4: This head and neck squamous cell carcinoma patient belongs to the high-risk group, indicating that this head and neck squamous cell carcinoma patient is not suitable for immunotherapy;

[0027] A5: This head and neck squamous cell carcinoma patient belongs to the low-risk group, indicating that this head and neck squamous cell carcinoma patient is suitable for immunotherapy.

[0028] Preferably, the sample includes tissues and body fluids.

[0029] Preferably, the body fluid is blood.

[0030] In the present invention, the determination of the expression levels of DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1 follows the established standard procedures well-known in the art (Sambrook, J. et al. (1989) Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Ausubel, F. M. et al. (2001) Current Protocols in Molecular Biology, Wiley & Sons, Hoboken, NJ). The determination can be carried out at the RNA level, for example, by transcriptome high-throughput sequencing, or by real-time fluorescence quantitative PCR technology to detect the cDNA level after RNA reverse transcription, etc. The sequences of DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1 disclosed in the present invention are all stored in the Gene Expression Omnibus database (http: / / www.ncbi.nlm.nih.gov / geo / ).

[0031] The beneficial effects of the present invention are:

[0032] (1) By using bioinformatics techniques, the present invention finds the markers DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1 related to the progression of head and neck squamous cell carcinoma, then constructs a prognostic and immunotherapy response applicability prediction model for head and neck squamous cell carcinoma based on these 8 genes, and further obtains a risk score. Through the risk score, the prognosis of head and neck squamous cell carcinoma and the applicability of immunotherapy can be predicted;

[0033] (2) The present invention proposes a prognostic model for predicting head and neck squamous cell carcinoma patients composed of 8 genes as biomarkers, and the present invention provides a reliable method for analyzing the prognosis and immunotherapy of head and neck squamous cell carcinoma patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In the drawings:

[0035] Figure 1 A is a KEGG enrichment analysis chart showing metabolites of the blue module;

[0036] Figure 1 B-C are distribution charts showing adenosine in HNSCC tissues;

[0037] Figure 2 A is a differential expression profile chart of 31 upstream regulators of purine metabolism in normal samples (n = 44) and tumor samples (n = 509) in the TCGA-HNSCC dataset;

[0038] Figure 2 B is a univariate Cox regression analysis chart of 23 purine metabolism regulators in the TCGA and merge-GEO datasets;

[0039] Figure 2 C is a biological pathway analysis chart of the Fadu cell line after AB680 treatment;

[0040] Figure 2 D is a receptor-ligand interaction bubble chart between different epithelial cells and myofibroblasts, INSR+ endothelial cells, and C1QA+ macrophages in the HNSCC single-cell dataset;

[0041] Figure 3 A-B are univariate Cox regression analysis and correlation analysis of 17 ligand-receptor genes in the merge-HNSCC cohort;

[0042] Figure 3 C is a chart showing two clusters determined according to receptor-ligand gene expression using ConsensusClusterPlus in the merge-HNSCC cohort;

[0043] Figure 3 D is a chart showing the difference in overall survival rate between two clusters in the merge-HNSCC cohort by Kaplan-Meier survival analysis;

[0044] Figure 4 A is a flowchart showing the construction of the LRS model;

[0045] Figure 5A-D are graphs showing the comparison of C-index of the LRS model with different published HNSCC prediction models in different datasets;

[0046] Figure 6 A-C are graphs showing the Kaplan-Meier analysis of survival of the LRS model in TCGA-HNSCC, merge-GEO, and ICGC-OSCC;

[0047] Figure 6 D is a pie chart showing the clinical trait differences between the high and low LRS score groups determined by Chi-square test;

[0048] Figure 6 E is a graph showing the univariate and multivariate Cox regression analysis of LRS score with patient HPV status, age, gender, tumor stage, and radiation therapy in the TCGA cohort;

[0049] Figure 7 A is a graph showing the expression levels of 8 signature genes in LRS verified by qRT-PCR in clinically collected tumor and adjacent normal tissues from the affiliated XX Medical Center of XX University;

[0050] Figure 7 B is a graph showing the Kaplan-Meier analysis of survival of the LRS model in the in-house collection cohort of XX Medical Center of XX University;

[0051] Figure 8 A is a heatmap showing the correlation between LRS score and representative immune activity;

[0052] Figure 8 B is a graph showing the correlation analysis between LRS score and cancer-related pathways in the TCGA-HNSCC and merge-GEO cohorts;

[0053] Figure 8 C-E are violin plots of the LRS subtypes and immune therapy response TIDE scores in the TCGA-HNSCC, merge-GEO, and ICGC-OSCC cohorts;

[0054] Figure 9 A is a graph showing the cell expression distribution patterns of 8 signature genes of the LRS model in the HNSCC single-cell dataset;

[0055] Figure 9 B is a graph showing the proportion of microenvironment cell infiltration after grouping by high and low LRS scores in the HNSCC single-cell dataset;

[0056] Figure 10A-B is a figure comparing the tumor microenvironment of different LRS subtypes by multiplex immunofluorescence staining in the cohort of patients collected within the XX Medical Center of XX University. Detailed implementation methods

[0057] The following packages involved in the following examples, such as the affy package, input package, limma package, WGCNA package, sva package, rms package, survcomp package, survival package, randomForestSRC package, and Scanpy package, are all prior arts and are sourced from

[0058] http: / / www.bioconductor.org;

[0059] or https: / / scanpy-tutorials.readthedocs.io / en / latest / pbmc3k.html etc. After loading, they are run in R and Python software.

[0060] Example 1: Collection of samples from the internal cohort of XX Medical Center of XX University and data acquisition

[0061] Samples from patients with head and neck squamous cell carcinoma (HNSCC), pre-cancer patients (including leukoplakia, lichen planus with or without dysplasia), and tumor-matched adjacent normal tissues were collected with the approval of the Ethics Committee of XX Medical Center of XX University. All patients who donated tissue samples provided written informed consent.

[0062] Example 2: Download and processing of HNSCC-related bulk data

[0063] Raw transcriptome microarray data and clinical data GSE65858 (n = 270), GSE41613 (n = 97), GSE42743 (n = 103), and ICGC-OSCC (n = 40) of HNSCC were downloaded from the Gene Expression Omnibus database (http: / / www.ncbi.nlm.nih.gov / geo / ). Transcriptome FPKM sequencing data and clinical data TCGA-HNSCC (n = 553) of HNSCC were downloaded from The Cancer Genome Atlas (TCGA).

[0064] The transcriptome FPKM sequencing data were processed using the affy package and input R package, and differential gene expression analysis was performed using the limma package. The sva R package was used to combine the GSE65858, GSE41613, and GSE42743 datasets to eliminate batch effects, and the integrated cohort was named the merge-GEO cohort. In addition, the integrated TCGA-HNSCC dataset was named the merge-HNSCC cohort.

[0065] Example 3: Downloading and Processing HNSCC-Related Single-Cell Data

[0066] An integrated analysis was performed on four publicly available single-cell RNA sequencing (scRNA-seq) datasets of HNSCC. The original count matrix of each dataset was processed using Scanpy (V1.9.1), followed by quality control analysis. In terms of gene filtering, we removed genes that were expressed in fewer than 50 cells or genes related to noise and dissociation artifacts. In terms of cell filtering, we adopted the following criteria:

[0067] (1) The number of genes expressed in each cell was between 500 and 6000;

[0068] (2) The mitochondrial RNA content was less than 20%;

[0069] (3) The total count of each cell did not exceed 40000.

[0070] Using DoubletDetection with default settings, potential doublet cells were identified and removed. After these steps, a total of 348668 high-quality cells were retained for subsequent analysis.

[0071] Subsequently, we normalized the gene expression matrix by normalizing the total UMI count of each cell to 10000 and performing a logarithmic transformation. We used the top 3000 variable genes for principal component analysis (PCA) and corrected the batch effects between different samples and datasets through the harmony algorithm (sigma = 0.07, lambda = 1.0, theta = 0.5). The neighborhood graph was constructed using the top 50 principal components (PCs), and a uniform manifold approximation and projection (UMAP) visualization graph was generated. In the first clustering, we identified the main cell types, including T cells, B cells, NK cells, plasma cells, monocytes, macrophages, dendritic cells (DCs), mast cells, fibroblasts, smooth muscle cells (SMCs), epithelial cells, salivary gland epithelium (SG Epithelium), endothelial cells, and lymphatic endothelial cells (LECs), through Louvain clustering based on classical markers (resolution = 0.3). In the second round of clustering, we further characterized the subpopulations in each cell type based on differentially expressed genes.

[0072] Example 4: Detection and Data Processing of HNSCC Spatial Metabolomics

[0073] 4.1 Detection of Metabolites in HNSCC Spatial Metabolomics

[0074] Embedded samples were stored at -80 °C before sectioning. The samples were sectioned into serial sagittal sections of approximately 10 μm using a cryomicrotome and thaw-mounted on positively charged desorption plates. The sections were stored at -80 °C before further analysis. Before mass spectrometry imaging analysis, the sections were dried at -20 °C for 1 hour and then at room temperature for 2 hours. Meanwhile, adjacent sections were left for hematoxylin-eosin (H&E) staining.

[0075] Briefly, the experiment was performed using the AFADESI-MSI platform (Beijing Victor Technology Co., Ltd., Beijing, China) and a Q-Orbitrap mass spectrometer. The solvent formulation in negative ion mode was acetonitrile (ACN) / H2O (8:2), and in positive ion mode was ACN / H2O (8:2). The solvent flow rate was 1.5 μL / min, the transfer gas flow rate was 45 L / min, the spray voltage was 7 kV, the distance between the sample surface and the sprayer was 3 mm, and the distance between the sprayer and the ion transfer tube was also 3 mm. The mass spectrometry resolution was 20,000, the mass range was 70 - 1200 Da, the automatic gain control target value was 2E6, the maximum injection time was 200 ms, the S-lens voltage was 55 V, and the capillary temperature was 350 °C. The MSI experiment continuously scanned the surface of the sample section in the x-direction at a constant rate of 0.2 mm / s and scanned in the y-direction with a vertical step size of 40 μm.

[0076] 4.2 HNSCC Spatial Metabolome Metabolite Data Processing

[0077] First, the raw data (.raw files) were converted to the imzML format using imzMLConverter and then imported into MSiReader (an open MSI software based on the Matlab platform) for ion image reconstruction. Background subtraction was performed using the Cardinal software package. All mass spectrometry (MS) images were normalized by the total ion count (TIC) per pixel. By aligning high-spatial-resolution H&E images, MS spectra of specific regions were extracted. Ions detected by AFADESI were annotated by the pySM pipeline and the internal SmetDB database. We obtained spatial metabolomics data (defined by SM1 - SM4), plotted the spatio-temporal distribution maps of metabolites, and performed UMAP clustering analysis to classify these metabolites.

[0078] 4.3 Construction of Weighted Gene Co-expression Network (WGCNA)

[0079] The WGCNAR software package was used to construct a data matrix of metabolite expression in different pathological regions of HNSCC spatial metabolomics. Metabolites with the top 25% variance were selected as the input dataset for subsequent WGCNA. Then, hierarchical clustering was used to remove abnormal samples, and a scale-free network was constructed. The dissimilarity of metabolites was calculated, and metabolites with similar expression profiles were modularized according to the dissimilarity. The correlation between the modules and the clinicopathological types of HNSCC was calculated.

[0080] 4.4 WGCNA identified modules related to tumor progression and further screened core metabolites.

[0081] To identify genes related to HNSCC progression, a co-expression network was constructed in HNSCC spatial metabolomics data using WGCNA. A soft threshold of β = 9 was selected to construct a scale-free network graph, and finally, 6 metabolite co-expression modules were constructed. By calculating the correlation between the gene co-expression modules and the pathological feature classification, it was found that the module-trait correlation heatmap showed that the MEblue module was highly correlated with the pathological evolution of HSNCC. KEGG enrichment analysis showed that the metabolites in the MEblue module were significantly enriched in the Purine metabolism pathway ( Figure 1 A), including adenosine, and the metabolite signal gradually increased with tumor evolution ( Figure 1 B-C), suggesting that Purine metabolism mediates the malignant transformation of HSNCC.

[0082] Example 5: NT5E is abnormally expressed in tumors, and its high expression is related to the prognosis of patients

[0083] We selected the upstream regulatory genes of metabolic enzymes according to purine metabolites in the blue module and the KEGG database. Using TCGA data for differential expression analysis, 23 genes that were abnormally expressed in tumor samples were found ( Figure 2 A). Then, we compared the prognostic value of the 23 genes in the TCGA-HNSCC and merge-GEO datasets. Univariate Cox regression analysis showed that only the change in NT5E expression significantly affected the prognosis of patients ( Figure 2 B).

[0084] AB680 is a commonly used NT5E inhibitor. Due to its long half-life and good tolerance, it has shown efficacy in inhibiting the tumor progression of pancreatic cancer and melanoma. We treated HNSCC cell lines with AB680 and performed RNA-seq high-throughput sequencing. The analysis showed that 281 genes were downregulated after AB680 treatment. Enrichment analysis showed that these differentially expressed genes (DEGs) were significantly enriched in processes such as extracellular matrix organization, cytokine-cytokine receptor interaction, regulation of chemokine production, and regulation of cell-cell adhesion ( Figure 2 C). To explore the effect of NT5E on the tumor cell TME, we designed a research strategy to cross the above DEGs with ligand-receptor genes in the CellPhoneDB database and perform cell communication analysis using the scRNA-seq dataset. The results showed that NT5E+ epithelial cells had strong interactions with myofibroblasts, INSR+ endothelial cells, and C1QA+ macrophages. Ligand-receptor pairs including FN1->NT5E, TNC->ITGA5, and SERPINE1->PLAUR were significantly enriched, indicating that NT5E may play a role in recruiting or activating immunosuppressive cells ( Figure 2 D). In summary, it is suggested that NT5E is a potential key molecule connecting purine metabolism and the formation of immunosuppressive TME in HNSCC.

[0085] Example 6: Ligand-receptor pair gene-related HNSCC risk molecular typing

[0086] 6.1 Selection of ligand-receptor pair gene-related HNSCC typing

[0087] Select 52 ligand-receptor pair genes enriched in Example 5 for further analysis. To evaluate the impact on patient prognosis, univariate Cox analysis was performed to determine 17 prognosis-related genes ( Figure 3 A-B). The ConsensuClusterPlus R software package was used to perform unsupervised clustering analysis on the merge-HNSCC database in Example 2. According to the analysis results, k = 2 was selected as the best subtype grouping, that is, HNSCC cases in the combined dataset were divided into C1 and C2 subtypes ( Figure 3 C).

[0088] 6.2 Identification of the characteristics of ligand-receptor pair gene-related subtypes

[0089] The survival R software package was used to perform prognosis analysis on the two subtypes and found that the survival prognosis outcome of the 2 subtypes was better than that of the 1 subtype ( Figure 3 D).

[0090] Example 7: Construction of the HNSCC prognosis risk model LRS

[0091] 7.1 Selection of differentially expressed genes

[0092] The "limma" R package was used to perform differential gene analysis on the two ligand-receptor pair-related subtypes (C1 and C2 subtypes) created in Example 6. With Log2FC < -0.585 or Log2FC > 0.585 (FDR < 0.05) as the threshold, a total of 524 differentially expressed genes were obtained. The 524 DEGs were screened and interacted through the Xboost and Lasso algorithms to obtain 65 feature selection genes, and further combined with survival information for Cox regression analysis to obtain 22 genes( Figure 4 A).

[0093] 7.2 Establishment of a prognostic risk model LRS by machine learning

[0094] Develop a consensus ligand-receptor-based signature (LRS) feature model with high accuracy and stability performance, integrating 10 machine learning algorithms, including random survival forest (RSF), elastic net (Enet), Lasso, ridge regression (Ridge), stepwise Cox regression (stepwise Cox), CoxBoost, Cox partial least squares regression (plsRcox), supervised principal component analysis (SuperPC), generalized boosted regression model (GBM), and survival support vector machine (survival-SVM). Among them, algorithms such as Lasso, stepwise Cox, CoxBoost, and RSF have the ability of feature selection. Therefore, we combined these algorithms to generate a consensus model. A total of 101 algorithm combinations were performed to fit the prediction model in the LOOCV (leave-one-out cross-validation) framework. The initial feature discovery was carried out in the TCGA-HNSCC dataset, and the RSF model was implemented through the randomForestSRC package. RSF has two parameters, ntree and mtry, where ntree represents the number of trees in the forest and mtry is the number of variables randomly selected for each node split. We used the LOOCV framework to perform a grid search on ntree and mtry, all combinations of (ntree, mtry) were formed, and the combination with the best C-index value was selected as the optimized parameter. The regression method considered censored data when formulating the inequality constraints for the support vector problem. The feature generation procedure is as follows:

[0095] (a) Univariate Cox regression to determine prognostic mRNAs in the TCGA-HNSCC cohort

[0096] (b) Then, 101 algorithm combinations were performed on the prognostic mRNAs to fit a prediction model based on the leave-one-out cross-validation (LOOCV) framework in the TCGA-HNSCC cohort;

[0097] (c) All models were detected in the merge-GEO dataset;

[0098] (d) For each model, the Harrell's concordance index (C-index) of all validation datasets was calculated, and the model with the highest average C-index was considered the best ( Figure 4 A).

[0099] In addition, we also compared the performance of LRS with other published HNSCC prognostic prediction models. Notably, LRS outperformed other models in almost every dataset (TCGA-HNSCC, GEO-meta, ICGC, GSE41613, GSE42743, and Cohort-meta) ( Figure 5 A-D).

[0100] The R code for the random survival forest is as follows:

[0101] Load the necessary packages

[0102] library(randomForestSRC)

[0103] library(survival)

[0104] library(dplyr)

[0105] library(tibble)

[0106] Set the random seed

[0107] seed <- 123

[0108] set.seed(seed)

[0109] Parameter settings

[0110] rf_nodesize <- 5

[0111] Initial training of the full-variable random survival forest

[0112]

[0113]

[0114] Extract variable importance and select the top 10 important variables

[0115] var_importance = fit_full$importance

[0116] selected_vars = names(sort(var_importance, decreasing = TRUE))[1:10]

[0117] Retrain the model using the filtered variables

[0118] formula_selected = as.formula(paste("Surv(OS.time,OS) ~ ", paste(selected_vars, collapse = "+")))

[0119] Determine the optimal number of trees using the filtered variables

[0120] set.seed(seed)

[0121] fit_selected = rfsrc(formula_selected,

[0122] data = est_dd,

[0123] ntree = 1000,

[0124] nodesize = rf_nodesize,

[0125] splitrule = 'logrank',

[0126] importance = TRUE,

[0127] proximity = TRUE,

[0128] forest = TRUE,

[0129] seed = seed)

[0130] Determine the optimal number of trees based on the error

[0131] best_selected = which.min(fit_selected$err.rate)

[0132] Reconstruct the final model using the optimal number of trees

[0133] set.seed(seed)

[0134] final_model = rfsrc(formula_selected,

[0135] data = est_dd,

[0136] ntree = best_selected,

[0137] nodesize = rf_nodesize,

[0138] splitrule = 'logrank',

[0139] importance = TRUE,

[0140] proximity = TRUE,

[0141] forest = TRUE,

[0142] seed = seed)

[0143] Validation set predicted risk score (RS)

[0144] rs = lapply(val_dd_list, function(x) {

[0145] pred = predict(final_model, newdata = x)

[0146] data.frame(x[, c("OS.time", "OS")], RS = pred$predicted)})

[0147] Calculate the C-index of the validation set

[0148] cc = data.frame(

[0149] Cindex = sapply(rs, function(x) {

[0150] summary(coxph(Surv(OS.time, OS) ~ RS, data = x))$concordance[1]})) %>%

[0151] rownames_to_column('ID')

[0152] cc$Model = 'RSF'

[0153] Summarize the results into result (note that result should be initialized first)

[0154] result = rbind(result, cc)

[0155] 7.3 Clinical applications of the LRS model

[0156] In the TCGA-HNSCC, merge-GEO, and ICGC-OSCC datasets, Kaplan-Meier survival analysis showed that patients in the low group had better survival advantages than those in the high group ( Figure 6 A-C). In addition, we found that the LRS score was significantly associated with survival status, HPV Status, Smoking Status, Tumor Stage, and Radiation modality in the TCGA dataset ( Figure 6 D). In the univariate Cox regression analysis of the TCGA data, the expression of LRS was observed to be statistically significant. In addition, LRS was considered an independent prognostic biomarker in the multivariate Cox proportional hazards regression model (for OS, HR = 13.25, 95% confidence interval (CI) = 8.87 to 19.77, P < 0.001; Figure 6 E).

[0157] Importantly, we collected a sample cohort from our hospital, performed qRT-PCR on the samples, and found that there were differences in the expression of 8 characteristic genes in different sample types in our hospital cohort ( Figure 7 A). Using the RSF method to construct a model based on the qRT-PCR test data of the samples in our hospital cohort, we found that a high LRS score in our hospital cohort was associated with poor prognosis ( Figure 7 B).

[0158] Example 8: Relationship between the LRS model and tumor microenvironment remodeling and immunotherapy response

[0159] An association analysis of LRS and classical gene signature set scores was performed between the merge-GEO and TCGA-HNSCC datasets, and a significant negative correlation was observed between the LRS score and immune-related marker pathways, including Co-stimulation T cell, TLS, TCR, BCR, NK cell, Inflammatory signature, and Antigen processing and presentation pathways, etc. In contrast, a significant positive correlation was found between various oncogenic signaling pathways and LRS, such as Hypoxia, Angiogenesis, CAF, WNT signaling, and Pro-tumor cytokines, etc. In addition, we found that the LRS score was mostly positively correlated with chemokines and matrix remodeling ( Figure 8 A-B).

[0160] TIDE predicts the immune escape ability of tumors by comprehensively evaluating the activities of these two mechanisms. Higher TIDE scores are associated with poorer immune checkpoint inhibition therapy. We also evaluated the immunotherapy responses of patients in the TCGA, GEO-meta, and ICGC cohorts and found that patients in the high LRS group had higher TIDE scores, indicating poor efficacy of immune checkpoint inhibitor therapy( Figure 8 C-E). The above results indicate that patients in the low-risk group benefit more from immunotherapy than those in the high-risk group.

[0161] We visualized the expression patterns of the eight signature genes in the LRS model in scRNA-seq data and found that their expression was heterogeneous across different cell clusters( Figure 9 A). When defining patient categories according to the LRS score, it is worth noting that the proportion of immune cells (such as T cells, B cells, and NK cells) was higher in the low LRS group, while the proportion of fibroblasts, endothelial cells, and epithelial cells was lower( Figure 9 B).

[0162] In addition, the results of multiplex immunofluorescence staining of tissues from non-treated patients in the internal collection cohort of XX Medical Center of XX University also confirmed that (antibodies: CD56 for NK cells, FAP for myofibroblasts, CD8 for lymphocytes, and PANCK for tumor epithelial cells) patients in the high LRS group had more myofibroblast infiltration, while the infiltration of immune cells CD8+T and NK cells decreased( Figure 10 A-B).

[0163] In summary, samples in the high LRS group have the characteristics of cold tumors, while samples in the low LRS group show a relatively activated immune microenvironment (hot tumors), suggesting that these patients have a good response rate to immunotherapy (such as PD-1 monoclonal antibody therapy).

[0164] Therefore, in this application, large-scale metabolomics and transcriptomics data are processed by random survival forests to automatically extract features closely related to cancer development, and then a precise prediction model is constructed. Lasso, ridge regression, and stepwise regression are used for feature selection to screen out the most diagnostically valuable features from a large number of metabolites and genes. Principal component analysis (PCA) converts high-dimensional data into low-dimensional data for subsequent analysis and visualization. Through these methods, we can effectively reduce noise and improve the performance of the model. To evaluate the generalization ability and stability of the model, we adopted cross-validation techniques such as K-fold cross-validation and leave-one-out cross-validation (LOOCV). These methods can effectively avoid overfitting problems and ensure that the constructed diagnostic model performs consistently on different datasets. By evaluating indicators such as C-index, area under the curve (AUC), accuracy, sensitivity, and specificity, the prediction performance of the model can be comprehensively measured.

Claims

1. A prognostic and immunotherapy response applicability prediction marker for head and neck squamous cell carcinoma, characterized in that: The biomarker is an mRNA biomarker, and the mRNA biomarker includes DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1.

2. Use of the predictive marker for predicting the prognosis of head and neck squamous cell carcinoma and the applicability of immunotherapy response according to claim 1, characterized in that: The applications include evaluating or predicting prognostic risk, predicting the applicability of immunotherapy, predicting survival rate, formulating a treatment / drug regimen, constructing a model for predicting the prognostic risk of head and neck squamous cell carcinoma, constructing a model for the applicability of immunotherapy, constructing a model for predicting the survival rate of head and neck squamous cell carcinoma, preparing a detection reagent or device for predicting the prognostic risk of head and neck squamous cell carcinoma, and preparing a detection reagent or device for predicting the survival rate of head and neck squamous cell carcinoma, or any combination of several of them.

3. A prognostic and immunotherapy response applicability prediction model for head and neck squamous cell carcinoma, characterized in that: Comprising the predictive biomarker for the prognosis of head and neck squamous cell carcinoma and the applicability prediction of immunotherapy response as claimed in claim 1, the method steps for constructing a predictive model for the prognosis of head and neck squamous cell carcinoma and the applicability of immunotherapy response are as follows: S1: Obtain a gene expression dataset: Obtain head and neck squamous cell carcinoma samples and the mRNA expression data of each sample. S2: Define the survival data of the patient as survival time and survival status, where death is defined as 1 and survival is defined as 0. S3: The random survival forest algorithm is adopted, which consists of N tree decision trees. Each tree is constructed according to the following rules: S31: Randomly select a target percentage of samples from the total samples collected in step S1 as bootstrap samples, and the remaining samples as out-of-bag samples for validating the model performance. S32: Randomly select some gene expression variables at each node, use the Log-rank statistic as the criterion to automatically select the best split point, split the samples within the node to maximize the difference in survival curves between nodes, and recursively partition the data until the number of samples in the node reaches the preset minimum value. S4: The cumulative risk function constructed for each tree is expressed as follows: During the decision tree splitting process of each node, according to the Log-rank statistic, the best split point is selected to divide the samples into subgroups with different risk levels, and the samples are gradually recursively partitioned until the requirement of the minimum number of samples in the node is met. The occurrence frequency of survival events within the node is used to estimate the node-specific risk, and finally, it is summarized into the overall cumulative risk function of the tree. h b (DKK1 i ,STC2 i ,INHBA i ,AMIGO2 i ,VEGFC i ,SPOCK1 i ,MT1E i ,FOXD1 i ) S5: For the overall risk score RS of patient i, the calculation method is the average of the cumulative risk functions of all trees. Where: RS i represents the predicted risk score of the i-th patient; N tree represents the total number of decision trees in the random survival forest; h b () represents that the cumulative risk function calculated by the b-th tree is automatically constructed by the algorithm based on gene expression levels; S6: Evaluate the prediction accuracy and generalization ability of the model through the consistency index C-index statistical index.

4. The prognostic and immunotherapy response applicability prediction model for head and neck squamous cell carcinoma according to claim 3, characterized in that: In step S3, N tree is 1000. In step S31, the self-help sample target percentage is 63% and the out-of-bag sample percentage is 37%. In step S32, the preset minimum value of the number of node samples is 5.

5. A prognostic and immunotherapy response applicability prediction system for head and neck squamous cell carcinoma, characterized in that: The system includes the prognostic and immunotherapy response applicability prediction model for head and neck squamous cell carcinoma described in claim 3. The system further includes a data input module, a model calculation module, and a result output module. The data input module is used to input the expression value results of DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1 into the model calculation module. The model calculation module outputs an RS value according to the input values, and defines the patient prediction risk score LRS = RS. The result output module is used to predict the prognostic risk and immunotherapy applicability according to the value of the patient prediction risk score LRS. i value, and defines the patient prediction risk score LRS = RS i . The result output module is used to predict the prognostic risk and immunotherapy applicability according to the value of the patient prediction risk score LRS.

6. The prognostic and immunotherapy response applicability prediction system for head and neck squamous cell carcinoma according to claim 5, characterized in that: If the predicted risk score LRS of the patient > 53.89, it is determined as a high-risk prognosis for head and neck squamous cell carcinoma, indicating a poor prognosis and poor immunotherapy response of the patient. If the predicted risk score of the patient ≤ 53.89, it is determined as a low-risk prognosis for head and neck squamous cell carcinoma, indicating a good prognosis and good immunotherapy response of the patient.

7. A prognostic and immunotherapy response applicability prediction kit for head and neck squamous cell carcinoma, characterized in that: The kit contains the model for predicting the prognostic risk and immunotherapy applicability of head and neck squamous cell carcinoma as claimed in claim 3.

8. The kit according to claim 7, characterized in that, A method for predicting the immunotherapy risk using the kit, the method steps are as follows: A1: Detect the expression levels of DKK1, STC2, INHBA, AMIGO2, VEGFC, SPOCK1, MT1E, and FOXD1 in the samples of head and neck squamous cell carcinoma patients, and perform normalization processing. A2: Substitute the gene expression levels obtained through standardization in step S1 into the model for predicting the prognosis risk and immunotherapy applicability of head and neck squamous cell carcinoma, and define the patient's predicted risk score LRS = RS i , and calculate the risk score LRS; A3: Determine whether the risk score LRS is greater than 53.

89. If the determination result is yes, execute step A4; if the determination result is no, execute step A5; A4: This patient with head and neck squamous cell carcinoma belongs to the high-risk group, indicating that this patient with head and neck squamous cell carcinoma is not suitable for immunotherapy; A5: This patient with head and neck squamous cell carcinoma belongs to the low-risk group, indicating that this patient with head and neck squamous cell carcinoma is suitable for immunotherapy.

9. The kit according to claim 8, wherein: The sample includes tissue and body fluid.

10. The kit according to claim 9, characterized in that: The body fluid is blood.