Marker gene for predicting esophageal squamous cell carcinoma immunotherapy efficacy and application thereof
Patent Information
- Application Number
- CN202310061322.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-20
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2043-01-20
AI Technical Summary
可见,当前的各种免疫治疗疗效预测的生物标志物在食管鳞癌不同免疫治疗方案中的预测作用仍存在不确定性,仍然需要寻找更加精准的预测生物标志物
[0030]本发明提供了预测食管鳞癌免疫治疗疗效的7个标记基因以及基于这7个标记基因的风险评分模型,通过ROC曲线验证了所述风险评分模型具有较高的准确性和特异性,能将免疫治疗高风险和免疫治疗低风险的患者显著分开,可用于预测食管鳞癌免疫治疗疗效。
Smart Images

Figure SMS_1 
Figure SMS_6 
Figure HDA0004061269270000011
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics, specifically relating to marker genes for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma and their applications. Background Technology
[0002] Currently, the treatment strategy for esophageal squamous cell carcinoma remains primarily surgery, supplemented by radiotherapy and chemotherapy. Early-stage esophageal squamous cell carcinoma can be removed endoscopically; however, most cases are not detected early and are often diagnosed at a locally advanced or even late stage by the time symptoms appear. The prognosis for patients with locally advanced esophageal squamous cell carcinoma undergoing surgery alone is poor, with a 5-year survival rate of less than 25%. Therefore, perioperative treatment of esophageal squamous cell carcinoma has become a research hotspot in recent years.
[0003] Immunotherapy, by improving the immunosuppressed state of the tumor microenvironment (TME), activates the patient's own immune system to kill tumor cells. Due to its fewer side effects and longer-lasting efficacy, it has become a hot research topic. Immune checkpoints are a series of molecules expressed on immune cells that regulate the degree of immune activation and prevent excessive activation of the immune system. Multiple clinical trials, including KEYNOTE-181, ESCORT, and KEYNOTE590, have confirmed the good efficacy and safety of immune checkpoint inhibitors (ICIs) for advanced esophageal squamous cell carcinoma, bringing new hope to patients with intermediate to advanced esophageal squamous cell carcinoma.
[0004] Currently, immunotherapy, primarily based on ICIs, has progressed from a later-line treatment for advanced esophageal squamous cell carcinoma to a second-line and first-line treatment. Treatment regimens have also expanded from monotherapy to combination therapy and neoadjuvant therapy. Although immune checkpoint inhibitor therapy has proven effective in various malignancies, its efficacy varies greatly among different tumor types, and even among different patients with the same tumor type. Tumor heterogeneity is one of the obstacles that must be overcome in tumor immunotherapy. The treatment effect of ICIs is influenced by tumor immunogenicity, the tumor microenvironment, and patient physical factors, with more than half of patients not benefiting from the treatment. Accurately identifying patients who may benefit from ICIs or those who show significant benefit from immunotherapy is currently the most pressing and challenging issue in clinical practice. Accurate predictive biomarkers are key to identifying potential beneficiaries of ICI treatment.
[0005] Programmed cell death protein ligand-1 (PD-L1) and tumor mutational burden (TMB) are common biomarkers for predicting the benefit of immunotherapy (ICIs), but neither can accurately predict the efficacy of ICIs in patients with esophageal squamous cell carcinoma. Furthermore, current immunotherapy predictive biomarkers under investigation can be broadly categorized into two types: one related to the tumor immune microenvironment, such as tumor-infiltrating lymphocytes, tertiary lymphoid structures, and T-cell inflammatory gene expression profiles; the other related to tumor cell molecular characteristics, such as high microsatellite instability / mismatch repair deficiency, tumor mutational burden, and neoantigen burden. Other biomarkers include serum non-coding RNA, DNA methylation, and gut microbiota. The predictive efficacy of these immunotherapy biomarkers in esophageal squamous cell carcinoma still requires further research data to confirm. Therefore, the predictive role of current biomarkers for various immunotherapy efficacy predictions in different immunotherapy regimens for esophageal squamous cell carcinoma remains uncertain, and more precise predictive biomarkers are still needed. Summary of the Invention
[0006] In view of this, the purpose of this invention is to provide marker genes for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma and their applications. A risk scoring model can be established using seven marker genes to predict the efficacy of immunotherapy for esophageal squamous cell carcinoma. Compared with existing prediction methods, this invention achieves stable and better predictive results using fewer genes, which is beneficial for screening patients who may benefit from immunotherapy, thus enabling patients to benefit more.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides marker genes for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma, including: CDC25B gene, INHBB gene, ARG1 gene, CTSW gene, HSPA2 gene, TNFRSF9 gene and LY6E gene.
[0009] In a second aspect, the present invention provides the application of marker genes for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma, the application including any one or a combination of several of the following: predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma, constructing a risk score model for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma, and preparing a kit for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma.
[0010] In a third aspect, the present invention provides a risk scoring model for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma based on the aforementioned marker gene, wherein the risk scoring model is as follows:
[0011] The risk score for each sample is calculated as follows: (0.104) × gene expression value of CDC25B + (0.0203) × gene expression value of INHBB + (0.015) × gene expression value of ARG1 + (-0.0355) × gene expression value of CTSW + (-0.0696) × gene expression value of HSPA2 + (-0.0749) × gene expression value of TNFRSF9 + (-0.101) × gene expression value of LY6E, where the gene expression value is log2(TPM+1) of the corresponding gene.
[0012] In a fourth aspect, the present invention provides a method for constructing the risk scoring model, the method comprising the following steps:
[0013] (1) Obtain gene expression data of all esophageal squamous cell carcinoma tumor tissue samples, wherein the samples belong to the pathological complete remission group or the non-pathological complete remission group;
[0014] (2) The gene expression data are randomly grouped into a training set and a validation set;
[0015] (3) In the training set, the Wilcoxon rank-sum test was used to compare the gene expression level of each gene in the pathological complete remission group and the non-pathological complete remission group. The FoldChange value and p value of each gene were calculated, and differentially expressed genes were screened according to the following criteria: genes with FoldChange ≥ 1.5 and p < 0.05 were defined as differentially expressed genes.
[0016] (4) In the training set, Logistic univariate logistic regression analysis was used to calculate the probability and 95% confidence interval of the association between gene expression and pathological remission results of all differentially expressed genes, and genes with p<0.05 were selected as key genes; Lasso regression analysis was used to select 10-fold cross-validation and leave-one-out cross-validation to further reduce key genes and screen out candidate genes for the model.
[0017] (5) Select the following 7 genes as modeling genes from the candidate genes in the model: CDC25B gene, INHBB gene, ARG1 gene, CTSW gene, HSPA2 gene, TNFRSF9 gene and LY6E gene.
[0018] (6) Based on Lasso regression analysis, the Lasso coefficients of the seven modeling genes were determined, and the following risk scoring model was established to predict the efficacy of immunotherapy in patients with esophageal squamous cell carcinoma:
[0019] The risk score for each sample is calculated as follows: (0.104) × gene expression value of CDC25B + (0.0203) × gene expression value of INHBB + (0.015) × gene expression value of ARG1 + (-0.0355) × gene expression value of CTSW + (-0.0696) × gene expression value of HSPA2 + (-0.0749) × gene expression value of TNFRSF9 + (-0.101) × gene expression value of LY6E, where the gene expression value is log2(TPM+1) of the corresponding gene.
[0020] In some specific examples of the present invention, in step (1), the gene expression data is obtained by the following methods: extracting total RNA from tumor samples of esophageal squamous cell carcinoma patients, quantifying the concentration of RNA and assessing fragment length; performing reverse transcription, complementary DNA synthesis and RNA library preparation on RNA fragments, capturing the RNA library and sequencing it; analyzing the sequencing data, quantifying gene expression, outputting a file containing TPM, and obtaining gene expression data after processing.
[0021] In some specific examples of this invention, the glmnet package of the R language software is used to perform Lasso regression analysis.
[0022] In a fifth aspect, the present invention provides a kit for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma, the kit comprising the aforementioned marker gene and the aforementioned risk score model.
[0023] In a sixth aspect, the present invention provides a method for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma using the aforementioned kit, comprising the following steps:
[0024] (1) Detect the gene expression of 7 marker genes in esophageal squamous cell carcinoma patient samples, including CDC25B gene, INHBB gene, ARG1 gene, CTSW gene, HSPA2 gene, TNFRSF9 gene and LY6E gene;
[0025] (2) Substitute the gene expression values corresponding to the gene expression of each gene in the sample obtained in step (1) into the risk scoring model to calculate the risk score of the sample; compare the risk score of the sample with the optimal threshold of the ROC curve and group the esophageal squamous cell carcinoma patient samples according to the comparison results: when the risk score of the sample is higher than the optimal threshold of the ROC curve, the esophageal squamous cell carcinoma patient belongs to the high-risk group, indicating that the esophageal squamous cell carcinoma patient is not suitable for immunotherapy and is a potential non-benefit population; when the risk score of the sample is lower than the optimal threshold of the ROC curve, the esophageal squamous cell carcinoma patient belongs to the low-risk group, indicating that the esophageal squamous cell carcinoma patient is suitable for immunotherapy and is a potential benefit population.
[0026] In some specific embodiments of the present invention, the sample is a tumor tissue sample.
[0027] In some specific embodiments of the present invention, the optimal threshold for the ROC curve is -1.04.
[0028] In this invention, based on gene expression data from esophageal squamous cell carcinoma tumor tissue samples, the Wilcoxon rank-sum test was used to compare the differences in gene expression levels between the pathological complete remission group and the non-pathological complete remission group. The FoldChange value and p-value for each gene were calculated, and differentially expressed genes were screened according to the criteria of FoldChange ≥ 1.5 and p < 0.05. Logistic univariate logistic regression analysis was used to calculate the probability and 95% confidence interval of the association between gene expression of all differentially expressed genes and pathological remission results, and genes with p < 0.05 were selected as key genes. Lasso regression analysis was used, with 10-fold cross-validation and leave-one-out cross-validation selected to further narrow down the key genes, resulting in candidate model genes. Combining gene function, seven genes were selected as modeling genes from the candidate genes. The Lasso coefficients of these seven modeling genes were determined based on Lasso regression analysis, and a risk scoring model was established to predict the efficacy of immunotherapy in patients with esophageal squamous cell carcinoma. ROC curve analysis verified that the risk scoring model has high accuracy and specificity.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] This invention provides seven marker genes for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma and a risk scoring model based on these seven marker genes. The risk scoring model has been verified by ROC curves to have high accuracy and specificity, and can significantly separate patients with high and low risk of immunotherapy. It can be used to predict the efficacy of immunotherapy for esophageal squamous cell carcinoma.
[0031] Compared with the gene set of eight key features already reported, and compared with the traditional method of judging the risk of immunotherapy based on PD-L1 features, this invention can more accurately predict the pathological remission of esophageal squamous cell carcinoma patients receiving immunotherapy combined with chemotherapy.
[0032] The prediction of the efficacy of immunotherapy for esophageal squamous cell carcinoma in patient samples by this invention is consistent with the immune infiltration status presented by 29 features representing the tumor immune microenvironment: the model predicts that samples in the low-risk group have higher immune infiltration and present a hot tumor state; the model predicts that samples in the high-risk group have lower immune infiltration and present a cold tumor state. This shows that the risk scoring model of this invention provides a reliable method for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma.
[0033] In summary, the risk scoring model of this invention, based on seven marker genes, can achieve stable and good predictive results with fewer genes, which is beneficial for screening patients who may benefit from immunotherapy and enabling patients to benefit more.
[0034] The gene sequences disclosed in this invention are all stored in the Gene Expression Database (http: / / www.ncbi.nlm.nih.gov / geo / ). Here, the seven marker genes and their GeneIDs are listed as follows: CDC25B gene (GeneID: 994), INHBB gene (GeneID: 3625), ARG1 gene (GeneID: 383), CTSW gene (GeneID: 1521), HSPA2 gene (GeneID: 3306), TNFRSF9 gene (GeneID: 3604), and LY6E gene (GeneID: 4061). Attached Figure Description
[0035] Figure 1A A volcano plot showing the comparison of gene expression levels of all genes in the pathological complete remission (pCR) group and the non-pathological complete remission (non-pCR) group of the training set.
[0036] Figure 1B The expression heatmaps of 111 differentially expressed genes in the pathological complete remission (pCR) group and the non-pathological complete remission (non-pCR) group of the training set are presented.
[0037] Figure 2 This is a graph showing the λ distribution of the Lasso regression results.
[0038] Figure 3 A forest plot showing the OR, 95% CI, and p-values of 15 candidate genes for the model is presented.
[0039] Figure 4 Expression heatmaps of seven modeling genes in the pathological complete remission (pCR) group and the non-pathological complete remission (non-pCR) group of the training set are presented.
[0040] Figure 5 The modeling genes (marker genes) and their corresponding Lasso coefficients are shown in the final established risk scoring model.
[0041] Figure 6 The receiver operating characteristic (ROC) curves calculated and plotted on the training set are presented.
[0042] Figure 7 Receiver operating characteristic curves (ROC) calculated and plotted on the validation set are presented.
[0043] Figure 8This is a radar chart plotted by comparing the risk scoring model score of this invention with the GSVA scores of gene sets of eight previously reported key features using ROC.
[0044] Figure 9 This is a stacked distribution diagram of the risk score grouping of the validation set samples calculated using the risk score model of the present invention and the pathological remission results of clinical testing.
[0045] Figure 10A This is a stacked distribution diagram of the risk score grouping and pathological remission results of 31 samples calculated using the risk score model of the present invention.
[0046] Figure 10B This is a stacking distribution diagram of CPS scores and clinically detected pathological remission results for 31 samples stained with PD-L1 22C3.
[0047] Figure 11 Information on the gene set with 29 features.
[0048] Figure 12A A heatmap plotted based on the GSVA scores corresponding to the gene set of 29 features in the training set samples.
[0049] Figure 12B A heatmap plotted based on the GSVA scores corresponding to the gene set of 29 features in the validation set samples. Detailed Implementation
[0050] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and specific examples. It should be understood that these embodiments are for illustrative purposes only and should not be construed as limiting the scope of the invention.
[0051] Unless otherwise specified in the following embodiments, the techniques or conditions described in the literature in this field or in accordance with the product manual shall apply.
[0052] The reagents or instruments used, unless otherwise specified by the manufacturer, are all conventional products that can be purchased on the market.
[0053] The R language software used in the following examples is the existing R software (version 4.1.2), which is available from https: / / www.r-project.org / .
[0054] Patient's condition:
[0055] We collected 114 formalin-fixed paraffin-embedded (FFPE) endoscopic biopsy tumor tissue samples from patients undergoing ESCC treatment at Zhejiang Cancer Hospital between 2020 and 2022. All patients received two cycles of PD-1 monoclonal antibody combined with chemotherapy and underwent radical esophagectomy for squamous cell carcinoma after the two cycles of treatment. All patients signed informed consent forms before specimen collection.
[0056] Based on the residual tumor in the pathological specimens from radical surgery, all patients were divided into a pathologic complete response (pCR) group and a non-pathologic complete response (non-pCR) group. There were 22 samples in the pathologic complete response group and 92 samples in the non-pathologic complete response group. Detailed patient (sample) information is shown in Table 1 below.
[0057] Table 1 Patient (Sample) Information
[0058]
[0059] *p<0.05 is considered statistically significant.
[0060] I. Library Construction and Sequencing
[0061] use Total RNA was extracted from FFPE tumor tissue samples using the FFPE RNA Extraction Kit (Cat.#8.02.24101X036G, AmoyDx, Xiamen, China). RNA concentration was quantified using Qubit (Thermo Fisher Scientific, Waltham, MA, USA). Fragment length was assessed using an Agilent 2100 Bioanalyzer and the RNA HS Kit (Cat.#5067-1511, Agilent). RNA was lysed at 95°C for 0–15 minutes based on the DV200 value estimated by the Agilent 2100 Bioanalyzer system. Then, it was used… Ultra TM II Directional RNA Library Prep Kitfor (Cat.#E7760L, NEB) This involves reverse transcription of RNA fragments, complementary DNA synthesis, and RNA library preparation. The RNA library is prepared by... Captured on a Master Panel (Cat.#8.06.0130, AmoyDx, Xiamen, China) for RNA expression detection and sequenced on an Illumina NovaSeq 6000 (San Diego, CA, USA).
[0062] II. Quantitative Gene Expression
[0063] STAR32 (versions 020, 201) was used to map paired RNA-seq reads and transcriptome annotations (Genecode version 20) to the human genome GRCh37 (hg19). Gene quantification was performed using RSEM 33 (version v1.2.28), and the output file contained gene_id, corresponding transcript_id, count, and TPM information. This data was then processed using R software to obtain gene expression data (used to analyze the expression matrix of differentially expressed genes).
[0064] III. Differential Gene Screening
[0065] Gene expression data from preoperative tumor tissue specimens of 114 patients with esophageal squamous cell carcinoma were used to randomly group them into training and validation sets using the caret package of R software. The training set consisted of 66 samples, and the validation set consisted of 48 samples.
[0066] In the training set, the Wilcoxon rank-sum test (using the wilcox.test function in R software) was used to compare the gene expression levels of each gene in the pathological complete remission (pCR) group and the non-pathological complete remission (non-pCR) group. The fold change and p-value of each gene were calculated, and differentially expressed genes (DEGs) were screened according to the following criteria: genes with fold change ≥ 1.5 and p-value < 0.05 were defined as differentially expressed genes (DEGs).
[0067] The comparison of all gene expression levels between the pathological complete remission (pCR) group and the non-pathological complete remission (non-pCR) group in the training set is presented in a volcano plot, as follows: Figure 1A As shown in the figure. The horizontal axis represents the fold change of genes to base log2, and the vertical axis represents the statistical significance (p-value) of gene differences to base log2. Each scatter point in the figure represents a gene. The two vertical dashed lines represent the threshold of 1.5 fold change, and the horizontal dashed line represents the threshold of p = 0.05. Therefore, the upper left quadrant represents downregulated genes with significant differences, and the upper right quadrant represents upregulated genes with significant differences. Figure 1AIt is evident that, compared to the non-pathological complete remission (non-pCR) group, the pathological complete remission (pCR) group contained both significantly upregulated and significantly downregulated genes. 111 genes were identified as differentially expressed genes (DEGs) meeting the criteria of FoldChange ≥ 1.5 and p-value < 0.05, of which 14 genes were highly expressed in the pCR group and 97 genes were highly expressed in the non-pCR group. Figure 1B The expression heatmaps of 111 differentially expressed genes in the pathological complete remission (pCR) group and the non-pathological complete remission (non-pCR) group of the training set are shown. Figure 1B The darker the red, the higher the level of expression; the darker the blue, the lower the level of expression.
[0068] IV. Key Gene Screening
[0069] In the training set, logistic univariate logistic regression analysis was performed using the rms package in R software to calculate the odds (OR) and 95% confidence interval (CI) of the association between gene expression and pathological remission outcomes for all 111 DEGs, and genes with p < 0.05 were screened as key genes. A total of 67 key genes were identified.
[0070] Pathological remission outcomes include pathological complete remission (pCR) and non-pathological complete remission (non-pCR).
[0071] In the training set, Lasso regression analysis was performed using the glmnet package in R software. 10-fold cross-validation and leave-one-out cross-validation were selected to further reduce the number of feature genes and screen important key genes as candidate genes for the model.
[0072] Figure 2 This is a graph showing the λ distribution of the Lasso regression results. The horizontal axis represents the Mean-squared Error corresponding to different numbers of genes (values on the graph) at different log(λ). The two vertical dashed lines in the graph represent the common optimal λ values lambda.min and lambda.1se, respectively. lambda.min refers to the λ value that yields the minimum mean of the objective parameter among all λ values, while lambda.1se refers to the λ value that yields the simplest model within a variance range of lambda.min. Figure 2 In the process, 18 important key genes were screened for lambda.min, and 15 important key genes were screened for lambda.1se. Finally, the 15 important key genes corresponding to lambda.1se were selected as candidate genes for the model.
[0073] The OR, 95% CI, and p-value of these 15 model candidate genes are as follows: Figure 3 As shown.
[0074] V. Risk Scoring Model Construction
[0075] Based on gene function, the following 7 genes were selected as modeling genes from 15 candidate genes: CDC25B gene, INHBB gene, ARG1 gene, CTSW gene, HSPA2 gene, TNFRSF9 gene, and LY6E gene.
[0076] Figure 4 Expression heatmaps of seven modeling genes in the training set for both pathological complete remission (pCR) and non-pathological complete remission (non-pCR) groups are presented. Figure 4 It is evident that ARG1, INHB, and CDC25B genes were expressed at low levels in the pCR group, while CTSW, HSPA2, TNFRSF9, and LY6E genes were expressed at high levels in the pCR group.
[0077] Lasso regression analysis was performed using the glmnet package in R software to determine the Lasso coefficients of seven modeling genes, and a risk score model was established to predict the efficacy of immunotherapy in patients with esophageal squamous cell carcinoma.
[0078] Seven modeling genes and their Lasso coefficients are as follows: Figure 5 As shown. According to Figure 5 The Lasso coefficients of the seven modeling genes (CDC25B, INHBB, ARG1, CTSW, HSPA2, TNFRSF9, and LY6E) were 0.104, 0.0203, 0.015, -0.0355, -0.0696, -0.0749, and -0.101, respectively. These seven modeling genes are marker genes associated with the efficacy of ESCC immunotherapy combined with chemotherapy.
[0079] The final risk scoring model is as follows:
[0080] The risk score for each sample is calculated as follows: (0.104) × gene expression value of CDC25B + (0.0203) × gene expression value of INHBB + 0.015 × gene expression value of ARG1 + (-0.0355) × gene expression value of CTSW + (-0.0696) × gene expression value of HSPA2 + (-0.0749) × gene expression value of TNFRSF9 + (-0.101) × gene expression value of LY6E. Here, the gene expression value is log2(TPM+1) of the corresponding gene.
[0081] Model evaluation is based on the area under the receiver operating characteristic curve (ROC) (AUC).
[0082] Figure 6The receiver operating characteristic (ROC) curves calculated and plotted using the pROC software package (version 1.18.0) on the training set are presented, and the AUC value of the training set is calculated to be 0.953 based on these curves.
[0083] VI. Performance Evaluation of the Risk Scoring Model (Part 1)
[0084] The established risk scoring model was validated on the validation set data: the receiver operating characteristic curve (ROC) was calculated and plotted using the pROC software package (version 1.18.0), and the area under the curve (AUC) was also calculated. Figure 7 The receiver operating characteristic (ROC) curves calculated and plotted using the pROC software package (version 1.18.0) in the validation set data matrix are presented, and the AUC value of the validation set is calculated to be 0.855 based on these curves.
[0085] VII. Performance Evaluation of the Risk Scoring Model (Part Two)
[0086] The predictive performance of the risk scoring model based on 7 marker genes in this invention was compared with that of a gene set of 8 important features reported previously (gene information in the gene set is shown in Table 2).
[0087] Specifically, the GSVA package in the R language software is used to calculate the scores of gene sets for each important feature on the validation set, and ROC curves are plotted. The area under the curve is then calculated as the AUC value. The AUC values are then compared horizontally with the scores of each important gene set using the risk scoring model of this invention. Figure 8 The radar chart shown is an example of a spider plot, created using the fmsb package in the R programming language.
[0088] Depend on Figure 8 As can be seen, in the training set, the AUC value of the risk scoring model of this invention is significantly higher than the AUC value of the gene set with 8 important features reported previously. This indicates that the risk scoring model of this invention, constructed based on 7 marker genes, has better predictive efficacy for pathological remission outcomes in esophageal squamous cell carcinoma patients receiving immunotherapy than the predictive performance of the gene set with 8 important features reported previously.
[0089] Table 2: Gene set information for 8 important features
[0090]
[0091] VIII. Performance Evaluation of the Risk Scoring Model (Part 3)
[0092] In the risk scoring model of this invention, the magnitude of the risk score represents the probability of risk associated with immunotherapy for patients with esophageal squamous cell carcinoma.
[0093] By detecting the gene expression of marker genes in esophageal squamous cell carcinoma patient samples and inputting the corresponding gene expression values into the aforementioned risk scoring model, the risk score of the patient samples can be calculated. The risk score is then compared with the optimal threshold recommended by the ROC curve of the training set, and the esophageal squamous cell carcinoma patient samples are grouped according to the comparison results.
[0094] When the risk score is higher than the optimal threshold recommended by the ROC in the training set, the patient with esophageal squamous cell carcinoma belongs to the high-risk group, indicating that the patient is not suitable for immunotherapy and is a potential non-benefit population; when the risk score is lower than the optimal threshold recommended by the ROC in the training set, the patient with esophageal squamous cell carcinoma belongs to the low-risk group, indicating that the patient is suitable for immunotherapy and is a potential benefit population.
[0095] Based on this, the efficacy or risk of immunotherapy in patients with esophageal squamous cell carcinoma can be predicted.
[0096] The risk scores of 48 samples in the validation set will be calculated using the risk scoring model of the present invention. Based on the comparison between the risk scores and the optimal threshold (-1.04) recommended by the ROC of the training set, the samples in the validation set will be divided into a high-risk group and a low-risk group.
[0097] The clustering distribution of risk score groups and clinically tested pathological remission results for 48 samples in the validation set calculated by the risk scoring model according to the present invention is shown in the figure below. Figure 9 As shown. Clinically, pathological remission results refer to the clinical classification of pathological samples into pathological complete remission (pCR) and non-pathological complete remission (non-pCR).
[0098] according to Figure 9 The results show that, with a pCR rate of 20.8% in the validation set clinical testing, the risk scoring model of this invention predicted 36 patients in the high-risk group, of whom 4 (11.1%) achieved pCR and 32 achieved non-pCR. The risk scoring model also predicted 12 patients in the low-risk group, of whom 6 (50%) achieved non-pCR and 6 achieved pCR. It is evident that the pCR rate predicted by the risk scoring model of this invention in the low-risk group was significantly higher than the overall pCR rate and the high-risk group pCR rate, while the pCR rate predicted by the risk scoring model of this invention in the high-risk group was significantly lower than the overall pCR rate and the low-risk group pCR rate. This indicates that the risk scoring model of this invention can more accurately predict the pathological remission of esophageal squamous cell carcinoma patients receiving immunotherapy combined with chemotherapy and can more accurately identify potential beneficiaries of immunotherapy.
[0099] IX. Performance Evaluation of the Risk Scoring Model (Part Four)
[0100] Of the 114 samples, 31 underwent PD-L1 immunohistochemical staining, and CPS scores were obtained for these 31 samples. The specific process is as follows:
[0101] Primary antibody PD-L1 (22C3) and EnVision FLEX / HRP secondary antibody kit (K8002) were purchased from DAKO.
[0102] DAKO staining platform: PT Link repair instrument, Link48 histochemical staining instrument.
[0103] Experimental procedure: Dewaxing with xylene for 10 min, retrieval with PDL1-specific retrieval solution (K8005) in PT Link retrieval instrument at 97℃ for 20 min, soaking in 3% H2O2 for 5 min, incubation with primary antibody PD-L1 (22C3) for 40 min, FLEX+Mouse (LINKER) for 30 min, FLEX / HRP for 30 min, DAB for 5 min, DAB for 5 min, Enhancer for 5 min, rinsing with DI water, counterstaining with hematoxylin for 1 min, dehydration with ethanol in stages, clearing with xylene, and mounting with neutral resin.
[0104] Result determination:
[0105] CPS is defined as the percentage of PD-L1-stained live tumor cells (partial or complete membrane staining of any intensity) and PD-L1-stained tumor-associated immune cells (lymphocytes, macrophages) (cell membrane or cytoplasmic staining of any intensity) out of all live tumor cells. The result is expressed as a value from 0 to 100 (when the calculated result exceeds 100, the final result is calculated as 100). The calculation formula is: CPS = (PD-L1 membrane-stained positive tumor cells + PD-L1 membrane-stained positive tumor-associated immune cells (lymphocytes, macrophages)) / Total number of tumor cells * 100
[0106] Interpretation steps:
[0107] ① Examine all well-preserved tumor areas at low magnification. Assess the total area of PD-L1 stained tumor cells. Note: Partial membrane staining or 1+ membrane staining may be difficult to see at low magnification. Ensure there are at least 100 viable tumor cells in the sample;
[0108] ② If there are fewer than 100 surviving tumor cells in the specimen, use a deep section of the same paraffin block or a tissue section of another paraffin block. There may be a sufficient number of tumor cells for PD-L1 expression assessment.
[0109] ③ At a higher magnification (20x), evaluate PD-L1 expression and calculate CPS:
[0110] * Determine the total number of surviving tumor cells with and without PD-L1 staining (CPS denominator);
[0111] * Determine the number of PD-L1 stained cells (tumor cells, lymphocytes, macrophages) (CPS molecules);
[0112] *Calculate CPS;
[0113] Note: Membrane staining assessment should be performed at a magnification not exceeding 20x. Readers of sections should not perform CPS calculations at a magnification of 40x.
[0114] Meanwhile, for these 31 samples, the risk score of each sample is calculated using the risk scoring model according to the present invention, and the samples are divided into a high-risk group and a low-risk group based on the comparison between the risk score and the optimal threshold (-1.04) recommended by the ROC of the training set.
[0115] The clustering distribution of risk score groups for 31 samples calculated according to the risk scoring model of the present invention with the pathological remission results of clinical testing is shown in the figure below. Figure 10A As shown. Figure 10A In the high-risk group, there were 21 people, of whom 95% were clinically assessed as non-pCR and 5% as pCR; in the low-risk group, there were 10 people, of whom 60% were clinically assessed as pCR and 40% as non-pCR.
[0116] A clustering plot of CPS scores grouped by PD-L1 22C3 histochemical staining and clinically tested pathological remission results for 31 samples is shown below. Figure 10B As shown. Figure 10B Among them, there were 4 cases in the PD-L1 expression negative group (CPS<1) and 27 cases in the PD-L1 expression positive group (CPS≥1). In the PD-L1 expression negative group, the proportion of non-pCR was 100%, while in the PD-L1 expression positive group, the proportion of pCR was only 26%, and the proportion of non-pCR was 74%.
[0117] Depend on Figure 10A and Figure 10B As can be seen, compared with the PD-L1 risk score, the risk model of the present invention can more accurately predict the pathological complete remission of esophageal squamous cell carcinoma patients receiving immunotherapy, and can more accurately identify potential beneficiaries of immunotherapy.
[0118] X. Performance Evaluation of the Risk Scoring Model (Part 5)
[0119] Twenty-nine features that influence the tumor immune microenvironment were selected for analysis.
[0120] The 29 signatures include: MHCI, MHCII, co-activation molecules, Effector cells, Effector cell traffic, NK cells, T cells, B cells, M1 signature, Th1 signature, antitumor cytokines, checkpoint molecules, Treg cells, Treg and Th2 traffic, neutrophil signature, granulocyte traffic, Immune Suppression by Myeloid Cells, Myeloid Cell Traffic, Tumor-associated Macrophages, Macrophage and Dendritic Cell Traffic, Th2 signature, and Protumor cytokines. The gene set for these 29 characteristics includes: cytokines, cancer-associated fibroblasts, matrix, matrix remodeling, angiogenesis, endothelium, tumor proliferation rate, and epithelial-mesenchymal transition (EMT) signature. A list of genes for these characteristics is provided below. Figure 11 As shown.
[0121] The GSVA score corresponding to the gene set of 29 features in the training set was calculated using the GSVA package in the R language software, and a heatmap was plotted. Figure 12A As shown.
[0122] The GSVA score corresponding to the gene set of 29 features in the validation set was calculated using the GSVA package in R software, and a heatmap was plotted. Figure 12B As shown.
[0123] Depend on Figure 12A and Figure 12BIt can be seen that in both the training and validation sets, the following results were observed: samples predicted by the model to be in the low-risk group showed higher immune infiltration and exhibited a hot tumor state; while samples predicted by the model to be in the high-risk group showed lower immune infiltration and exhibited a cold tumor state.
[0124] As can be seen, the risk scoring model constructed based on seven marker genes in this invention has been strongly supported and validated by data in the validation set, demonstrating high accuracy and specificity. It can significantly separate patients at high and low risk of immunotherapy and can be used to predict the efficacy of immunotherapy for esophageal squamous cell carcinoma. Compared with the gene sets of eight important features previously reported, and compared with traditional methods for judging immunotherapy efficacy based on PD-L1 features, this invention can more accurately predict the pathological remission of esophageal squamous cell carcinoma patients receiving immunotherapy. Furthermore, the prediction of the efficacy of immunotherapy for esophageal squamous cell carcinoma in this invention is consistent with the immune infiltration status presented by 29 features representing the tumor immune microenvironment, confirming the reliability of the prediction. In summary, the risk scoring model constructed based on the expression of seven genes in this invention can achieve stable and good predictive results using fewer genes, which is beneficial for screening patients who may benefit from immunotherapy, enabling patients to better benefit from the treatment.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A marker gene for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma, characterized in that, The combination of the following 7 genes: CDC25B, INHB, ARG1, CTSW, HSPA2, TNFRSF9, and LY6E.
2. A kit for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma, characterized in that, The kit contains detection reagents for detecting the seven genes in the marker genes as described in claim 1.