A molecular typing method and a marker combination for predicting the response of colorectal cancer patients to immunotherapy
By using transcriptomic analysis of 200 core genes and multi-omics integration methods, a classifier for predicting the response to immunotherapy in colorectal cancer was constructed. This solves the problems of insufficient prediction accuracy and high detection cost in existing technologies, achieving efficient and accurate prediction of immunotherapy response, which is suitable for clinical application.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIN HUA HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
- Filing Date
- 2026-06-25
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, the accuracy of single biomarkers in predicting the response to immunotherapy for colorectal cancer is insufficient. Existing molecular subtyping methods are not specifically designed for immunotherapy prediction. Machine learning models are prone to overfitting in small sample training scenarios and have limited cross-platform generalization ability. Multi-omics integration methods have high detection costs and complex technical processes, making them difficult to promote and apply in routine clinical practice.
Transcriptome analysis of 200 core genes was used, and unsupervised clustering was performed by combining the nearest template prediction (NTP) algorithm and integrated nonnegative matrix factorization (IntNMF). Iterative feature selection was performed by combining support vector machine recursive feature elimination (SVM-RFE) to construct a classifier for predicting the response to immunotherapy in colorectal cancer. Transcriptome expression data were detected by RNA sequencing or gene chip to reduce detection costs and improve prediction accuracy.
It significantly improves prediction accuracy, reduces testing costs, enhances clinical accessibility, exhibits good cross-platform robustness, can identify patients who benefit from immunotherapy, has independent prognostic predictive value, and complements existing molecular subtyping methods well.
Smart Images

Figure CN122493946A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biomedicine and precision medicine, specifically to a molecular subtyping method for predicting the response to immunotherapy in colorectal cancer patients, a combination of biomarkers for implementing the method, a classifier construction method, and their applications. Background Technology
[0002] Colorectal cancer (CRC) is the third leading cause of cancer incidence and the second leading cause of cancer death worldwide. Global cancer statistics for 2024 show that more than 1.9 million new CRC cases and more than 900,000 deaths occur annually. In recent years, immune checkpoint inhibitors (ICIs) have achieved breakthrough efficacy in various solid tumors, especially in patients with microsatellite instability-high (MSI-H) or deficient mismatch repair (dMMR) CRC. Anti-PD-1 / PD-L1 and anti-CTLA-4 antibodies have shown significant clinical benefits, with objective response rates (ORR) reaching 30%-50% (Le et al., New England Journal of Medicine, 2015; Overman et al., The Lancet Oncology, 2017). However, clinical practice has observed significant heterogeneity in immunotherapy response even within the MSI-H / dMMR patient population, with approximately 30%-50% of MSI-H patients developing primary resistance. On the other hand, microsatellite stable (MSS) patients account for 95% of all colorectal cancer patients, and while most are insensitive to immunotherapy monotherapy, some still benefit from it. Currently, there is a lack of effective methods for early prediction and identification, and a molecular subtyping approach capable of predicting immunotherapy efficacy is urgently needed in clinical practice. Therefore, developing biomarkers and predictive models that can accurately predict immunotherapy response in colorectal cancer patients is of significant clinical importance and urgent need for optimizing treatment decisions, avoiding ineffective treatments, and reducing the waste of medical resources.
[0003] Currently, commonly used biomarkers for predicting immunotherapy in clinical practice can be mainly divided into the following categories: First, protein-level biomarkers, represented by the expression level of programmed death-ligand 1 (PD-L1), which are quantitatively assessed using the Combined Positive Score (CPS) obtained through immunohistochemical detection. Second, genomic-level biomarkers, such as tumor mutational burden (TMB) and microsatellite instability (MSI) status. Third, immune cell biomarkers, such as the density of tumor-infiltrating lymphocytes (TILs), tertiary lymphoid structures (TLS), and specific immune subset biomarkers such as CXCL13-positive T cell characteristics. In addition, recent studies have found that some novel biomarkers, such as the dynamic changes in circulating tumor DNA (ctDNA) in the blood and the composition of the gut microbiome, are also associated with immunotherapy response. However, all of the above biomarkers have their own limitations. PD-L1 expression exhibits high heterogeneity both spatially (between primary and metastatic lesions) and temporally (before and after treatment). There is a lack of unified standards and comparability among different detection platforms (22C3, 28-8, SP142, etc.) and different detection antibodies. Furthermore, its predictive efficacy varies significantly across different cancer types and treatment regimens (Herbst et al., Nature, 2014; Rimm et al., JAMA Oncology, 2017). The PD-L1 expression score (CPS score) is poorly predictive in colorectal cancer. There is no consensus on the threshold for TMB as a predictive biomarker (ranging from 10 mut / Mb to 16 mut / Mb). Different sequencing platforms and panel sizes show insufficient consistency in TMB estimation, and its predictive value in CRC is weaker than that in melanoma and non-small cell lung cancer (Yarchoan et al., New England Journal of Medicine, 2017; Goodman et al., Molecular Cancer Therapeutics, 2017). MSI status only covers about 5% of CRC patients, making it impossible to predictively stratify the vast majority of MSS-type patients. Single biomarkers often fail to fully characterize the complex state of the tumor microenvironment (TME), resulting in limited predictive accuracy. Multiple studies have shown that the predictive AUC of single biomarkers in CRC is generally between 0.50 and 0.65, far from meeting the needs of clinical decision-making.
[0004] In recent years, immunotherapy prediction strategies based on multi-omics integrated analysis and molecular subtyping have attracted widespread attention. The core idea of this technical approach is to construct a molecular classification system that can comprehensively reflect the state of the tumor immune microenvironment by integrating multi-dimensional molecular information such as the genome, transcriptome, proteome, and epigenome, thereby achieving accurate prediction of immunotherapy response.
[0005] In transcriptomic molecular subtyping studies, the Consensus Molecular Subtypes (CMS) for colorectal cancer is the most representative work (Guinney et al., Nature Medicine, 2015). The CMS subtype classifies CRC into four subtypes: CMS1 (immune type, approximately 14%), characterized by immune activation, microsatellite instability, and high mutational burden, considered associated with a better response to immunotherapy; CMS2 (epithelial type, approximately 37%), characterized by WNT and MYC pathway activation; CMS3 (metabolic type, approximately 13%), characterized by metabolic dysregulation; and CMS4 (mesenchymal type, approximately 23%), characterized by TGF-β activation, mesenchymal invasion, and angiogenesis. However, the CMS subtype was initially designed based on prognostic analysis of surgically resected samples and was not specifically designed for immunotherapy prediction. Its predictive efficacy is limited, and it cannot effectively identify individuals who benefit from immunotherapy in the majority subtypes such as CMS2 and CMS4. Furthermore, CMS typing requires whole transcriptome testing and a complex classification algorithm (CMScaller), which presents a high barrier to clinical translation.
[0006] Regarding gene biomarker combinations, several studies have attempted to construct multi-gene expression profiles for predicting immunotherapy. Ayers et al. (Journal of Clinical Investigation, 2017) reported a T-cell inflammatory gene expression profile (GEP) based on IFN-γ-related genes (including CXCL9, CXCL10, CXCL11, IDO1, STAT1, etc.), which showed some predictive ability for immunotherapy response in various solid tumors, but its predictive efficacy in CRC has not been fully validated. Jiang et al. (Nature Medicine, 2018) proposed the TIDE (Tumor Immune Dysfunction and Exclusion) algorithm based on tumor microenvironment immune cell scoring, which predicts immunotherapy response by integrating information from two dimensions: immune dysfunction and immune rejection. However, the TIDE algorithm relies on a predefined gene set and prior knowledge of immune cell components, resulting in unstable generalization performance across different datasets. Furthermore, recent studies have explored classifiers based on specific cell subpopulation marker genes. For example, Gong et al. (Discover Oncology, 2026) constructed a CRC immunotherapy response classifier using a random forest algorithm based on 39 common marker genes from three inflammation-related cell subpopulations: CEMIP-positive monocytes, CCL4-positive neutrophils, and MMP3-positive fibroblasts. The classifier achieved an AUC of 0.718 on the IMvigor210 cohort, but its gene set was relatively small and lacked sufficient validation on a CRC-specific immunotherapy cohort. These findings provide valuable references for this invention, but still suffer from insufficient predictive accuracy, need for improved stability, and limited cross-platform generalization capabilities.
[0007] In feature selection and classification algorithms, Support Vector Machine-Recursive Feature Elimination (SVM-RFE), a classic feature selection method, has been widely used in bioinformatics for dimensionality reduction and biomarker screening of gene expression data. It effectively eliminates redundant features and reduces the risk of overfitting (Guyon et al., Machine Learning, 2002). The Nearest Template Prediction (NTP) algorithm, first proposed by Hoshida (PLoS ONE, 2010), is a flexible classification strategy based on single samples. It predicts the category by calculating the correlation between the test sample and a preset template, exhibiting natural robustness to cross-platform technology differences and has been applied in various molecular classification systems such as CMS genotyping. The Integrative Non-negative Matrix Factorization (IntNMF) algorithm can jointly analyze multiple data modalities and extract latent patterns, demonstrating advantages in multi-omics integrated clustering (Chalise and Fridley, PLoS ONE, 2017). However, there are currently no precedents for integrating IntNMF multi-omics clustering, SVM-RFE iterative feature screening, and NTP single-sample classification into a complete technology chain, especially in the specific technical problem of predicting the response to colorectal cancer immunotherapy, where no technical solution has been found that combines the three technologies.
[0008] In summary, the existing technologies have the following technical problems that urgently need to be solved: (1) The prediction accuracy of single biomarkers (PD-L1, TMB, MSI, etc.) is insufficient, with AUC generally between 0.50 and 0.65; (2) Existing molecular typing methods (such as CMS typing) are not specifically designed for immunotherapy prediction, and have blind spots in identifying the population that will benefit from immunotherapy; (3) Existing machine learning-based prediction models are prone to overfitting in small sample training scenarios, and have limited cross-platform generalization ability; (4) Although multi-omics integration methods can theoretically improve prediction accuracy, their detection costs are high and the technical process is complex, making it difficult to promote and apply in routine clinical practice. Therefore, developing a simple, accurate, and clinically accessible method for predicting the response to colorectal cancer immunotherapy has important technical value and clinical significance. Summary of the Invention
[0009] This invention aims to address the following technical problems in existing technologies: First, the predictive accuracy of existing single biomarkers (including PD-L1 CPS score, tumor mutational burden (TMB), microsatellite instability (MSI) status, and CXCL13-positive T cell characteristics) for colorectal cancer immunotherapy response is generally insufficient, with AUC values mostly between 0.50 and 0.65, far from meeting the needs of precise clinical decision-making; Second, existing molecular subtyping methods, such as CMS subtyping, are not specifically designed for immunotherapy prediction, and have blind spots in identifying patients who will benefit from immunotherapy, especially MSS-type patients; Third, existing machine learning-based prediction models are prone to overfitting in small-sample training scenarios, have limited cross-platform generalization ability, and exhibit significant differences in predictive performance between different studies; Fourth, while existing multi-omics-integrated prediction methods can theoretically improve prediction accuracy, their detection costs are high, and the technical processes are complex, requiring systematic detection across multiple dimensions such as the genome, transcriptome, and proteome, making them difficult to promote and apply in routine clinical practice.
[0010] To address the aforementioned technical problems, this invention aims to provide a molecular subtyping method and biomarker combination for colorectal cancer immunotherapy response that features high predictive accuracy, controllable detection costs, strong clinical accessibility, and good cross-platform robustness.
[0011] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0012] This invention provides a method for predicting the response to immunotherapy in colorectal cancer patients, comprising the following steps:
[0013] S1: Obtain transcriptomic expression data from tumor tissue samples of patients with colorectal cancer to be tested;
[0014] S2: Using a prediction template containing 200 core genes, perform Nearest Template Prediction (NTP) analysis on the transcriptome expression data to classify patients into C1 immune-resistant or C2 immune-sensitive subtypes.
[0015] S3: Output the immunotherapy response prediction results based on the subtype classification results, where the C2 subtype is predicted to be an immunotherapy responder and the C1 subtype is predicted to be an immunotherapy non-responder.
[0016] Further, the 200 core gene prediction templates mentioned in step S2 include 100 C1 subtype-specific genes and 100 C2 subtype-specific genes; wherein the C1 subtype-specific genes are selected from EPM2A, EFCAB6, PCDHB12, AKAP6, PHYHIP, CCDC40, UGT2B10, HHAT, CCT6B, ASB12, PTPRT, TNRC6C, KCNB1, MYEF2, AP3B2, DGLUCY, ZMAT1, C DNF, TEX52, SLC10A4, RYR1, DCLK2, RGMB, KCNB2, PRTG, ZNF662, GP1BA, KIF19, ACCS, WDR17, OR52N1, PDE8B, G REM2, EPX, SLC13A2, KCNH6, DGKB, SHISA9, ZNF648, CDH22, PCDHB11, PCDHB10, FAM174B, RAB9B, MPPED2, QRFPR ,CCDC188,ARMCX4,SLAIN1,PCDHB9,ZNF559-ZNF177,MYRIP,CCDC149,SHF,ANKRD24,THEG,BRINP1,GSTZ1,C RYM, SOD3, ZGLP1, MATN3, FZD9, SMPDL3B, ERBB4, KSR2, CLCA2, CORO2B, PYROXD2, TSPOAP1, SDK2, SLC25A47, K LF8, WNK4, LYRM9, PACRG, KLHL13, OR5AN1, POPDC2, GALNT8, SCG3, TMEM232, CACNA2D2, CCDC169, NRIP2, DES, DOC2A, ZNF554, FN3K, ADH1A, DSCAML1, DHRS11, ARSG, CAPN3, CLMN, PRRT1B, NKAIN3, TMOD4, EIF4EBP3, KCNA6;C2 subtype-specific genes were selected from ANXA4, SQLE, FAM216B, PSTPIP2, MSRB1, CCSAP, IRAK2, CDA, KIR2DL1, ETV7, CPS1, PARP14, OR2I1P, IGF2BP3, GRINA, GSTP1, CPNE4, CNIH4, ATAD2, SERPINB5, TAAR5, S100A10, CYRIB, UBE2L6, PARP9, PRR19, MX1, EMX1, RSAD2, HOXC9, UBD, OAS3, TRPV5, LAIR2, EPRS1, CMPK2, CENPW, CACYBP, DNAJA1, CD68, YWHAZ, OAS1, PFDN2, CENPL, ZNF296, DCAF13, CCT3, DTX3L, and TAP1. TDO2, SLC4A11, GRHL3, MX2, FXYD5, CCL17, LBR, NAT8, SLC2A1, PSMC4, PSMD2, DRAP1, OAS2, H2B U1, C1QTNF9B, NEK2, VSNL1, IFIT1, HSPH1, CKS1B, FBXO45, CTSZ, PLA2G7, KIF14, H2AC18, POMP, DEGS1, DSCC1, PSMD14, KRT38, IFIT3, H2AC20, STIP1, DTL, PKP1, PRDX1, NCF2, UBE2T, PLA2G2F , CLEC4A, H4C6, ACTL8, APOBEC3A, UBE2C, PSMD8, LIN9, S100A11, FOSL1, PLEK2, S100A8, FBXO6. ;
[0017] Furthermore, the transcriptome expression data described in step S1 is obtained through RNA sequencing or gene chip detection, and the detection platform includes, but is not limited to, Illumina RNA-seq and Nanostring nCounter.
[0018] Further, the NTP algorithm in step S2 includes: constructing 200 gene expression template vectors for C1 and C2 subtypes, calculating the Pearson correlation coefficient between the transcriptome expression profile of the sample to be tested and the C1 and C2 templates, and assigning the sample to the subtype corresponding to the template with the highest correlation coefficient and meeting the preset significance threshold.
[0019] Furthermore, the immunotherapy includes immune checkpoint inhibitor therapy, wherein the immune checkpoint inhibitor is selected from anti-PD-1 antibody, anti-PD-L1 antibody and / or anti-CTLA-4 antibody; and the colorectal cancer includes microsatellite stable (MSS) and microsatellite highly unstable (MSI-H) colorectal cancer.
[0020] Secondly, the present invention provides a biomarker combination for predicting immunotherapy response in colorectal cancer patients, the biomarker combination consisting of the aforementioned 200 core genes. The biomarker combination predicts immunotherapy response in colorectal cancer patients by detecting the transcriptomic expression levels of the 200 core genes in tumor tissue.
[0021] Thirdly, this invention provides a method for constructing an immunotherapy response prediction classifier. This method integrates multi-omics data and employs a three-layer technology chain: IntNMF unsupervised clustering, SVM-RFE iterative feature selection, and NTP single-sample classification. The method includes the following steps:
[0022] A1: Obtain multi-omics data from the training sample queue, including genomic mutation data, transcriptome expression data, and single-cell transcriptome data;
[0023] A2: Feature screening was performed on the multi-omics data to obtain a candidate feature set: genomic features were selected from genes with a mutation frequency >3 and a distribution difference of P<0.25 between responders and non-responders; transcriptomic features were selected from genes expressed in at least 20% of samples and significantly associated with tumor regression rate (Spearman's rho>0.3, P<0.1); single-cell features were selected from IMT (Inflammation / Migration Tumor) and IAB (ITGAX+ Activated B) cell populations that were significantly enriched in responders at baseline, as well as scores for five core functional features: MHC-I, MHC-II, IFN response, cytoskeleton, and adhesion.
[0024] A3: The Integral Nonnegative Matrix Factorization (IntNMF) algorithm is used to perform unsupervised integrated clustering on the candidate feature set to determine the optimal number of clusters k=2, and obtain two subtypes: C1 immune resistance and C2 immune sensitivity. The IntNMF algorithm parameters are set as follows: maxiter=200, st.count=20, and the modal weights wt=c(1,1,1).
[0025] A4: Differential expression analysis was performed on the C1 and C2 subtypes. Genes with log2FC>0.5 and corrected P<0.10 were selected as subtype-specific genes. Gene set enrichment analysis (GSEA) was used to confirm the differences in immune pathways between subtypes.
[0026] A5: The Support Vector Machine Recursive Feature Elimination (SVM-RFE) algorithm is used, combined with Monte Carlo Cross-Validation (MCCV) for iterative feature selection. Through 100 nested resampling (each time the training set is divided into 70% / 30% hierarchical partitions and internal 5-fold cross-validation is used to optimize the SVM parameter C), the selection frequency of each gene in 100 iterations is counted, and the top 100 genes with the highest selection frequency in each subtype are selected to construct a prediction template containing 200 core genes.
[0027] A6: The nearest template prediction (NTP) algorithm is used to assign subtypes to the test samples based on the 200 core gene prediction templates.
[0028] This invention provides an immunotherapy response prediction system, comprising:
[0029] The data acquisition module is used to acquire transcriptome expression data of tumor tissue samples from patients with colorectal cancer to be tested. The transcriptome expression data is obtained by RNA sequencing or gene chip detection.
[0030] The classification module is used to perform nearest template prediction (NTP) analysis on the transcriptome expression data using a prediction template containing 200 core genes, classifying patients into C1 immune-resistant or C2 immune-sensitive types; the 200 core genes include 100 C1 subtype-specific genes and 100 C2 subtype-specific genes.
[0031] The output module is used to output the prediction results of immunotherapy response.
[0032] This invention provides the application of the aforementioned biomarker combination in the preparation of a kit for predicting immunotherapy response in colorectal cancer patients. The kit contains specific primer pairs and / or probes for detecting the expression levels of the aforementioned 200 core genes.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] (1) Significantly improved prediction accuracy. The IRMS molecular typing method of this invention achieved an AUC of 0.87 in the training cohort (sensitivity 74%, specificity 100%) and an AUC of 0.70 in the external colorectal cancer validation cohort (sensitivity 60%, specificity 80%), which is significantly better than existing single biomarkers such as CPS score (AUC=0.62), TMB (AUC=0.55), and CXCL13⁺ T cell characteristics (AUC=0.67) (Delong test, P<0.05 for all). The net reclassification improvement index (NRI) of IRMS was 0.42 relative to CPS and 0.51 relative to TMB.
[0035] (2) Low detection cost and strong clinical accessibility. This invention only requires the detection of transcriptome expression profiles of 200 core genes, supports multiple detection platforms such as RNA-seq and Nanostring nCounter (cross-platform classification consistency rate of 91.7%, Kappa=0.83), and does not require whole genome sequencing or multi-omics detection, which greatly reduces the detection cost and technical threshold.
[0036] (3) Synergistic effect of multi-dimensional technologies. This invention is the first to integrate information from three dimensions—genomic mutation, transcriptome expression, and single-cell functional characteristics—using the IntNMF algorithm to comprehensively characterize the tumor immune microenvironment. Through an iterative feature screening strategy of SVM-RFE+MCCV (100 times), the average selection frequency of 200 core genes in 100 resamplings was >60 times, ensuring the stability and generalization ability of the classifier. After excluding single-cell features, the AUC decreased from 0.70 to 0.58 (a decrease of 17.1%), verifying the necessity of multi-omics integration.
[0037] (4) It has independent prognostic value. Multivariate Cox regression analysis confirmed that IRMS classification is a prognostic factor independent of clinicopathological variables (adjusted HR=0.35, 95% CI: 0.15-0.81, p=0.013). The median PFS for C2 subtype patients was 15.6 months, and for C1 subtype it was 4.2 months (HR=0.28, p=0.004).
[0038] (5) It has good complementarity with the existing CMS classification. IRMS successfully classified all CMS1 into C2 and all CMS3 into C1, while identifying potential immunotherapy beneficiaries missed by the CMS classification in CMS2 (37.5%) and CMS4 (25%), thus effectively supplementing the CMS classification.
[0039] (6) The NTP algorithm has a significantly lower risk of overfitting in predictive performance than common machine learning algorithms such as SVM, random forest, XGBoost and logistic regression. It is also naturally robust to cross-platform technology differences and is more suitable for clinical application.
[0040] (7) The 200-gene panel was optimized and verified: After comparing the classifier performance of 20 to 500 genes, it was confirmed that 200 genes (100 per subtype) achieved the best balance between prediction performance and scheme stability. The AUC of the 20-gene scheme was only 0.52, while the scheme with more than 300 genes showed overfitting. Attached Figure Description
[0041] Figure 1a. Overall map of immunotherapy response molecular subtype (IRMS). Displayed from top to bottom: clinical information, gene mutations, gene expression profiles, and functional characteristic scores.
[0042] Figure 1 b. The bar chart shows the distribution of treatment responses for different IRMS subtypes in the FIRM cohort.
[0043] Figure 1 c. Comparison of AUC values for different predictive metrics in the FIRM cohort.
[0044] Figure 1 d. Schematic diagram of the IRMS classifier construction process.
[0045] Figure 1 e. The bar chart shows the distribution of treatment responses among different IRMS subtypes in the Laurent-Puig (2023) colorectal cancer immunotherapy cohort.
[0046] Figure 1 f, Comparison of AUC values for different predictive metrics in the Laurent-Puig (2023) cohort.
[0047] Figure 1 g. Kaplan–Meier curves compare overall survival (OS) and progression-free survival (PFS) in patients with C1 and C2 subtypes in the Laurent-Puig (2023) cohort.
[0048] Figure 1 h, the correspondence between IRMS subtypes and consensus molecular subtypes (CMS) of colorectal cancer in the FIRM cohort.
[0049] AUC is calculated based on the receiver operating characteristic (ROC) curve. Statistical significance is determined using Fisher's exact test. Figure 1 b、 Figure 1 e) and log-rank test for survival analysis ( Figure 1g) Assessment. CPS, combined positive score; AUC, area under the curve; TLS, tertiary lymphoid structures; TCR, T cell receptor; TMB, tumor mutation burden; SVM-RFE, support vector machine-recursive feature elimination; MCCV, Monte Carlo cross-validation; NTP, nearest template prediction; CMS, consensus molecular subtypes; IMT, inflammation / migration tumor cells; IAB, ITGAX⁺ activated B cells (ITGAX). + Activated B cells).
[0050] Figure 2 a. The dot plot shows the cluster predictor index (blue) and gap statistics (red) for different numbers of clusters (k).
[0051] Figure 2 b. The heatmap shows the application of the IRMS classifier in the Laurent-Puig (2023) colorectal cancer cohort. The x-axis represents the predicted subtype, i.e., C1 or C2; the y-axis represents the 200 template genes used for prediction.
[0052] Figure 2 c. A multivariate Cox regression analysis of the relationship between clinical variables, mutation status, and IRMS subtype and progression-free survival (PFS) in the Laurent-Puig (2023) cohort. This analysis included stage II–III patients.
[0053] Figure 2 d– Figure 2 Performance evaluation of k, IRMS in predicting immunotherapy response in multiple cohorts, including: VanAllen, 2015 melanoma cohort ( Figure 2d); Gide, 2019 melanoma cohort ( Figure 2 e); Hu, 2023 Non-small Cell Lung Cancer (NSCLC) Cohort ( Figure 2 f); Jung, 2019 NSCLC Queue ( Figure 2 g); Uppaluri, 2020 Head and Neck Squamous Cell Carcinoma (HNSC) Cohort ( Figure 2 h); Pusztai, 2021 Breast Cancer Cohort ( Figure 2 i); Kim, 2018 Gastric Cancer Cohort ( Figure 2 j); and Li, 2024 hepatocellular carcinoma (HCC) cohort ( Figure 2 k). The ROC curve on the left is used to evaluate predictive performance; the bar chart on the right shows the distribution of treatment response in different IRMS subtypes.
[0054] Figure 2 l– Figure 2 o, Kaplan–Meier curves compare overall survival (OS) or progression-free survival (PFS) of C1 and C2 subtypes in different cohorts, including VanAllen, 2015 melanoma cohort ( Figure 2 (l), Nathanson, 2017 melanoma cohort ( Figure 2 m), Gide, 2019 melanoma cohort ( Figure 2 n) and Jung, 2019 NSCLC Queue ( Figure 2 o).
[0055] AUC was calculated based on the receiver operating characteristic (ROC) curve. Statistical significance was determined using multivariate Cox regression analysis. Figure 2 c) Fisher's exact test ( Figure 2 d– Figure 2 k) and log-rank test ( Figure 2 l– Figure 2o) Evaluation is performed. Mut, mutation; HR, hazard ratio; NSCLC, non-small cell lung cancer; HNSC, head and neck squamous cell carcinoma; HCC, hepatocellular carcinoma; OS, overall survival; PFS, progression-free survival; IRMS, immunotherapy response molecular subtype.
[0056] Figure 3 a. The heatmap shows the consistency matrix generated by the IntNMF method under the condition of k = 2, which is used to evaluate the stability of molecular typing results. The closer the matrix value is to 1, the higher the probability that the corresponding sample is consistently assigned to the same cluster in multiple independent runs, indicating stronger cluster stability.
[0057] Figure 3 b. The silhouette coefficient plot illustrates clustering performance and is used to evaluate the consistency within clusters and the degree of separation between different clusters. The silhouette coefficient ranges from -1 to 1, where the closer the value is to 1, the higher the degree of matching between the sample and its cluster, and the better the separation from other clusters.
[0058] Figure 3 c. The heatmap illustrates the differences in immune-related pathway scores between the C1 and C2 subtypes. Immune-related pathways are grouped according to functional categories. Pathway scores were calculated using Gene Set Variation Analysis (GSVA). Detailed Implementation
[0059] Example 1: Construction of an IRMS molecular typing model
[0060] 1.1 Patient Cohort and Sample Collection
[0061] This embodiment utilizes pre-treatment tissue samples from 27 patients who underwent neoadjuvant immunochemotherapy (nICT) for colorectal cancer surgery to construct an IRMS molecular subtyping model. Patient inclusion criteria were: pathologically confirmed locally advanced rectal cancer (LARC), clinical stage II-III, receiving neoadjuvant immunochemotherapy (including anti-PD-1 antibody combined with mFOLFOX6 regimen), possessing complete pre-treatment tumor tissue biopsy samples, and evaluable treatment response data. The patient cohort characteristics were as follows: median age 58 years (range 32-75 years), 16 males and 11 females. Treatment response was assessed using the tumor regression rate (TRG) standard, with 14 responders (TRG 0-1) and 13 non-responders (TRG 2-3).
[0062] 1.2 Multi-omics Data Acquisition and Preprocessing
[0063] Collect the following three types of omics data from all patients:
[0064] (1) Genomic mutation data: The whole exome sequencing WES platform (Illumina NovaSeq 6000, PE150, average depth 200×) was used to detect somatic mutations using Mutect2 (v.4.2), filtering out streptavidin mutations and germline mutations, and retaining non-synonymous mutations in the exon regions.
[0065] (2) Bulk RNA sequencing data: Strand-specific RNA-seq library construction was performed using Illumina TruSeq StrandedTotal RNA Library Prep, and PE150 sequencing was performed on the NovaSeq 6000 platform, yielding approximately 50 M clean reads per sample. Alignment to the GRCh38 reference genome was performed using STAR (v.2.7.10), gene expression quantification was performed using featureCounts (v.2.0.3), and normalization was performed using TPM (Transcripts Per Million).
[0066] (3) Single-cell transcriptome sequencing data: Single-cell capture and library construction were performed using the 10x Genomics Chromium platform and v3 kit, and sequencing was performed on NovaSeq 6000. The target number of cells per sample was 8,000-10,000. Raw data processing and quantification were performed using CellRanger (v.7.0), and downstream analysis was performed using Seurat (v.4.3).
[0067] 1.3 Multi-omics feature screening
[0068] To ensure sufficient variables are included to achieve stable subtyping, the statistical thresholds for feature selection were appropriately relaxed. The feature selection results for each modality are as follows:
[0069] (1) Genomic characteristics: Genes with a mutation frequency >3 and a distribution difference of P<0.25 between responders and non-responders were selected, resulting in 32 mutant genes. Representative genes include LPHN3, DBH, DSCAM, KMT2D, APOB, MUC16, SYNE1, etc.
[0070] (2) Transcriptome characteristics: Genes expressed in at least 20% of the samples and significantly associated with tumor regression rate (Spearman's rho>0.3, P<0.1) were selected, resulting in a total of 515 genes. Representative genes include immune-related genes such as CA1, CLCA1, CXCL10, CXCL11, IDO1, IFNG, and NKG7.
[0071] (3) Single-cell characteristics: IMT (Inflammation / Migration Tumor) and IAB (ITGAX+Activated B) cell populations that were significantly enriched in responders at baseline were selected, along with scores for five core functional characteristics: MHC-I, MHC-II, IFN response, cytoskeleton, and adhesion. Functional characteristic scores were quantified using the GSVA algorithm (v.2.0.5) with default Kimono parameters.
[0072] 1.4 IntNMF Unsupervised Integration Clustering
[0073] The optimal number of clusters was determined using the `getClustNum` function from the R package `MOVICS` (v.0.99.17). The cluster prediction index (CPI) was 0.92 at k=2, but dropped sharply to 0.71 at k=3; the Gap statistic peaked at 1.24 at k=2 and was 0.86 at k=3. Combining these two indicators, the optimal number of clusters was determined to be k=2.
[0074] The IntNMF multi-omics clustering algorithm (v.1.3.0) was used to integrate and analyze three data modalities. Parameter settings: maxiter=200, st.count=20, weights for each omics wt=c(1,1,1). By calculating the convergence of the reconstruction error of each omics during the iteration process, it was found that the objective function tended to stabilize after about 150 iterations (convergence tolerance <10^-4).
[0075] Robustness assessment of clustering results:
[0076] (1) Consistency matrix analysis: Under the condition of k=2, the mean of the diagonal blocks of the consistency matrix is 0.91 (C1) and 0.93 (C2), and the mean of the off-diagonal blocks is 0.12, indicating that the clustering assignment is highly stable.
[0077] (2) Contour analysis: The average contour width of subtype C1 is 0.58, the average contour width of subtype C2 is 0.64, and the overall average contour width is 0.61 (>0.5 is a good clustering threshold).
[0078] This led to the definition of two subtypes with different biological characteristics: C1, the immune-resistant type (13 cases in total, including 8 non-responders and 5 responders), and C2, the immune-sensitive type (14 cases in total, including 9 responders and 5 non-responders). In the FIRM cohort, all non-responders were classified into subtype C1 (62%, 8 / 13), while subtype C2 accounted for 64% (9 / 14) of all responders.
[0079] 1.5 Differential expression analysis among subtypes and preliminary gene screening
[0080] Differential expression analysis was performed between the C1 and C2 subtypes using Bulk RNA-seq data (DESeq2v.1.38.3). Genes with log2FC>0.5 and corrected P<0.10 (Benjamini-Hochberg correction) were identified as subtype-specific genes. Results: 665 genes were specifically expressed in C1 (upregulated in C1), and 332 genes were specifically expressed in C2 (upregulated in C2). C2 upregulated genes were significantly enriched in immune activation pathways such as IFN-γ response (NES=2.34, FDR<0.001), antigen presentation (NES=2.18, FDR<0.001), and T cell activation (NES=2.05, FDR=0.002); C1 upregulated genes were enriched in immunosuppression-related pathways such as extracellular matrix tissue (NES=1.89, FDR=0.008) and epithelial-mesenchymal transition (NES=1.76, FDR=0.015).
[0081] 1.6 SVM-RFE+MCCV Iterative Feature Selection
[0082] To eliminate redundant features and optimize computational efficiency, a screening process based on random resampling and recursive feature elimination (RFE) was developed. The specific workflow is as follows:
[0083] (1) In each MCCV iteration, the discovery dataset (27 cases) is randomly divided into a training set (70%, 19 cases) and a test set (30%, 8 cases) to ensure that the ratio of responders / non-responders between the two groups is balanced (stratified sampling).
[0084] (2) In each training set, the penalty parameter C of the linear kernel SVM is optimized by combining internal 5-fold hierarchical cross-validation with grid search (candidate values: 10^-3, 10^-2, 10^-1, 1, 10, 10^2, 10^3). The optimal C values are concentrated between 10^-1 and 1.
[0085] (3) Execute the SVM-RFE algorithm under optimal parameters, sort according to the absolute value of feature weights, recursively eliminate the features with the smallest weights, and retain the top 100 features in each iteration.
[0086] (4) The nested resampling process is repeated 100 times independently, performing 100 recursive feature eliminations. In each feature elimination process, 5-fold cross-validation is used to find the optimal hyperparameter C.
[0087] (5) The selection frequency of each gene in 100 iterations was counted, and the top 100 genes with the highest selection frequency in each subtype were selected (selection frequency threshold: C1 gene > 60 times, C2 gene > 55 times), and a total of 200 core gene prediction templates were constructed.
[0088] 1.7 NTP Algorithm Subtype Prediction
[0089] Subtype assignment is performed in an external validation queue using the nearest template prediction (NTP) algorithm. NTP is a flexible classification strategy based on single samples, robust to cross-platform technology differences. Its core computational steps are as follows:
[0090] (1) Template construction: Based on the iterative feature screening of SVM-RFE+MCCV in the training set, template genes for C1 subtype and C2 subtype (100 each) were obtained.
[0091] (2) Sample matching: For the expression profile of 200 genes of the sample to be tested, calculate the similarity between it and the C1 template and the C2 template.
[0092] (3) Subtype assignment: The sample is assigned to the subtype corresponding to the template with the highest similarity and the significance threshold (FDR<0.05, based on 1000 permutation tests).
[0093] (4) Quality control: Samples that fail to reach the significance threshold (too low similarity or mismatch between the two templates) are marked as unclassified and excluded from further analysis. The classification success rate is above 92.3% across all analysis cohorts.
[0094] Example 2: Performance evaluation of the IRMS classifier in a colorectal cancer validation cohort
[0095] 2.1 Verify queue characteristics
[0096] The IRMS classifier constructed in Example 1 was independently validated in the colorectal cancer immunotherapy cohort disclosed at Laurent-Puig 2023 (hereinafter referred to as the LP2023 cohort). This cohort included 64 patients with stage II-III colorectal cancer who received immunotherapy, with 42 responders.
[0097] 2.2 Predictive Performance Evaluation
[0098] Association analysis between IRMS classification results and treatment response showed:
[0099] (1) Of the 64 successfully classified samples, 29 were of subtype C2 and 35 were of subtype C1. 86% (25 / 29) of the C2 subtype were responders and 52% (17 / 35) of the C1 subtype were responders.
[0100] (2) Compared with the C1 subtype, patients in the C2 subtype showed a significantly better treatment response (Fisher exact test, p=0.006).
[0101] (3) Classifier performance indicators: AUC=0.70 (95%CI: 0.58-0.82), sensitivity=60%, specificity=80%, positive predictive value PPV=76.5%, negative predictive value NPV=91.7%.
[0102] (4) Performance comparison with existing immunotherapy biomarkers: CPS score AUC=0.58 (95%CI: 0.45-0.71), TMB AUC=0.55 (95%CI: 0.42-0.68), CXCL13+ T cell characteristic AUC=0.62 (95%CI: 0.49-0.75). The AUC of IRMS was significantly better than the above biomarkers (Delong test, all P<0.05).
[0103] 2.3 Survival Analysis
[0104] Long-term survival follow-up data analysis:
[0105] Patients classified as C2 had significantly better overall survival (OS) and progression-free survival (PFS) than those classified as C1 (OS: p = 0.017; PFS: p = 0.004; Figure 1g). Multivariate Cox analysis further confirmed that IRMS is an independent prognostic factor, suggesting its potential to predict long-term survival benefit in colorectal cancer patients receiving immunotherapy (Figure 2c).
[0106] Example 3: Comparative Analysis of IRMS and CMS Typing
[0107] 3.1 Correspondence between IRMS and CMS
[0108] In the FIRM cohort, IRMS was compared with the consensus molecular subtyping of colorectal cancer (CMS). CMS classification was performed using CMScaller (v.0.99.17) based on bulk RNA-seq data, and the NTP algorithm was used to match CMS1-CMS4 templates.
[0109] The cross-correlation analysis results are as follows: all CMS1 (immune type, n=5) cases were classified into C2 (immunely sensitive type); of CMS2 (epithelial type, n=8) cases, 5 cases were classified into C1 and 3 cases into C2; all CMS3 (metabolic type, n=6) cases were classified into C1 (immunely resistant type); and of CMS4 (mesenchymal type, n=8) cases, 6 cases were classified into C1 and 2 cases into C2. A significant association was found between IRMS and CMS subtypes (Fisher's exact test, p=0.005).
[0110] 3.2 Complementarity Analysis
[0111] Notably, IRMS identified a subset of potential immunotherapy responders in the CMS2 and CMS4 subtypes:
[0112] (1) Of the 5 CMS2 patients, 3 were classified as C2 subtype by IRMS, and 3 of them were actual immunotherapy responders.
[0113] (2) Of the 8 CMS4 patients, 4 were classified as C2 subtype by IRMS, of which 4 were actual responders.
[0114] (3) Overall, IRMS identified 7 C2 subtype patients in CMS2 and CMS4, all of whom were actual immunotherapy responders, while CMS typing could not identify this group of beneficiaries. This indicates that IRMS provides additional predictive information on top of traditional CMS typing and can be used to screen potential beneficiaries missed by CMS typing.
[0115] Comparative Example 1: Comparison of Predictive Performance of IRMS and Single Biomarkers
[0116] 1.1 Comparison of Scheme Design
[0117] In the FIRM cohort (27 cases), the predictive performance of IRMS was compared with four conventional single immunotherapy biomarkers: CPS score (PD-L1 combined positive score, detected by 22C3 antibody), TMB (tumor mutation burden, calculated from whole-exome sequencing data), MSI status (analyzed by MSIsensor), and CXCL13+ T cell characteristics (expression score based on RNA-seq). The optimal cutoff value for each biomarker was determined in the training set using the Youden Index.
[0118] 1.2 Performance Comparison
[0119] The performance metrics of the various prediction methods are compared below:
[0120] (1) IRMS: AUC=0.87, sensitivity=74%, specificity=100%, positive predictive value PPV=100%, negative predictive value NPV=72.7%, accuracy=85.2%.
[0121] (2) CPS score (cutoff value = 10): AUC = 0.62, sensitivity = 57%, specificity = 69%, PPV = 66.7%, NPV = 60.0%, accuracy = 62.9%.
[0122] (3) TMB (cutoff value = 12 mut / Mb): AUC = 0.55, sensitivity = 50%, specificity = 62%, PPV = 58.3%, NPV = 53.3%, accuracy = 55.6%.
[0123] (4) CXCL13+ T cell characteristics (cutoff value = median expression level): AUC = 0.67, sensitivity = 64%, specificity = 69%, PPV = 69.2%, NPV = 64.3%, accuracy = 66.7%.
[0124] IRMS significantly outperformed the aforementioned single biomarkers across all key metrics (Delong test comparison of AUC: IRMS vs CPS, P=0.008; IRMS vs TMB, P=0.003; IRMS vs CXCL13+ feature, P=0.021). Net reclassification improvement index (NRI) analysis also showed that the overall NRI of IRMS relative to CPS score was 0.42 (P=0.009), and the NRI relative to TMB was 0.51 (P=0.004).
[0125] Comparative Example 2: Typing control without using single-cell features
[0126] 2.1 Comparison Scheme
[0127] In the IRMS construction process, a control experiment was set up: single-cell transcriptome-derived features (IMT, IAB and five functional features of MHC-I, MHC-II, IFN response, cytoskeleton and adhesion) were excluded, and only 32 genomic mutant genes and 515 bulk transcriptome genes were used for IntNMF clustering (k=2). The remaining steps were consistent with Example 1.
[0128] 2.2 Comparison Results
[0129] The control classification results after excluding single-cell features differed significantly from the complete IRMS protocol:
[0130] (1) Decreased cluster stability: After excluding single-cell features, the mean of the diagonal blocks of the consistency matrix decreased to 0.78 (C1) and 0.82 (C2), while the original scheme was 0.91 and 0.93 respectively; the average contour width decreased from 0.61 to 0.52 (a decrease of 14.8%).
[0131] (2) Decreased predictive performance: In the LP2023 validation cohort, the AUC of the control protocol decreased to 0.58 (0.70 for the full IRMS protocol), the sensitivity decreased to 45%, and the specificity decreased to 65%.
[0132] (3) Reduced differentiation of immune pathways: The GSVA scores of immune-related pathways between C1 and C2 subtypes were significantly reduced across the board. Among them, the difference in IFN-γ response pathway increased from P=0.002 in the original protocol to P=0.08, losing statistical significance.
[0133] These results indicate that single-cell derived functional features (especially IMT and IAB cell population features) play a crucial role in accurately characterizing the state of the tumor immune microenvironment and improving the predictive performance of classifiers.
[0134] Comparative Example 3: Performance Comparison of Classifiers with Different Numbers of Genes
[0135] 3.1 Comparison Scheme
[0136] To verify the optimality of the 200-gene panel, control schemes with different numbers of genes were set up: the top 20, 50, 100, 150, 200, 300, and 500 genes were selected from the 100 iterations of SVM-RFE to construct prediction templates, and all schemes were compared under the same training / testing partition.
[0137] 3.2 Comparison Results
[0138] The predictive performance of different gene count schemes is as follows:
[0139] (1) 20-gene scheme: AUC=0.52, poor consistency among genes, the average selection frequency of each gene in 100 iterations was only 18 times, indicating that the small gene panel is not stable enough under the MCCV framework.
[0140] (2) 50-gene scheme: AUC=0.58, the performance is improved but still insufficient.
[0141] (3) 100-gene scheme (50 C1 / 50 C2 each): AUC=0.63, the consistency of the classification results with the complete scheme is 71%.
[0142] (4) 150 gene scheme (75 C1 / 75 C2 each): AUC=0.67, taxonomic consistency is 85%.
[0143] (5) 200-gene scheme (the scheme of this invention): AUC=0.70, the average selection frequency of each gene in 100 iterations is >60 times, and the stability is excellent.
[0144] (6) 300 gene scheme: AUC=0.69, the performance is no longer improved compared with the 200 gene scheme, and the detection cost is increased.
[0145] (7) 500 gene scheme: AUC=0.65, showing a downward trend in performance, indicating that too many genes introduce noise.
[0146] The results showed that 200 core genes (100 per subtype) achieved the optimal balance between prediction performance and scheme stability. Too few genes could not stably capture subtype features, while too many introduced noise.
[0147] Example 4: Cross-platform stability verification of a 200-gene panel
[0148] To evaluate the cross-platform applicability of the IRMS classifier, IRMS classification was performed using RNA-seq and Nanostring Counter data from the same batch of samples, and the classification consistency was compared. Analysis of 12 samples with both RNA-seq and Nanostring Counter data showed a classification consistency rate of 91.7% (11 / 12) between the two platforms, with a Kappa coefficient of 0.83 (95% CI: 0.56–1.00), indicating that the IRMS classifier has good robustness across different technology platforms.
Claims
1. A method for predicting immunotherapy response in colorectal cancer patients, characterized in that, Includes the following steps: S1: Obtain transcriptomic expression data from tumor tissue samples of patients with colorectal cancer to be tested; S2: Using a prediction template containing 200 core genes, perform nearest template prediction (NTP) analysis on the transcriptome expression data to classify patients into C1 immune-resistant or C2 immune-sensitive subtypes. S3: Output the immunotherapy response prediction results based on the subtype classification results, where the C2 subtype is predicted to be an immunotherapy responder and the C1 subtype is predicted to be an immunotherapy non-responder; the colorectal cancer includes microsatellite stable (MSS) and microsatellite highly unstable (MSI-H) colorectal cancer.
2. The method according to claim 1, characterized in that, The 200 core gene prediction templates mentioned in step S2 include: 100 C1 subtype-specific genes selected from EPM2A, EFCAB6, PCDHB12, AKAP6, PHYHIP, CCDC40, UGT2B10, HHAT, CCT6B, ASB12, PTPRT, TNRC6C, KCNB1, MYEF2, AP3B2, DGLUCY, ZMAT1, CDNF, TEX52, SLC10A4, and RYR1. , DCLK2, RGMB, KCNB2, PRTG, ZNF662, GP1BA, KIF19, ACCS, WDR17, OR52N1, PDE8B, GREM2, EPX, SLC13A2, K CNH6, DGKB, SHISA9, ZNF648, CDH22, PCDHB11, PCDHB10, FAM174B, RAB9B, MPPED2, QRFPR, CCDC188, ARMCX 4. SLAIN1, PCDHB9, ZNF559-ZNF177, MYRIP, CCDC149, SHF, ANKRD24, THEG, BRINP1, GSTZ1, CRYM, SOD3, Z GLP1, MATN3, FZD9, SMPDL3B, ERBB4, KSR2, CLCA2, CORO2B, PYROXD2, TSPOAP1, SDK2, SLC25A47, KLF8, WNK 4. LYRM9, PACRG, KLHL13, OR5AN1, POPDC2, GALNT8, SCG3, TMEM232, CACNA2D2, CCDC169, NRIP2, DES, DOC2 A. ZNF554, FN3K, ADH1A, DSCAML1, DHRS11, ARSG, CAPN3, CLMN, PRRT1B, NKAIN3, TMOD4, EIF4EBP3, KCNA6;And 100 C2 subtype-specific genes, selected from ANXA4, SQLE, FAM216B, PSTPIP2, MSRB1, CCSAP, IRAK2, CDA, KIR2DL1, ETV7, CPS1, PARP14, OR2I1P, IGF2BP3, GRINA, GSTP1, CPNE4, CNIH4, ATAD2, SERPINB5, TAAR5, S100A10, CYRIB, UBE2L6, PARP9, PRR19, MX1, EMX1, RSAD2, HOXC9, UBD, OAS3, TRPV5, LAIR2, EPRS1, CMPK2, CENPW, CACYBP, DNAJA1, CD68, YWHAZ, OAS1, PFDN2, CENPL, ZNF296, DCAF13, CCT3, DTX3L, T AP1, TDO2, SLC4A11, GRHL3, MX2, FXYD5, CCL17, LBR, NAT8, SLC2A1, PSMC4, PSMD2, DRAP1, OAS2, H2BU1, C1QTNF9B, NEK2, VSNL1, IFIT1, HSPH1, CKS1B, FBXO45, CTSZ, PLA2G7, KIF14, H2AC18, POM P, DEGS1, DSCC1, PSMD14, KRT38, IFIT3, H2AC20, STIP1, DTL, PKP1, PRDX1, NCF2, UBE2T, PLA2G2 F. CLEC4A, H4C6, ACTL8, APOBEC3A, UBE2C, PSMD8, LIN9, S100A11, FOSL1, PLEK2, S100A8, FBXO6. ; 3. The method according to claim 1, characterized in that, The transcriptome expression data in step S1 is obtained by RNA sequencing or gene chip detection; the NTP algorithm in step S2 includes: calculating the similarity measure between the transcriptome expression profile of the sample to be tested and the C1 and C2 subtype templates, and classifying the sample into the subtype corresponding to the template with the highest similarity; before step S2, a batch effect correction step is also included for the transcriptome expression data.
4. The method according to claim 1, characterized in that, The immunotherapy includes immune checkpoint inhibitor therapy, wherein the immune checkpoint inhibitor is selected from anti-PD-1 antibody, anti-PD-L1 antibody and / or anti-CTLA-4 antibody, or the immunotherapy is neoadjuvant immunochemotherapy (nICT).
5. The application of a combination of biomarkers in the preparation of a kit for predicting immunotherapy response in colorectal cancer patients, characterized in that, The biomarker combination comprises the 200 core genes as described in claim 2.
6. A method for constructing an immunotherapy response prediction classifier, characterized in that, Includes the following steps: A1: Obtain multi-omics data from the training sample cohort, including genomic mutation data, transcriptome expression data, and single-cell transcriptome data; A2: Feature screening was performed on the multi-omics data to obtain a candidate feature set: genomic features were selected from genes with a mutation frequency >3 and a distribution difference of P<0.25 between responders and non-responders; transcriptomic features were selected from genes expressed in at least 20% of the samples and significantly associated with tumor regression rate (Spearman's rho>0.3, P<0.1); single-cell features were selected from IMT (Inflammation / Migration Tumor) and IAB (ITGAX+Activated B) cell population functional characteristic scores that were significantly enriched in responders at baseline; A3: The Integral Nonnegative Matrix Factorization (IntNMF) algorithm is used to perform unsupervised clustering on the candidate feature set to determine the optimal number of clusters k=2, and obtain two subtypes: C1 immune-resistant and C2 immune-sensitive. A4: Differential expression analysis was performed on the two subtypes, and genes with log2FC>0.5 and corrected P<0.10 were selected as subtype-specific genes; A5: The Support Vector Machine Recursive Feature Elimination (SVM-RFE) algorithm is used, combined with Monte Carlo Cross-Validation (MCCV) for iterative feature selection. The top 100 genes with the highest selection frequency in each subtype are determined through multiple nested resampling, and a prediction template containing 200 core genes is constructed. A6: The nearest template prediction (NTP) algorithm is used to assign subtypes to the test samples based on the 200 core gene prediction templates.
7. The method according to claim 6, characterized in that, The single-cell functional characteristic score mentioned in step A2 includes five core functional characteristic scores: MHC-I, MHC-II, IFN response, cytoskeleton, and adhesion, which are quantified using the GSVA algorithm. The IntNMF algorithm in step A3 is set with the following parameters: maxiter=200, st.count=20, and the weights of each data mode wt=c(1,1,1); the nested resampling process in step A5 is repeated independently 100 times, each time the training set is randomly divided into 70% training set and 30% test set, and the SVM parameters are optimized by combining internal 5-fold hierarchical cross-validation with grid search.
8. The method according to claim 6, characterized in that, In step A5, the SVM-RFE algorithm uses a linear kernel support vector machine to recursively eliminate the features with the smallest weights until the top 100 features are retained. The NTP algorithm described in step A6 includes: introducing a permutation test (n=1000) in non-colorectal cancer samples, and retaining only samples with a false discovery rate (FDR) <0.25 for subtype assignment.
9. An immunotherapy response prediction system, characterized in that, include: The data acquisition module is used to acquire transcriptome expression data from tumor tissue samples of patients with colorectal cancer to be tested; A classification module is used to perform nearest template prediction (NTP) analysis on the transcriptome expression data using a prediction template containing 200 core genes to classify patients into C1 immune-resistant or C2 immune-sensitive types, wherein the 200 core genes are as described in claim 2. The output module is used to output the prediction results of immunotherapy response.
10. The system according to claim 9, characterized in that: The classification module also includes a preprocessing unit for standardizing and normalizing transcriptome expression data.