Random forest classifier and molecular subtype classifier used for molecular typing of esophageal squamous cell carcinoma, and method for constructing them.
A random forest classifier using 314 genes unifies molecular typing for esophageal squamous cell carcinoma, identifying four subtypes (ECMS1-4) through multi-omics data fusion, improving treatment efficacy and survival by personalizing therapy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SHANXI MEDICAL UNIV
- Filing Date
- 2025-07-14
- Publication Date
- 2026-05-01
AI Technical Summary
Existing molecular typing systems for esophageal squamous cell carcinoma lack broad applicability and consistency due to regional variations and differences in algorithms, hindering personalized treatment plans and leading to treatment resistance and metastasis.
A random forest classifier incorporating 314 characteristic genes identifies four consensus molecular subtypes (ECMS1-4) through a method involving whole-genome sequencing, transcriptome analysis, small RNA sequencing, and methylation data fusion, using spectral clustering and consensus molecular typing to unify subtype classification.
The method provides a unified molecular typing system for esophageal squamous cell carcinoma, revealing subtype-specific characteristics and therapeutic targets, enhancing treatment efficacy and survival rates by personalizing therapy based on molecular subtypes.
Smart Images

Figure 2026073927000001_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of bioinformatics and computational biology, and specifically relates to a random forest classifier and a molecular subtype classifier used for molecular typing of esophageal squamous cell carcinoma, as well as a construction method.
Background Art
[0002] Esophageal cancer is one of the common cancer types in the world. The number of newly reported cases worldwide in 2020 exceeded 6 million, the number of death cases exceeded 5 million, its prognosis survival rate is relatively poor, and it seriously threatens the physical health of the general public. It is classified into esophageal adenocarcinoma (EAC) or esophageal squamous cell carcinoma (ESCC) according to histological classification. Since the early symptoms of esophageal squamous cell carcinoma are not obvious, esophageal squamous cell carcinoma patients are often in the terminal stage when they are first definitely diagnosed, with a low surgical resection rate, poor treatment effect, and a 5-year survival rate of only 10-15%. For other types of tumors, esophageal squamous cell carcinoma lacks effective and specific targeted therapeutic drugs. Although conventional surgery, radiotherapy and chemotherapy have achieved great progress, the treatment effect is limited, and local recurrence and distant metastasis are still important causes of treatment failure in esophageal squamous cell carcinoma. In addition, there is a high degree of heterogeneity between the inside of esophageal squamous cell carcinoma and patients, with diversity in its molecular biology and clinical characteristics. The means of classification based on traditional histological features, tumor size, staging, etc. cannot provide the connotation of intrinsic molecular biology, making it difficult to determine whether patients can benefit from a specific treatment plan, which is an important cause of treatment resistance, recurrence and metastasis, and has become a bottleneck restricting esophageal cancer treatment research. The treatment benefits of esophageal squamous cell carcinoma patients are limited, and the survival rate of patients with esophageal squamous cell carcinoma remains low.
[0003] Improving the level of care for esophageal squamous cell carcinoma and extending patient survival time is of great significance to improving the health level and quality of life of the general public, and has significant social and economic value. Molecular typing is the most effective policy for deciphering tumor heterogeneity and is a crucial strategy for achieving accurate treatment. Patients with tumors of different molecular subtypes generally have their own specific clinical characteristics, such as prognosis, outcome, or drug sensitivity. Molecular typing of tumors has already been used to guide clinical treatment; for example, breast cancer treatment guidelines recommend different treatment policies for five different molecular subtypes. It is necessary to find molecular typing methods for esophageal squamous cell carcinoma, interpret the subtype-specific regulatory mechanisms of esophageal squamous cell carcinoma, and accurately design personalized treatment plans based on this. In recent years, molecular typing research related to esophageal squamous cell carcinoma has been conducted, examining the heterogeneity of ESCC by exploring gene expression, copy number variation, and / or mutation spectrum. For example, two tumor analysis subtypes identified based on RNA expression profiles of Asian populations have been reported. Subtype I is associated with immune responses, while subtype II is primarily associated with ectoderm development, glycolysis processes, and cell proliferation. However, these studies are based on mono-omics data, which ignore the heterogeneity of other omics data. The Cancer and Tumor Gene Map (TCGA) project conducted a multi-omics molecular typing study of esophageal squamous cell carcinoma and determined three subtypes (C1-C3) through a comprehensive multi-omics analysis of 90 ESCC patients (including only Asian and Vietnamese populations). C1 is associated with alterations in the NRF2 pathway. C2 is associated with high-frequency somatic mutations in NOTCH1 or ZNF750. C3 is characterized by activation of the PI3K pathway. However, the typing system developed in this study exhibits regional variations and lacks broad applicability. Differences in algorithms and patient cues mean that the results of different typing studies are not the same, hindering the clinical application of molecular typing systems. Therefore, it is urgent to establish a consensus molecular subtype system within existing molecular typing systems. [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] This invention provides a random forest classifier used for molecular typing of esophageal squamous cell carcinoma, the classifier incorporates a total of 314 characteristic genes, each represented by the following Ensembl database numbers: ENSG00000177272, ENSG00000122224, ENSG00000078589, ENSG00000132185, ENSG00000105369, ENSG00000177455, ENSG00000147138, ENSG00000007312, ENSG00000172724, EN SG00000107317, ENSG00000121895, ENSG00000012124, ENSG00000163564, ENSG00000117091, ENSG00000156738, ENSG00000117215, ENSG00000161405 , ENSG00000169442, ENSG00000113088, ENSG00000137078, ENSG00000254838, ENSG00000175463, ENSG00000162739, ENSG00000164691, ENSG00000185 905, ENSG00000182162, ENSG00000048462, ENSG00000100079, ENSG00000163534, ENSG00000166501, ENSG00000128218, ENSG00000132465, ENSG00000 005844, ENSG00000160654, ENSG00000079263, ENSG00000143297, ENSG00000113263, ENSG00000173200, ENSG00000168918, ENSG00000124406, ENSG00 000136573, ENSG00000183918, ENSG00000115085, ENSG00000122188, ENSG00000122122, ENSG00000081237, ENSG00000117090, ENSG00000110777, ENS G00000137265, ENSG00000026751, ENSG00000115165, ENSG00000167984, ENSG00000110448, ENSG00000035403, ENSG00000162777, ENSG00000167077,ENSG00000122550、ENSG00000185811、ENSG00000100351、ENSG00000106948、ENSG00000167286、ENSG00000134242、ENSG00000143851、ENSG00000110934、ENSG00000073792、ENSG00000153283、ENSG00000247774、ENSG00000198821、ENSG00000117322、ENSG00000010671、ENSG00000182866、ENSG00000196684、ENSG00000128815、ENSG00000269404、ENSG00000167208、ENSG00000172673、ENSG00000076662、ENSG00000124256、ENSG00000170476、ENSG00000197943、ENSG00000181847、ENSG00000009790、ENSG00000124772、ENSG00000099958、ENSG00000072858、ENSG00000153885、ENSG00000172578、ENSG00000057019、ENSG00000182963、ENSG00000108405、ENSG00000198851、ENSG00000145649、ENSG00000172215、ENSG00000164283、ENSG00000179583、ENSG00000185862、ENSG00000187498、ENSG00000013725、ENSG00000176105、ENSG00000095585、ENSG00000105122、ENSG00000164938、ENSG00000047457、ENSG00000042980、ENSG00000073849、ENSG00000101082、ENSG00000173905、ENSG00000146676、ENSG00000153563、ENSG00000105374、ENSG00000106263、ENSG00000187678、ENSG00000163577、ENSG00000168421、ENSG00000143067、ENSG00000166428、ENSG00000111711、ENSG00000205744、ENSG00000171388、ENSG00000173198、ENSG00000101333、ENSG00000147689、ENSG00000117308、ENSG00000168398、ENSG00000163220、ENSG00000173156、ENSG00000162946、ENSG00000165272、ENSG00000070018、ENSG00000213160、ENSG00000114790、ENSG00000153048、ENSG00000105355、ENSG00000162366、ENSG00000175591、ENSG00000143546、ENSG00000242808、ENSG00000005884、ENSG00000162409、ENSG00000157379、ENSG00000060558、ENSG00000111424、ENSG00000148344、ENSG00000142669、ENSG00000189334、ENSG00000178038、ENSG00000153294、ENSG00000163596、ENSG00000147697、ENSG00000110047、ENSG00000030582、ENSG00000146267、ENSG00000144749、ENSG00000186847、ENSG00000073910、ENSG00000115008、ENSG00000164520、ENSG00000092820、ENSG00000158315、ENSG00000142627、ENSG00000159348、ENSG00000115009、ENSG00000062038、ENSG00000171236、ENSG00000158552、ENSG00000124145、ENSG00000124107、ENSG00000134070、ENSG00000088726、ENSG00000177191、ENSG00000176788、ENSG00000187583、ENSG00000146072、ENSG00000128596、ENSG00000011009、ENSG00000116871、ENSG00000099992、ENSG00000182612、ENSG00000100097、ENSG00000150093、ENSG00000133816、ENSG00000078098、ENSG00000206538、ENSG00000166922、ENSG00000145431、ENSG00000196923、ENSG00000115414、ENSG00000167123、ENSG00000072110、ENSG00000165617、ENSG00000122786、ENSG00000072682、ENSG00000161638、ENSG00000122884、ENSG00000120708、ENSG00000087116、ENSG00000164294、ENSG00000070404、ENSG00000128805、ENSG00000122870、ENSG00000163430、ENSG00000142173、ENSG00000120149、ENSG00000198910、ENSG00000113721、ENSG00000133110、ENSG00000080573、ENSG00000157227、ENSG00000162745、ENSG00000105472、ENSG00000117385、ENSG00000113140、ENSG00000164574、ENSG00000151135、ENSG00000142552、ENSG00000143502、ENSG00000164111、ENSG00000149257、ENSG00000166250、ENSG00000177570、ENSG00000143226、ENSG00000184012、ENSG00000107819、ENSG00000172638、ENSG00000120820、ENSG00000083444、ENSG00000142156、ENSG00000108602、ENSG00000163297、ENSG00000149380、ENSG00000130635、ENSG00000152049、ENSG00000112769、ENSG00000100196、ENSG00000115363、ENSG00000142192、ENSG00000101825、ENSG00000110723、ENSG00000106333、ENSG00000079150、ENSG00000163520、ENSG00000111817、ENSG00000128833、ENSG00000166130、ENSG00000105974、ENSG00000026025、ENSG00000138316、ENSG00000151617、ENSG00000140682、ENSG00000221968、ENSG00000131459、ENSG00000140416、ENSG00000164465、ENSG00000218336、ENSG00000168542、ENSG00000154096、ENSG00000134013、ENSG00000182667、ENSG00000142227、ENSG00000145817、ENSG00000151388、ENSG00000127249、ENSG00000165915、ENSG00000196968、ENSG00000164692、ENSG00000087303、ENSG00000184828、ENSG00000143387、ENSG00000177119、ENSG00000091136、ENSG00000134668、ENSG00000128595、ENSG00000183876、ENSG00000124225、ENSG00000196344、ENSG00000158186、ENSG00000137573、ENSG00000258947、ENSG00000088836、ENSG00000131236、ENSG00000171401、ENSG00000163359、ENSG00000129009、ENSG00000123130、ENSG00000174705、ENSG00000153561、ENSG00000150630、ENSG00000170558、ENSG00000163209、ENSG00000106624、ENSG00000182492、ENSG00000168487、ENSG00000145113、ENSG00000126458、ENSG00000135127、ENSG00000149557、ENSG00000139211、ENSG00000016602、ENSG00000129514、ENSG00000035862、ENSG00000113083、ENSG00000114353、ENSG00000140279、ENSG00000167397、ENSG00000162576, ENSG00000116132, ENSG00000106397, ENSG00000039560, ENSG00000006327, ENSG00000214711, ENSG00000136694, , ENSG00000185873, ENSG00000189377, ENSG00000183844.
[0005] The present invention further provides a molecular subtype classifier for esophageal squamous cell carcinoma constructed using the random forest classifier described above, and the resulting consensus molecular subtypes are ECMS1, ECMS2, ECMS3, and ECMS4, with the following characteristics of the four molecular subtypes: ECMS1: metabolic pathway abnormalities, NFE2L2 activation, and selection of drugs used as NFE2L2 inhibitors; ECMS2: elevated classical signaling pathways in tumors, low degree of methylation; ECMS3: fewer copy number mutation events, low tumor mutation burden, high PD-1 expression, and benefit from immunosuppressant therapy; ECMS4: activation of the epithelial-stromal transformation pathway, low overall survival and disease-free survival, and benefit from epithelial-stromal transformation and VEFG inhibitor therapy.
[0006] The present invention further provides a method for constructing a molecular subtype classifier for esophageal squamous cell carcinoma, the method being: To obtain whole-genome DNA, transcriptome, and small RNA from esophageal squamous cell carcinoma, perform whole-genome sequencing, preprocess the high-throughput sequencing data, and obtain pathological omics data for esophageal squamous cell carcinoma, Based on the obtained omics data, namely copy number, transcriptome, small RNA, and methylation data, the omics data will be fused to perform subtype identification and identify molecular subtypes of esophageal squamous cell carcinoma. To collect typing tags from public datasets of esophageal squamous cell carcinoma, perform consensus molecular typing analysis, and obtain consensus molecular typing subtypes, Applying hypergeometric tests, Fisher's tests, and t-tests to identify subtype specific molecular features, Based on differential gene analysis and gene set enrichment analysis (GSEA), the relationship between each subtype and immunity, metabolism, classical tumor pathways, and epithelial-stromal transformation will be identified, respectively. We will perform immunomicroenvironmental analysis using R packet deconvolution and ESTIMATE, and survival analysis using Kaplan-Meier. This includes identifying molecular subtypes and characteristics of esophageal squamous cell carcinoma based on the analysis results, and constructing a molecular subtype classifier for esophageal squamous cell carcinoma.
[0007] Furthermore, preprocessing high-throughput sequencing data specifically involves the following: For whole-genome DNA sequencing, FASTP software is used to perform quality control filtering on the original sequencing data, filtering out low-quality reads, removing sequencing linker sequences, and removing relatively short reads. BWA software is used to match the sequencing reads with the human reference genome hg38 to obtain SAM files. Samtools is used to convert the SAM files to BAM files. Sambamba is used to sort the BAM files by coordinate position. The reads are repeated using the MarkDuplicates flag of the Picard toolkit, and the cnv_facets tool is used to detect somatic copy number variations (SCNVs) in tumor samples and paired normal samples. For transcriptome sequencing, FASTP software is used to perform quality control filtering on the original sequencing data, filtering out low-quality reads, removing sequencing linker sequences, removing relatively short reads, and correcting some bases based on paired read overlap regions. The quality-controlled reads are then matched against the human reference genome hg38 sequence using the STAR tool, and a data BAM file is formed. Gene expression levels are estimated using the RSEM tool, referencing annotations according to Ensembl version 98, using only matched reads. Gene expression values are then standardized as TPM. For whole-genome bisulfite sequencing reads, quality control filtering is first performed using FASTP software, and then matching is performed using the Bismark tool. The resulting BAM file is then processed to remove duplicate reads, and methylation signals are extracted using the Bismark methylation extractor. For small RNA sequencing data, the linker sequence is first removed using FASTP software, the short sequences are compared with the small RNA database miRBase and the human reference genome hg38 sequence using miRDeep2, and the expression levels of each small RNA are obtained using the miRDeep2 quantitative module.
[0008] Regarding the aforementioned copy count data, after extracting the log ratio value of the data from the cnv_facets result, redundant copy count areas are removed based on the CNregions function in iClusterPlus of the R packet. Regarding the methylation data, the methylation levels of all sites in all samples are integrated into a unified data matrix using the R packet methylKit. The methylation levels of the 2kb upstream and 0.5kb downstream of the gene transcription start site are calculated. The specific calculation method involves dividing the number of methylated sequences measured at all CpG sites within the gene promoter region by the total number of sequences. If 20% of the samples have missing values or low coverage in the promoter region, the data within that promoter region is not considered. Gene expression levels are first standardized as TPM (total productivity per million reads), and gene expression levels for all samples are imported using the R packet tximport to obtain a data matrix related to gene expression. Small RNA expression is standardized to CPM (number of small RNAs per million reads), and the small RNA expression values for all samples are read to obtain a small RNA data matrix. An example of the data format for each omics is as follows: rows represent data for a specific omics, columns represent specific samples, a total of n samples, and a total of four omics data matrices. If the matrix is a copy number data matrix, the features are copy number variant regions, the data are log ratio values obtained by matching tumor samples with normal samples in those regions, and the number of features is the number of regions after redundancy has been removed. If the matrix is a gene expression matrix, then the features are genes, the data are gene expression levels (specifically log2(TPM+1)), and the number of features is the number of genes. If the matrix is a methylation data matrix, the features are promoter regions, the data are the methylation levels of those promoter regions, the range is 0 to 1, and the number of features is the number of available promoters. If the matrix is a small RNA expression matrix, the features are the expression levels of the small RNAs, specifically log2(CPM+1), and the number of features is the total number of small RNAs. Further filtering was performed based on the absolute deviation of the median value to select the top 1000 genes with the most mutations, the top 500 small RNAs, the top 1000 promoters, and all copy number regions after removing redundancy. For each omics data, the Z-value was standardized, the Euclidean distance between two samples was calculated at the single-omics data level, a single-omics level similarity network was constructed using R-packet SNF, and network fusion was performed under the conditions of hyperparameters alpha = 0.75, K = 30, and T = 20.
[0009] Based on the analysis results, the method for identifying molecular subtypes and characteristics of esophageal squamous cell carcinoma is as follows: A spectral clustering algorithm is used to obtain molecular typing tags for each sample. To evaluate the contribution of each omics to the subtype, edges with the top 5% similarity weights are selected. If the difference in variable weights of multiple data types is less than 10%, it is considered that the edge between two samples in the network is contributed by multiple omics data. If the edge with the highest weight is greater than the edge with the next highest weight, it is considered that the evidence for the edge between two samples in the network comes from a specific single omics. A similarity network fusion method based on multi-omics data is then used to obtain four multi-omics molecular subtypes: MESCC1, MESCC2, MESCC3, and MESCC4.
[0010] Furthermore, typing tags from seven published typing studies (PubMed IDs: 29127303, 32398863, 32580137, 31048097, 28052061, 28052061, 36584672) were collected to construct consensus molecular typing input data. 85% of the samples are randomly selected, and hypergeometric tests are performed on the correlation between each typing system. A p-value of less than 0.05 indicates a correlation between subtypes. A correlation network of 26 interconnected nodes (each node representing a molecular subtype from a specific typing study) is constructed. The Markov clustering algorithm (MCL) is applied to identify clusters in the network. This process is repeated 100 times, and the frequency with which molecular subtypes from different typing systems are classified into a unified cluster is statistically calculated to construct a consensus frequency matrix (26*26). Further identification using MCL yields consensus molecular typing ECMS1, ECMS2, ECMS3, and ECMS4. Hypergeometric tests are performed to identify core ECMS samples (p-value less than 0.2), and a classifier is trained to predict the consensus molecular subtype for samples with a p-value greater than 0.2, yielding the consensus molecular subtype for all samples.
[0011] This invention utilizes high-throughput sequencing to sequence whole-genome DNA, transcriptome, and small RNA from tumor tissues of 152 patients with esophageal squamous cell carcinoma, and to sequence whole-genome methylation. Extensive preprocessing of the high-throughput data is performed using quality control software, sequence matching and processing software, mutation detection software, and transcriptome and small RNA quantification software to obtain omics data for esophageal squamous cell carcinoma patients. Based on a similarity network fusion method, four types of omics data are fused to perform subtype identification, identifying four multi-omics molecular subtypes (MESCC1-4) of esophageal squamous cell carcinoma. Furthermore, four consensus molecular subtypes (ECMS1-4) are identified through consensus molecular typing studies. Molecular features of subtype specificity are identified using hypergeometric tests, Fisher's tests, and t-tests. Based on differential gene analysis and gene set enrichment analysis (GSEA), the relationships between each subtype and immunity, metabolism, classical tumor pathways, and epithelial-stromal transformation are identified. Subtype ECMS1 enriched NFE2L2 amplification events. NFE2L2 expression levels in MESCC2 were higher than in the other three subtypes, indicating activation of NFE2L2 in MESCC2. ECMS3 had fewer copy number events and a relatively low tumor mutation burden (TMB). KMT2D mutations were less common in patients with the ECMS3 subtype. The ECMS3 subtype is characterized by higher immunoactivity, including myeloid suppressor cells (MDSCs), natural killer cells (NKCs), inflammatory helper T cells (Th17 cells), and the complement system. ECMS1 is characterized by dysfunction of various metabolic pathways, including cytochrome, glutathione, sucrose, glycoside, and uronic acid metabolic pathways. According to GSEA, the characteristics of ECMS2 were revealed to be upregulation of classical cell cycle-related pathways (WNT, MYC signaling pathway, KEGG cell cycle pathway). The ECMS4 subtype exhibits mesenchymal characteristics, characterized by TGFβ substrate rearrangement, the VEGF pathway, and epithelial-stromal transposition (EMT), and is called the EMT subtype. The EMT subtype shows significantly higher expression of epithelial-stromal related genes compared to the other three subtypes.Furthermore, the GSEA results based on proteomics data were similar to the characteristics of GSEA molecules based on transcriptional sequencing data, revealing that the four consensus molecular subtypes share the same molecular pattern at the protein level.
[0012] Immunomicroenvironmental analysis was performed using R-packet deconvolution and ESTIMATE. ESTIMATE results showed that the immune score of ECMS3 was significantly higher than the other three subtypes, and the matrix score of ECM4 was significantly higher than that of ECMS1 and ECMS2. Low tumor purity was observed in ECMS3 and ECMS4, suggesting potential interaction between the tumor and the progressing tumor microenvironment. TIMER immunoinfiltration analysis also showed relatively high immunoinfiltration content in MESCC1, with components such as B cells, T cells CD4+, and T cells CD8+ being present in ECMS3. Tumor-associated macrophages can promote EMT in ESCCs, and ECMS4 and ECMS3 showed higher macrophage infiltration rates compared to the classical subtype ECMS2.
[0013] Survival analysis using Kaplan-Meier was performed, and the results of the survival analysis of the four subtypes showed that patients with the MESCC4 subtype exhibited the worst overall survival (OS) and disease-free survival (DFS). This invention constructs a molecular typing classifier based on the characteristic genes and gene expression levels of the four subtypes at the transcriptome level. Applying this classifier, the clinical relevance of the four subtypes was verified, and based on the analysis of clinical trial data, the ECMS3 subtype was significantly correlated with the group that benefited from immune checkpoint inhibitor therapy, and the expression level of PDCD1 (PD-1) was significantly higher than that of the other three groups, indicating that patients with ECMS3 may benefit from immune checkpoint inhibitor therapy. Furthermore, the association analysis of drug sensitivity and gene expression showed that the data were obtained from CTRP cell line data (823 cell lines). Negative correlations indicated potential drug sensitivity. VER-155008 showed a significant negative correlation with NFE2L2 expression, indicating that necrosulfonamide is a potential treatment policy targeting NFE2L2 in ECMS1 patients. Furthermore, cerulenin and PF-3758309 may be used to target MYC and VEGFC in ECMS2 and ECMS4.
[0014] This invention provides a multi-omics molecular subtype classification method and system for esophageal squamous cell carcinoma to analyze its heterogeneity, as well as molecular and clinical characteristics and classifiers for each subtype. It also provides a consensus molecular typing classification method and system for esophageal squamous cell carcinoma, and aims to unify the typing systems to make molecular typing more widely adopted. [Brief explanation of the drawing]
[0015] [Figure 1] This shows the origins and distributions supporting the omics similarity evidence for each subtype of the MESCC typing system. [Figure 2] The consensus frequency matrix is the frequency at which molecular subtypes are divided into the same cluster. [Figure 3] It is an esophageal squamous cell carcinoma molecular subtype network based on eight typing systems. [Figure 4] It is the molecular characteristics of the consensus molecular subtype. In the figure: a is the spectrum of subtype mutations and copy number characteristics, b is the statistical distribution of copy number variations between subtypes, c is the distribution of tumor mutation burden between subtypes, d is the KMT2D mutation distribution, e is the distribution of BST2 methylation levels, f is the NFE2L2 expression distribution, and e is the distribution of related pathway enrichment scores. [Figure 5] It is the mutation frequency of genes in different datasets. [Figure 6] The distribution differences of epithelial-mesenchymal transition (EMT), VEGF, and cell cycle-related genes are shown in a heatmap. In the figure: a shows the distribution difference of EMT-related genes in a heatmap, b shows the distribution difference of VEGF-related genes in a heatmap, and c shows the distribution difference of cell cycle-related genes in a heatmap. [Figure 7] It is the distribution of the immune microenvironment in each consensus molecular subtype. In the figure: a-i are the immune score, stromal score, tumor purity, B cells, CD4+ T cells, CD8+ T cells, neutrophils, macrophages, and myeloid dendritic cells, respectively. [[ID=...]] [Figure 8] It is the correlation analysis curve between the subtype tags and survival periods of samples in GSE53625. Here, a shows the correlation analysis curve between the subtype tags and overall survival period, and b shows the correlation analysis curve between the subtype tags and disease-free survival period. [Figure 9] It is the overall survival of the consensus molecular subtype in the dataset GSE53625 and the disease-free survival period in the dataset TCGA. In the figure: a is the overall survival period of the consensus molecular subtype in the dataset GSE53625, and b is the disease-free survival period of the consensus molecular subtype in the dataset TCGA. [Figure 10] In it, a-d are the distributions of sensitivities of the chemotherapeutic drugs Cisplatin, Vinorelbine, Paclitaxel, and Docetaxel in each consensus molecular subtype, respectively. [Figure 11] The correlation between the drug sensitivity and gene expression of cell lines. In the figure, a is the correlation distribution between related drugs and small molecule sensitivity and the expression level of NFE2L2, b is the distribution of VER-155008 sensitivity and the expression level of NFE2L2, c is the correlation distribution between related drugs and small molecule sensitivity and the expression level of MYC, d is the distribution of cerulenin sensitivity and the expression level of MYC, e is the correlation distribution between related drugs and small molecule sensitivity and the expression level of VEGCF, and f is the distribution of PF-3758309 sensitivity and the expression level of VEGFC. [Figure 12] The clinical treatment response in each consensus molecular subtype. In the figure: a is the GSE45670 cohort; b is the internal cohort; pCR: pathological complete remission; <pCR: less than pCR; PR: partial response; SD: disease stability; PD: disease progression.
Embodiments for Carrying out the Invention
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. It is obvious that the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the present invention.
[0017] Unless otherwise defined, all technical terms and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art. The citations disclosed in this specification and the materials they cited are incorporated by reference.
[0018] Equivalent technologies of the specific embodiments described that can be understood by those skilled in the art through ordinary experiments are included in this application.
[0019] Unless otherwise specified, the experimental methods in the following examples are conventional methods. Unless otherwise specified, the equipment used in the following examples is standard laboratory equipment. Unless otherwise specified, the experimental materials used in the following examples are those that can be purchased at a regular store.
[0020] 1. Data preprocessing: For whole-genome DNA sequencing, FASTP software was used to perform quality control filtering on the original sequencing data, filtering out low-quality reads, removing sequencing linker sequences, and removing relatively short reads. BWA software was used to match the sequencing reads to the human reference genome hg38 to obtain SAM files. Samtools was used to convert the SAM files to BAM files, and Sambamba was used to sort the BAM files by coordinate position. Reads were repeated using the MarkDuplicates flag in the Picard toolkit. The GATK tool and a known human variant database were used to correct base mass and minimize system errors by the sequencer. Mutations were identified against normal samples using Mutect2 single-sample patterns, and PON files were created. Mutation patterns were then detected using Mutect2 normal-tumor tissue pairing, and somatic mutations (point mutations and insertion / deletion mutations) were detected in each sample. Further filtering of mutations was performed based on sequence context and sample contamination. Somatic mutations were annotated using ANNOVAR software to obtain relevant information such as the region where the mutation occurred, the gene, its frequency in public databases, and its impact on amino acids. Subsequent analysis filtered for mutations located in non-genetic regions, mutations with an ExAC_EAS frequency > 0.0.5, or mutations with synonymous mutations. Somatic copy number variations (SCNVs) were detected in normal samples paired with tumor samples using the cnv_facets tool. Mutations and SCNVs were analyzed using the R packet "mafTools". In this example, GISTIC was used to identify significantly amplified or deleted genomic regions.
[0021] For transcriptome sequencing, FASTP software is used to perform quality control filtering on the original sequencing data, filtering out low-quality reads, removing sequencing linker sequences, removing relatively short reads, and correcting some bases based on paired read overlap regions. After quality control, the reads are matched against the human reference genome hg38 sequence using the STAR tool to form a data BAM file. Gene expression levels are estimated using the RSEM tool, referring to annotations according to Ensembl (version 98), using only matched reads, and gene expression values are standardized as TPM.
[0022] For whole-genome bisulfite sequencing reads, quality control filtering is first performed using FASTP software, followed by matching using the Bismark tool. The resulting BAM file is then processed to remove duplicate reads, and methylation signals are extracted using the Bismark methylation extractor.
[0023] For small RNA sequencing data, the linker sequence is first removed using FASTP software, the short sequences are compared with the small RNA database miRBase and the human reference genome hg38 sequence using miRDeep2, and the expression levels of each small RNA are obtained using the miRDeep2 quantitative module.
[0024] 2. Identification of molecular subtypes by a method based on the fusion of multi-omics data and similarity networks: The present invention identifies molecular subtypes based on four types of omics data: copy number, transcriptome, small RNA, and methylation data.
[0025] Here, the log ratio value of the copy count data is extracted from the cnv_facets result, and then redundant copy count regions are removed based on the CNregions function in iClusterPlus of the R packet.
[0026] For methylation data, the R packet methylKit is used to integrate the methylation levels of all sites in all samples into a unified data matrix. The methylation level within the gene promoter region (2kb upstream and 0.5kb downstream of the gene transcription start site) is then calculated. The specific calculation method involves dividing the number of methylated sequences measured at all CpG sites within the gene promoter region by the total number of sequences. If 20% of the samples have missing values or low coverage in the promoter region, the data within that region is not considered.
[0027] Gene expression data is first standardized as the number of transcripts per million reads (TPM), and the gene expression levels of all samples are imported using the R packet tximport to obtain a data matrix related to gene expression. Small RNA expression is standardized to the number of small RNAs per million reads (CPM). The small RNA expression values for all samples are read to obtain a small RNA data matrix, and before data processing, log2 processing is first performed on TPM and CPM. An example of the data format for each omics is a matrix where rows are data for a specific omics, columns are specific samples, for a total of n samples, and for a total of four omics data matrices. If it is a copy number data matrix, the features are copy number variant regions, the data are log ratio values obtained by matching tumor samples and normal samples in that region, and the number of features is the number of regions after redundancy is removed. If the matrix is a gene expression matrix, the features are genes, the data are gene expression levels, specifically log2(TPM+1), and the number of features is the number of genes. If it is a methylation data matrix, the features are promoter regions, the data are the methylation level of the promoter region, the range is 0 to 1, and the number of features is the number of available promoters. If it is a small RNA expression matrix, the features are small RNA expression levels, specifically log2(CPM+1), and the number of features is the total number of small RNAs.
[0028] To further screen for biologically significant features, we perform further filtering based on the absolute deviation of the median value to select the top 1000 most mutated genes, top 500 small RNAs, top 1000 promoters, and all copy number regions after removing redundancy. For each omics data, we standardize the Z-value and calculate the Euclidean distance between two samples at the single-omics data level. We construct a single-omics level similarity network using R-packet SNF and perform network fusion under the hyperparameters alpha = 0.75, K = 30, and T = 20. We obtain molecular typing tags for each sample using a spectral clustering algorithm. To assess the contribution of each omics to a subtype, we select edges with similarity weights in the top 5%. If the difference in variable weights across multiple data types is less than 10%, it is considered that the edge between two samples in the network is contributed by multiple omics data. If the edge with the highest weight is greater than the edge with the next highest weight, it is considered that the evidence for the edge between two samples in the network comes from a specific single omics. Four molecular subtypes (MESCC1-4) were obtained using a similarity network fusion method based on multi-omics data.
[0029] Furthermore, typing tags from seven published typing studies (PubMed IDs: 29127303, 32398863, 32580137, 31048097, 28052061, 28052061, 36584672) were collected to construct consensus molecular typing input data. 85% of the samples are randomly selected, and hypergeometric tests are performed on the correlation between each typing system. A p-value of less than 0.05 indicates a correlation between subtypes. A correlation network of 26 interconnected nodes (each node representing a molecular subtype from a specific typing study) is constructed. The Markov clustering algorithm (MCL) is applied to identify clusters in the network. This process is repeated 100 times, and the frequency with which molecular subtypes from different typing systems are classified into a unified cluster is statistically calculated to construct a consensus frequency matrix (26*26). Further identification using MCL yields consensus molecular typing ECMS1, ECMS2, ECMS3, and ECMS4. Hypergeometric tests are performed to identify core ECMS samples (p-value less than 0.2), and a classifier is trained to predict the consensus molecular subtype for samples with a p-value greater than 0.2, yielding the consensus molecular subtype for all samples.
[0030] III. Characterization of each molecular subtype: Differential gene analysis is performed between a specific subtype and three other subtypes using the R packet LIMMA to obtain differentially expressed genes between each molecular subtype and the other three subtypes. Then, gene set enrichment analysis is performed using the R packet HTSanalyzeR2 with bootstrap set to 1000 iterations. Kaplan-Meier analysis of survival time is performed using the R packet survival. The tumor microenvironment of each subtype, including B cells, CD4+, CD8+, neutrophils, macrophages, and dendritic cells, is analyzed using the TIMER method with the R packet deconvolution, and the immune score and substrate score of each subtype are further evaluated using the ESTIMATE tool. Molecular characteristics of subtype specificity are identified using hypergeometric tests, Fisher's tests, and t-tests.
[0031] IV. Construction of a subtype classifier: Through difference analysis, the top 50 genes with up-adjusted multiplicative differences for each subtype are selected to represent the characteristic genes of that subtype. Since the characteristic genes need to be present in multiple datasets simultaneously, crossover analysis with genes from other datasets is necessary to determine the final characteristic genes.
[0032] The expression values of characteristic genes are extracted, z-score standardization is performed, and then, based on the Random Forest algorithm, the correctly predicted samples are selected as core samples based on the accuracy of the Random Forest model's predictions. The Random Forest classifier is then rebuilt, and the number of Random Forest classifier trees is set to 500. This classifier can be used to predict molecular subtypes of esophageal squamous cell carcinoma in other transcriptome datasets.
[0033] The classifier incorporates 314 characteristic genes (indicated by Ensembl numbers): ENSG00000177272, ENSG00000122224, ENSG00000078589, ENSG00000132185, ENSG00000105369, ENSG00000177455, ENSG00000147138, ENSG00000007312, ENSG00000172724, ENSG00000107317, ENSG00000121895, ENSG00000012124, ENSG0000016 3564, ENSG00000117091, ENSG00000156738, ENSG00000117215, ENSG00000161405, ENSG00000169442, ENSG00000113088, ENSG00000137078, ENSG00 000254838, ENSG00000175463, ENSG00000162739, ENSG00000164691, ENSG00000185905, ENSG00000182162, ENSG00000048462, ENSG00000100079, EN SG00000163534, ENSG00000166501, ENSG00000128218, ENSG00000132465, ENSG00000005844, ENSG00000160654, ENSG00000079263, ENSG000001432 97, ENSG00000113263, ENSG00000173200, ENSG00000168918, ENSG00000124406, ENSG00000136573, ENSG00000183918, ENSG00000115085, ENSG00000 122188, ENSG00000122122, ENSG00000081237, ENSG00000117090, ENSG00000110777, ENSG00000137265, ENSG00000026751, ENSG00000115165, ENSG 00000167984, ENSG00000110448, ENSG00000035403, ENSG00000162777, ENSG00000167077, ENSG00000122550, ENSG00000185811, ENSG00000100351,ENSG00000106948、ENSG00000167286、ENSG00000134242、ENSG00000143851、ENSG00000110934、ENSG00000073792、ENSG00000153283、ENSG00000247774、ENSG00000198821、ENSG00000117322、ENSG00000010671、ENSG00000182866、ENSG00000196684、ENSG00000128815、ENSG00000269404、ENSG00000167208、ENSG00000172673、ENSG00000076662、ENSG00000124256、ENSG00000170476、ENSG00000197943、ENSG00000181847、ENSG00000009790、ENSG00000124772、ENSG00000099958、ENSG00000072858、ENSG00000153885、ENSG00000172578、ENSG00000057019、ENSG00000182963、ENSG00000108405、ENSG00000198851、ENSG00000145649、ENSG00000172215、ENSG00000164283、ENSG00000179583、ENSG00000185862、ENSG00000187498、ENSG00000013725、ENSG00000176105、ENSG00000095585、ENSG00000105122、ENSG00000164938、ENSG00000047457、ENSG00000042980、ENSG00000073849、ENSG00000101082、ENSG00000173905、ENSG00000146676、ENSG00000153563、ENSG00000105374、ENSG00000106263、ENSG00000187678、ENSG00000163577、ENSG00000168421、ENSG00000143067、ENSG00000166428、ENSG00000111711、ENSG00000205744、ENSG00000171388、ENSG00000173198、ENSG00000101333、ENSG00000147689、ENSG00000117308、ENSG00000168398、ENSG00000163220、ENSG00000173156、ENSG00000162946、ENSG00000165272、ENSG00000070018、ENSG00000213160、ENSG00000114790、ENSG00000153048、ENSG00000105355、ENSG00000162366、ENSG00000175591、ENSG00000143546、ENSG00000242808、ENSG00000005884、ENSG00000162409、ENSG00000157379、ENSG00000060558、ENSG00000111424、ENSG00000148344、ENSG00000142669、ENSG00000189334、ENSG00000178038、ENSG00000153294、ENSG00000163596、ENSG00000147697、ENSG00000110047、ENSG00000030582、ENSG00000146267、ENSG00000144749、ENSG00000186847、ENSG00000073910、ENSG00000115008、ENSG00000164520、ENSG00000092820、ENSG00000158315、ENSG00000142627、ENSG00000159348、ENSG00000115009、ENSG00000062038、ENSG00000171236、ENSG00000158552、ENSG00000124145、ENSG00000124107、ENSG00000134070、ENSG00000088726、ENSG00000177191、ENSG00000176788、ENSG00000187583、ENSG00000146072、ENSG00000128596、ENSG00000011009、ENSG00000116871、ENSG00000099992、ENSG00000182612、ENSG00000100097、ENSG00000150093、ENSG00000133816、ENSG00000078098、ENSG00000206538、ENSG00000166922、ENSG00000145431、ENSG00000196923、ENSG00000115414、ENSG00000167123、ENSG00000072110、ENSG00000165617、ENSG00000122786、ENSG00000072682、ENSG00000161638、ENSG00000122884、ENSG00000120708、ENSG00000087116、ENSG00000164294、ENSG00000070404、ENSG00000128805、ENSG00000122870、ENSG00000163430、ENSG00000142173、ENSG00000120149、ENSG00000198910、ENSG00000113721、ENSG00000133110、ENSG00000080573、ENSG00000157227、ENSG00000162745、ENSG00000105472、ENSG00000117385、ENSG00000113140、ENSG00000164574、ENSG00000151135、ENSG00000142552、ENSG00000143502、ENSG00000164111、ENSG00000149257、ENSG00000166250、ENSG00000177570、ENSG00000143226、ENSG00000184012、ENSG00000107819、ENSG00000172638、ENSG00000120820、ENSG00000083444、ENSG00000142156、ENSG00000108602、ENSG00000163297、ENSG00000149380、ENSG00000130635、ENSG00000152049、ENSG00000112769、ENSG00000100196、ENSG00000115363、ENSG00000142192、ENSG00000101825、ENSG00000110723、ENSG00000106333、ENSG00000079150、ENSG00000163520、ENSG00000111817、ENSG00000128833、ENSG00000166130、ENSG00000105974、ENSG00000026025、ENSG00000138316、ENSG00000151617、ENSG00000140682、ENSG00000221968、ENSG00000131459、ENSG00000140416、ENSG00000164465、ENSG00000218336、ENSG00000168542、ENSG00000154096、ENSG00000134013、ENSG00000182667、ENSG00000142227、ENSG00000145817、ENSG00000151388、ENSG00000127249、ENSG00000165915、ENSG00000196968、ENSG00000164692、ENSG00000087303、ENSG00000184828、ENSG00000143387、ENSG00000177119、ENSG00000091136、ENSG00000134668、ENSG00000128595、ENSG00000183876、ENSG00000124225、ENSG00000196344、ENSG00000158186、ENSG00000137573、ENSG00000258947、ENSG00000088836、ENSG00000131236、ENSG00000171401、ENSG00000163359、ENSG00000129009、ENSG00000123130、ENSG00000174705、ENSG00000153561、ENSG00000150630、ENSG00000170558、ENSG00000163209、ENSG00000106624、ENSG00000182492、ENSG00000168487、ENSG00000145113、ENSG00000126458、ENSG00000135127、ENSG00000149557、ENSG00000139211、ENSG00000016602、ENSG00000129514、ENSG00000035862、ENSG00000113083、ENSG00000114353、ENSG00000140279、ENSG00000167397、ENSG00000162576、ENSG00000116132、ENSG00000106397、ENSG00000039560, ENSG00000006327, ENSG00000214711, ENSG00000136694, ENSG00000185873, ENSG00000189377, ENSG00000183844. ,
[0034] V. Correlation analysis between drug sensitivity and treatment response: Based on the oncoPredict packet, the sensitivity of a sample to a specific chemotherapy drug is predicted. For subtype-specific genes, correlation analysis is performed between gene expression levels and drug sensitivity using drug sensitivity data and gene expression data from the CTRP2 database, and drug molecules with a significant negative correlation are identified as potential candidate drugs.
[0035] Based on the constructed random forest classifier, preoperative radiotherapy and esophageal squamous cell carcinoma subtypes in the common dataset GSE45670 are predicted.
[0036] Based on copy number data from patients receiving autoimmune suppressants, redundant copy number regions are removed using the CNregions function in iClusterPlus of the R packet to obtain a copy number data frame of the sample, with the value being the log ratio. A random forest classifier for predicting immune response is constructed, with 500 trees in the classifier. This can be used to predict the molecular subtype of esophageal squamous cell carcinoma samples with copy number data and to further study the correlation between treatment response and subtype.
[0037] VI. Experimental results: 1. Multi-omics data similarity network fusion identifies four subtypes: Based on the similarity network fusion method, four subtypes are identified, and the similarity evidence for each subtype comes from single-omics or multi-omics levels. The origin and distribution of the similarity evidence are shown in Figure 1.
[0038] 2. Study of Consensus Molecular Subtypes: Based on the consensus molecular typing identification method described in the present invention, a 26*26 consensus frequency matrix is obtained, as shown in Figure 2, where the rows and columns represent molecular subtypes of the typing system, and the color represents the frequency at which two molecular subtypes are divided into the same cluster in 100 85% sampling repeated analyses. Based on the consensus frequency matrix (network), Markov clustering identifies four subtypes (ECMS1-4), and the type of subtype included in each consensus molecular typing is as shown in Figure 3, where each node represents the molecular subtype of the corresponding typing system, its size is proportional to the sample size, and the width of the side represents the consensus frequency between two subtypes.
[0039] To obtain the consensus molecular subtype for each sample, a hypergeometric test is used to analyze the significance of each sample belonging to each molecular subtype. Samples with a P-value less than 0.2 are designated as core samples, and the consensus molecular subtype tags of samples with a P-value of 0.2 or greater are predicted. Finally, the consensus molecular subtype for all samples is obtained.
[0040] 3. Subtype Molecular and Clinical Characteristics: As shown in Figure 4a, similar to previous studies, the gene with the highest frequency of mutations and copy number variations was TP53 (91%), followed by CDKN2A (84%), NFE2L2 (48%), KMT2C (39%), KMT2D (30%), NOTCH1 (18%), CSMD3 (16%), PIK3CA (14%), and LRP1B (12%). Among these, CDKN2A and FAT1 contained many copy number deletion events, and NFE2L2 contained many copy number amplification times.
[0041] ECMS1 enriches NFE2L2 copy number validation events (P<0.01, 22 / 32 vs 44 / 120, chi-squared test), and simultaneously, NFE2L2 is expressed at relatively high levels within ECMS1 (Figure 4f).
[0042] ECMS3 tumors are characterized by a low number of copy number mutation events (Figure 4b) and a relatively low tumor mutational burden (TMB) (Figure 4c), which is consistent with the CNV2 subtype defined in Du et al. Non-synonymous mutations in KMT2D are associated with poor prognosis, and the mutation frequency of this gene is not the same across the four consensus molecule subtypes, with ECMS3 subtype tumors having fewer KMT2D mutations (Figure 4d). Mutation frequencies in the TCGA dataset, the dataset according to the present invention (SXM-I), and other Chinese esophageal squamous cell carcinoma datasets (SXM-II) are compared (Figure 5). TP53 showed a high frequency in Chinese individuals, but there was no significant difference between datasets (P=0.22). Compared to the TCGA-ESCC queue, CDKN2A (chi-squared test, P=0.0019) and FAT1 (P=0.046, chi-squared test) had even higher mutation frequencies in Chinese individuals. Conversely, the mutation frequency of NEF2L2 is low (P=0.045, chi-squared test).
[0043] Tumors of the ECMS1 subtype are characterized by dysfunction of multiple metabolic pathway subtypes, including cytochrome, glutathione, sucrose, glycoside, and uronic acid metabolic pathways (Figure 4g). NFE2L2 has previously been shown to be involved in metabolic reprogramming and may partially explain the phenotype of metabolic dysfunction in ECMS1.
[0044] ECMS2 is characterized by dysfunction of classical tumor pathways such as WNT, MYC, and the cell cycle (Figure 4g). ECMS3 is associated with immune activation and includes high expression of myeloid suppressor cells (MDSCs), mature keratinocytes (NKCs), inflammatory adjuvant T cells (Th17 cells), and PD-1 (Figure 4g). The ECMS4 subtype exhibited stromal filling characteristics characterized by TGFβ activation, the VEGF pathway, and epithelial-stromal transformation (EMT) (Figure 4g). This invention consistently found that core characteristic genes of the EMT and VEGF pathways were significantly more highly expressed in ECMS4 compared to the other three subtypes, and that core characteristic genes of the cell cycle were overexpressed in ECMS2 (Figure 6c).
[0045] Using ESTIMATE, immune and substrate scores were calculated based on deconvolution of tumor transcriptome data, consistent with GSEA results. Tumors of the ECMS3 subtype showed higher immune scores (Figure 7a), and tumors of the ECMS4 subtype had higher substrate content (Figure 7b). The tumor purity of ECMS3 and ECMS4 was lower (Figure 7c), suggesting that these tumors may be more susceptible to the influence of the tumor microenvironment. Further analysis using TIMER also showed higher immune infiltration of B cells, CD4+, and CD8+ T cells in the ECMS3 subtype (Figure 7d-i). Similar molecular characteristics were observed in gene set enrichment analysis based on proteomics data from 90 esophageal squamous cell carcinomas.
[0046] Survival analysis of four subtypes revealed that patients with the ECMS4 subtype exhibited the worst overall survival and disease-free survival. To verify this correlation, the present invention constructed a subtype classifier based on random forests, predicted the subtype tags of samples in the independent dataset GSE53625, and performed clinical correlation analysis, finding a significant correlation with overall survival (Figure 8a) and a significant correlation with disease-free survival in the TCGA dataset (Figure 8b).
[0047] 4. Drug sensitivity analysis: Based on the publicly available cell line drug database CTRP data and the oncoPredict package, the sensitivity of tumor samples to specific chemotherapeutic drugs is predicted. As a result of the analysis, it was found that ECMS3-type tumors are resistant to the treatment of Cisplatin and Docetaxel, while ECMS1-type tumors are resistant to the treatment of Paclitaxel (Figure 10). The present invention discovered that NFE2L2, MYC, PDCD1 (PD-1), and VEGFC are significantly upregulated in the ECMS1-4 subtypes respectively. Similarly, based on the data of esophageal cancer cell lines in the CTRP database, through the correlation analysis between drug sensitivity and gene expression, it was revealed that the negative correlation indicates the possibility of drug sensitivity. VER-155008 shows a significant negative correlation with NFE2L2 expression, indicating that Necrosulfonamide is a potential treatment policy targeting NFE2L2 in ECMS1 patients (Figure 11a). Furthermore, cerulenin and PF-3758309 may be used to target MYC and VEGFC in ECMS2 and ECMS4 (Figure 11b and Figure 11c).
[0048] Based on the data analysis of clinical trials, there was no significant correlation between the ECMS3 subtype and the complete pathologic response (<pCR) to the new neoadjuvant chemoradiotherapy (Figure 12a), which was shown to be consistent with the results obtained from the computational analysis. However, the expression level of PDCD1 (PD-1) in ECMS3 was significantly higher than that in the other three groups (Figure 4g). The ECMS3 subtype was significantly correlated with the treatment benefit group in immune checkpoint inhibitor therapy (Figure 12b), indicating that patients with ECMS3 may benefit from immune checkpoint inhibitor therapy.
[0049] Finally, while the above embodiments have been used solely to illustrate the technical concepts of the present invention and not to limit them, those skilled in the art should understand that they may modify the technical concepts described in the above embodiments or equally replace some or all of their technical features, and that such modifications or replacements will not cause the essence of the corresponding technical concepts to deviate from the scope of the technical concepts of the embodiments of the present invention.
Claims
1. A random forest classifier used for molecular typing of esophageal squamous cell carcinoma, the classifier incorporates a total of 314 characteristic genes, each represented by the Ensembl database number as follows: ENSG00000177272, ENSG00000122224, ENSG00000078589, ENSG00000132185, ENSG00000105369, ENSG00000177455, ENSG00000147138, ENSG00000007312, ENSG00000172724, ENSG0 0000107317, ENSG00000121895, ENSG00000012124, ENSG00000163564, ENS G00000117091, ENSG00000156738, ENSG00000117215, ENSG00000161405, E NSG00000169442, ENSG00000113088, ENSG00000137078, ENSG00000254838 , ENSG00000175463, ENSG00000162739, ENSG00000164691, ENSG0000018590 5, ENSG00000182162, ENSG00000048462, ENSG00000100079, ENSG00000163 534, ENSG00000166501, ENSG00000128218, ENSG00000132465, ENSG000000 05844, ENSG00000160654, ENSG00000079263, ENSG00000143297, ENSG0000 0113263, ENSG00000173200, ENSG00000168918, ENSG00000124406, ENSG000 00136573, ENSG00000183918, ENSG00000115085, ENSG00000122188, ENSG0 0000122122, ENSG00000081237, ENSG00000117090, ENSG00000110777, ENS G00000137265, ENSG00000026751, ENSG00000115165, ENSG00000167984, E NSG00000110448, ENSG00000035403, ENSG00000162777, ENSG00000167077,ENSG00000122550、ENSG00000185811、ENSG00000100351、ENSG00000106948、ENSG00000167286、ENSG00000134242、ENSG00000143851、ENSG00000110934、ENSG00000073792、ENSG00000153283、ENSG00000247774、ENSG00000198821、ENSG00000117322、ENSG00000010671、ENSG00000182866、ENSG00000196684、ENSG00000128815、ENSG00000269404、ENSG00000167208、ENSG00000172673、ENSG00000076662、ENSG00000124256、ENSG00000170476、ENSG00000197943、ENSG00000181847、ENSG00000009790、ENSG00000124772、ENSG00000099958、ENSG00000072858、ENSG00000153885、ENSG00000172578、ENSG00000057019、ENSG00000182963、ENSG00000108405、ENSG00000198851、ENSG00000145649、ENSG00000172215、ENSG00000164283、ENSG00000179583、ENSG00000185862、ENSG00000187498、ENSG00000013725、ENSG00000176105、ENSG00000095585、ENSG00000105122、ENSG00000164938、ENSG00000047457、ENSG00000042980、ENSG00000073849、ENSG00000101082、ENSG00000173905、ENSG00000146676、ENSG00000153563、ENSG00000105374、ENSG00000106263、ENSG00000187678、ENSG00000163577、ENSG00000168421、ENSG00000143067、ENSG00000166428、ENSG00000111711、ENSG00000205744、ENSG00000171388、ENSG00000173198、ENSG00000101333、ENSG00000147689、ENSG00000117308、ENSG00000168398、ENSG00000163220、ENSG00000173156、ENSG00000162946、ENSG00000165272、ENSG00000070018、ENSG00000213160、ENSG00000114790、ENSG00000153048、ENSG00000105355、ENSG00000162366、ENSG00000175591、ENSG00000143546、ENSG00000242808、ENSG00000005884、ENSG00000162409、ENSG00000157379、ENSG00000060558、ENSG00000111424、ENSG00000148344、ENSG00000142669、ENSG00000189334、ENSG00000178038、ENSG00000153294、ENSG00000163596、ENSG00000147697、ENSG00000110047、ENSG00000030582、ENSG00000146267、ENSG00000144749、ENSG00000186847、ENSG00000073910、ENSG00000115008、ENSG00000164520、ENSG00000092820、ENSG00000158315、ENSG00000142627、ENSG00000159348、ENSG00000115009、ENSG00000062038、ENSG00000171236、ENSG00000158552、ENSG00000124145、ENSG00000124107、ENSG00000134070、ENSG00000088726、ENSG00000177191、ENSG00000176788、ENSG00000187583、ENSG00000146072、ENSG00000128596、ENSG00000011009、ENSG00000116871、ENSG00000099992、ENSG00000182612、ENSG00000100097、ENSG00000150093、ENSG00000133816、ENSG00000078098、ENSG00000206538、ENSG00000166922、ENSG00000145431、ENSG00000196923、ENSG00000115414、ENSG00000167123、ENSG00000072110、ENSG00000165617、ENSG00000122786、ENSG00000072682、ENSG00000161638、ENSG00000122884、ENSG00000120708、ENSG00000087116、ENSG00000164294、ENSG00000070404、ENSG00000128805、ENSG00000122870、ENSG00000163430、ENSG00000142173、ENSG00000120149、ENSG00000198910、ENSG00000113721、ENSG00000133110、ENSG00000080573、ENSG00000157227、ENSG00000162745、ENSG00000105472、ENSG00000117385、ENSG00000113140、ENSG00000164574、ENSG00000151135、ENSG00000142552、ENSG00000143502、ENSG00000164111、ENSG00000149257、ENSG00000166250、ENSG00000177570、ENSG00000143226、ENSG00000184012、ENSG00000107819、ENSG00000172638、ENSG00000120820、ENSG00000083444、ENSG00000142156、ENSG00000108602、ENSG00000163297、ENSG00000149380、ENSG00000130635、ENSG00000152049、ENSG00000112769、ENSG00000100196、ENSG00000115363、ENSG00000142192、ENSG00000101825、ENSG00000110723、ENSG00000106333、ENSG00000079150、ENSG00000163520、ENSG00000111817、ENSG00000128833、ENSG00000166130、ENSG00000105974、ENSG00000026025、ENSG00000138316、ENSG00000151617、ENSG00000140682、ENSG00000221968、ENSG00000131459、ENSG00000140416、ENSG00000164465、ENSG00000218336、ENSG00000168542、ENSG00000154096、ENSG00000134013、ENSG00000182667、ENSG00000142227、ENSG00000145817、ENSG00000151388、ENSG00000127249、ENSG00000165915、ENSG00000196968、ENSG00000164692、ENSG00000087303、ENSG00000184828、ENSG00000143387、ENSG00000177119、ENSG00000091136、ENSG00000134668、ENSG00000128595、ENSG00000183876、ENSG00000124225、ENSG00000196344、ENSG00000158186、ENSG00000137573、ENSG00000258947、ENSG00000088836、ENSG00000131236、ENSG00000171401、ENSG00000163359、ENSG00000129009、ENSG00000123130、ENSG00000174705、ENSG00000153561、ENSG00000150630、ENSG00000170558、ENSG00000163209、ENSG00000106624、ENSG00000182492、ENSG00000168487、ENSG00000145113、ENSG00000126458、ENSG00000135127、ENSG00000149557、ENSG00000139211、ENSG00000016602、ENSG00000129514、ENSG00000035862、ENSG00000113083、ENSG00000114353、ENSG00000140279、ENSG00000167397、ENSG00000162576、ENSG00000116132、ENSG00000106397、ENSG00000039560、ENSG00000006327、ENSG00000214711、ENSG00000136694、ENSG00000185873、ENSG00000189377、ENSG00000183、 A random forest classifier characterized by 844.
2. The molecular subtype classifier for esophageal squamous cell carcinoma is constructed based on the random forest classifier described in claim 1, and the resulting molecular subtypes are ECMS1, ECMS2, ECMS3, and ECMS4, characterized in that ECMS1 is characterized by metabolic pathway abnormalities, NFE2L2 activation, and selection of drugs used as NFE2L2 inhibitors, ECMS2 is characterized by elevated classical signaling pathways in the tumor and a low degree of methylation, ECMS3 is characterized by fewer copy number mutation events, low tumor mutation burden, high PD-1 expression, and benefit from immunosuppressant therapy, and ECMS4 is characterized by activation of the epithelial-stromal transformation pathway, low overall survival and disease-free survival, and benefit from epithelial-stromal transformation and VEFG inhibitor therapy.
3. A method for constructing a molecular subtype classifier for esophageal squamous cell carcinoma, wherein a random forest classifier is used to divide the characteristic genes of esophageal squamous cell carcinoma into four molecular subtypes, ECMS1, ECMS2, ECMS3, and ECMS4, based on pathological omics data of esophageal squamous cell carcinoma, where the characteristics of ECMS1 are metabolic pathway abnormalities, NFE2L2 activation, and the selection of drugs used as NFE2L2 inhibitors; the characteristics of ECMS2 are elevated classical signaling pathways in the tumor and a low degree of methylation; the characteristics of ECMS3 are few copy number mutation events, low tumor mutation burden, high PD-1 expression, and benefit from immunosuppressant therapy; and the characteristics of ECMS4 are activation of the epithelial-stromal transformation pathway, low overall survival and disease-free survival, and benefit from epithelial-stromal transformation and VEFG inhibitor therapy. The method for constructing the molecular subtype classifier for esophageal squamous cell carcinoma is as follows: To obtain whole-genome DNA, transcriptome, and small RNA from esophageal squamous cell carcinoma, perform whole-genome sequencing, preprocess the high-throughput sequencing data, and obtain pathological omics data for esophageal squamous cell carcinoma, Based on the obtained omics data, namely copy number data, transcriptome data, small RNA and methylation data, the omics data will be fused to perform subtype identification and identify molecular subtypes of esophageal squamous cell carcinoma. To collect typing tags from public datasets of esophageal squamous cell carcinoma, perform consensus molecular typing analysis, and obtain consensus molecular typing subtypes, To identify subtype specific molecular features by applying hypergeometric tests, Fisher's test, and t-tests, Based on differential gene analysis and gene set enrichment analysis (GSEA), the relationship between each subtype and immunity, metabolism, classical tumor pathways, and epithelial-stromal transformation will be identified, respectively. Immunomicroenvironmental analysis will be performed using R-packet deconvolution and ESTIMATE, and survival analysis will be performed using Kaplan-Meier. A construction method comprising: identifying molecular subtypes and characteristics of esophageal squamous cell carcinoma based on the analysis results, and constructing a molecular subtype classifier for esophageal squamous cell carcinoma.
4. Preprocessing high-throughput sequencing data specifically involves the following: For whole-genome DNA sequencing, FASTP software is used to perform quality control filtering on the original sequencing data, filtering out low-quality reads, removing sequencing linker sequences, and removing relatively short reads. BWA software is used to match the sequencing reads with the human reference genome hg38 to obtain SAM files. Samtools is used to convert the SAM files to BAM files. Sambamba is used to sort the BAM files by coordinate position. The MarkDuplicates flag of the Picard toolkit is used to repeat the reads, and the cnv_facets tool is used to detect somatic cell copy number mutations (SCNVs) in tumor samples and paired normal samples. For transcriptome sequencing, FASTP software is used to perform quality control filtering on the original sequencing data, filtering out low-quality reads, removing sequencing linker sequences, removing relatively short reads, and correcting some bases based on paired read overlap regions. After quality control, the reads are matched against the human reference genome hg38 sequence using the STAR tool to form a data BAM file. Gene expression levels are estimated using the RSEM tool, referring to annotations according to Ensembl version 98, using only matched reads, and gene expression values are standardized as TPM. For whole-genome sequencing reads, first, quality control filtering is performed using FASTP software, and then matching is performed using the Bismark tool. The resulting BAM file is then processed to remove duplicate reads, and methylation signals are extracted using the Bismark methylation extractor. The construction method according to claim 3, characterized in that, for small RNA sequencing data, linker sequences are first removed using FASTP software, short sequences are compared with the small RNA database miRBase and the human reference genome hg38 sequence using miRDeep2, and the expression levels of each small RNA are obtained using the quantitative module of miRDeep2.
5. Regarding the aforementioned copy count data, after extracting the log ratio value of the data from the cnv_facets result, the redundant copy count area is removed based on the CNregiments function in iClusterPlus of the R packet. The construction method according to claim 3, characterized in that, regarding methylation data, the methylation levels of all sites in all samples are integrated into a unified data matrix using R packet methylKit, and the methylation levels of the 2kb upstream and 0.5kb downstream of the gene transcription start site are calculated, the specific calculation method being to divide the number of methylated sequences measured at all CpG sites in the gene promoter region of the sample by the total number of sequences, and if 20% of the sample has missing values or low coverage in the promoter region, the data within the promoter region is not considered.
6. The gene expression values are first standardized as the number of transcripts per million reads (TPM), and the gene expression levels of all samples are imported using the R packet tximport to obtain a data matrix related to gene expression. The expression of small RNAs is standardized to CPM (number of small RNAs per million reads), the small RNA expression values for all samples are read, and a small RNA data matrix is obtained. If the matrix is a copy number data matrix, the features are copy number variant regions, the data are log ratio values obtained by matching tumor samples with normal samples in those regions, and the number of features is the number of regions after redundancy has been removed. If the matrix is a gene expression matrix, then the features are genes, the data are gene expression levels (specifically log2(TPM+1)), and the number of features is the number of genes. If the matrix is a methylation data matrix, the features are promoter regions, the data are the methylation levels of those promoter regions, the range is 0 to 1, and the number of features is the number of available promoters. If the matrix is a small RNA expression matrix, the features are the expression levels of the small RNAs, specifically log2(CPM+1), and the number of features is the total number of small RNAs. The construction method according to claim 4, characterized by further filtering based on the absolute deviation of the median value to select the top 1000 genes with the most mutations, the top 500 small RNAs, the top 1000 promoters, and all copy number regions after redundancy has been removed, standardizing the Z value for each omics data, calculating the Euclidean distance between two samples at the single omics data level, constructing a single omics-level similarity network using R-packet SNF, and performing network fusion under the conditions of hyperparameters alpha = 0.75, K = 30, and T = 20.
7. Based on the analysis results, the method for identifying molecular subtypes and characteristics of esophageal squamous cell carcinoma is as follows: Using a spectral clustering algorithm, molecular typing tags are obtained for each sample, and to evaluate the contribution to each omics subtype, edges with the top 5% similarity weights are selected. If the difference in variable weights across multiple data types is less than 10%, it is considered that the edge between two samples in the network is contributed by multiple omics data. If the edge with the highest weight is greater than the edge with the next highest weight, it is considered that the evidence for the edge between two samples in the network comes from a specific single omics. A similarity network fusion method based on multi-omics data is used to obtain four multi-omics molecular typing subtypes: MESCC1, MESCC2, MESCC3, and MESCC4. The construction method according to claim 4, further characterized by collecting typing tags from seven published typing studies PubMed IDs: 29127303, 32398863, 32580137, 31048097, 28052061, 28052061, and 36584672, constructing consensus molecular typing input data, randomly selecting 85% of the samples, performing hypergeometric tests on the correlation between each typing system, applying a Markov clustering algorithm to identify clusters in the network, repeating this 100 times, statistically determining the frequency with which molecular subtypes of different typing systems are classified into a unified cluster, constructing a 26*26 consensus frequency matrix, and further identifying consensus molecular typings ECMS1, ECMS2, ECMS3, and ECMS4 using a Markov clustering algorithm.