An esophageal squamous carcinoma molecular typing marker based on proteomics and non-targeted metabolomics combination and application thereof

By combining DIA proteomics and untargeted metabolomics, a molecular subtyping model for esophageal squamous cell carcinoma was constructed, which solved the problem of lack of metabolomics features in existing subtyping methods, realized accurate subtyping and prognostic assessment of esophageal squamous cell carcinoma patients, and provided a new method for diagnosis, treatment and prognostic evaluation.

CN117089616BActive Publication Date: 2026-04-28SHANXI MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANXI MEDICAL UNIV
Filing Date
2023-08-07
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing molecular subtyping methods for esophageal squamous cell carcinoma lack metabolomics subtyping features. There is no molecular subtyping method for esophageal squamous cell carcinoma based on a combination of proteomics and metabolomics, which makes it difficult to reveal the metabolic heterogeneity of the microenvironment and its regulatory mechanisms, thus affecting precision treatment and prognostic assessment.

Method used

Using a combination of DIA proteomics and untargeted metabolomics, a classifier model was constructed by detecting four proteins (hexokinase 3, Sec1 family domain protein 1, cytoscleroprotein 5, and signal sequence receptor α subunit) and two metabolites (creatine and 2-deoxy-D-glucose) to classify esophageal squamous cell carcinoma into types I, II, and III. Similarity network fusion was used for subtyping, and a classifier model was constructed for molecular subtyping.

Benefits of technology

It enables precise classification of esophageal squamous cell carcinoma patients, with type I having the best prognosis and type III having the worst prognosis, providing new methods for diagnosis, treatment, and prognostic evaluation, and significantly distinguishing patients' overall survival and disease-free survival.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117089616B_ABST
    Figure CN117089616B_ABST
Patent Text Reader

Abstract

The application belongs to the field of biomedical technology, and aims to solve the problems of lack of metabolic typing characteristics in the current ESCC molecular typing method, lack of esophageal squamous cell carcinoma molecular typing method based on proteomics and metabolomics, and provides an esophageal squamous carcinoma ESCC molecular typing marker based on DIA proteomics and non-targeted metabolomics and its application. The marker is a combination of 4 proteins and 2 metabolites, the 4 proteins are: hexokinase 3 (HK3), Sec1 family domain protein 1 (SCFD1), septin 5 (SEPTIN5), and signal sequence receptor alpha subunit (SSR1); the 2 metabolites are creatine and 2-deoxy-D-glucose (2'DG). The application fills the gap of the molecular typing method based on proteomics and metabolomics in the field of esophageal squamous carcinoma research, and the proposed molecular typing model can be used as a method for diagnosis, treatment and prognosis evaluation of esophageal squamous carcinoma patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical technology, specifically relating to a molecular subtyping biomarker for esophageal squamous cell carcinoma (ESCC) based on a combination of DIA proteomics and non-targeted metabolomics, and its application. Background Technology

[0002] Esophageal cancer is one of the most common malignant tumors of the digestive system. Globally, there are more than 600,000 new cases of esophageal cancer each year, and about 540,000 deaths, ranking sixth in mortality among all cancers [1]. Compared with other malignant tumors, the targeted therapy drugs for ESCC are currently very limited. In recent years, immunotherapy, represented by PD-1 / PD-L1 inhibitors, has played a significant role in the treatment of ESCC, bringing new opportunities for ESCC. However, for second-line treatment, immunotherapy can only achieve an objective response rate of about 20%. For example, pembrolizumab did not have an overall survival benefit compared with paclitaxel second-line treatment in patients with PD-L1 CPS≥1, but showed better safety. For first-line treatment, the effective response time was only maintained for 7 months. The objective response rate of neoadjuvant immunotherapy was not significantly improved compared with neoadjuvant chemoradiotherapy [4-6].

[0003] Given the high heterogeneity of ESCC and its microenvironment, molecular subtyping in omics research is an effective method for analyzing the heterogeneity of tumors and their microenvironment. Studying the molecular characteristics of different subtypes is of great significance for understanding the molecular mechanisms of tumor development, discovering new molecular changes and metabolic characteristics with potential therapeutic targets, and helping to achieve precision medicine.

[0004] Currently, there are several molecular subtypes of ESCC based on different omics levels internationally. For example, TCGA classifies esophageal cancer into three subtypes based on genomic and methylmic characteristics. In subtype 1, NFE2L2 mutations are associated with poor prognosis and tolerance to radiotherapy and chemotherapy. Patients in subtype 2 have a higher mutation rate in ZNF750 or NOTCH1 genes, while those in subtype 3 have a higher level of DNA methylation [7]. However, this study has a small sample size and does not include Chinese patients, so it is unclear whether its subtypes are applicable to Chinese patients with esophageal squamous cell carcinoma. Cui et al. classified patients into three subtypes based on whole-genome sequencing of 508 Chinese ESCC patients: NFE2L2 mutant, RTK-RAS-MYC copy number amplification, and double-negative [8]. Li et al. classified ESCC into two subtypes based on proteomics and phosphorylation modification omics. In subtype 2, spliceosome and ribosomal proteins are upregulated, and three potential drugs may be effective for patients in subtype 2 [9]. In a recent study, Liu et al. combined whole-genome, epigenome, transcriptome and proteomics data to classify ESCC patients into four subtypes: immunomodulatory (IM), immunosuppressive (IS), cell cycle activated (CCA) and NRF2 activated (NRFA), and formulated treatment strategies for different subtypes of esophageal squamous cell carcinoma

[10] . Although the above molecular subtyping studies have revealed the heterogeneity of ESCC and its microenvironment from multiple perspectives, they have mostly focused on the level of genomic variation and tumorigenesis mechanism, and it is difficult to establish a direct correspondence between them and tumor phenotypes. On the other hand, due to the lack of metabolomics subtyping features, the above studies have not revealed the metabolic heterogeneity and regulatory mechanisms in the microenvironment.

[0005] Metabolomics is a method for studying biological systems at the metabolite level. It mainly studies the level and distribution of metabolites (such as glucose, amino acids, lipids, etc.) in organisms and their association with physiological and pathological processes, revealing the metabolic characteristics of tissues, cells, and organisms and the interaction between metabolites. Metabolomics has broad application prospects in tumor research. It can provide important information for early diagnosis, treatment guidance and prognostic assessment of tumors, identify cancer biomarkers and driving factors of tumor development, and is one of the important components of tumor research and precision medicine

[11] .

[0006] Tumor development and progression require metabolic reprogramming of tumor cells. Tumor cells autonomously change their flux through various metabolic pathways (such as carbohydrate, lipid, and amino acid metabolism) to meet the increased bioenergy and biosynthetic needs and reduce the oxidative stress required for tumor cell proliferation and survival. Changes in tumor cell metabolism can also cause changes in metabolites in the tumor microenvironment

[12] . Changes in the metabolic characteristics of the tumor microenvironment can usually reshape the immune response within the tumor and affect the function of immune cells. Some targeted cancer metabolism drugs have been shown to help tumor immunotherapy

[13] . For example, CB-839, an inhibitor of glutaminase-1 (GLS1), enhances the function of immune cells by blocking the glutamine metabolic pathway in cancer cells. The use of this drug as a single agent and in combination with PD-1 inhibitors has entered clinical trials for various tumors such as renal cell carcinoma and breast cancer

[14] . High levels of indole-2,3 dioxygenase (IDO) in cancer cells can reduce the availability of tryptophan in the tumor microenvironment and inhibit the tumor-killing function of CD8+ T cells; it can also break down tryptophan into kynurenine, inducing immunosuppression of dendritic and regulatory T cells; multiple early phase I / II clinical trials have also confirmed that small molecule IDO inhibitors can improve the efficacy of immunotherapy for malignant tumors such as melanoma [15,16].

[0007] Proteins and metabolites are both downstream of life activities and are closely related. As the carriers of life activities, proteins can indicate the life activities that will occur in the body, while metabolites indicate the life activities that have already occurred in the body and reflect the microenvironment in which the cell is located. Compared with the gene level, changes in proteins and metabolites can better reflect phenotypic characteristics. At present, there are some studies combining proteomics and metabolomics internationally: Geiger et al., through joint analysis of proteomics and metabolomics, revealed that high levels of intracellular L-arginine can affect the activation, differentiation, and survival of T cells, thereby enhancing the body's anti-tumor ability, and discovered three L-arginine-sensitive transcription factors that are closely related to the survival of T cells

[17] . Wozniak et al. analyzed serum samples from 294 patients with Staphylococcus aureus bacteremia (SAB), revealed early prediction and pathogenicity characteristics, and established pathogenic characteristics and multivariate models that can accurately predict the mortality of SAB patients, providing a powerful tool for guiding personalized treatment of SAB patients

[18] . Ringel et al. analyzed tumor tissues from high-fat diet (HFD) mice, mapped the overall metabolic changes in tumor immune infiltration, and discovered the characteristics of fatty acid uptake and oxidation by HFD tumor cells. They revealed that obesity damages the number of CD8+ T cells and anti-tumor activity in the tumor microenvironment of mice, thereby accelerating tumor growth

[19] . The above studies have demonstrated the feasibility of combined proteomics and metabolomics analysis. However, there are currently no reports on the combined proteomics and metabolomics typing of ESCC. Therefore, combining proteomics and metabolomics in ESCC to study the changes in proteins and metabolites in the ESCC microenvironment, and performing molecular typing of ESCC based on changes in microenvironment proteins and metabolism, will help reveal the metabolic heterogeneity of the microenvironment and its regulatory mechanisms, thereby guiding clinical diagnosis and treatment, and achieving precise treatment and prognostic assessment for patients.

[0008] References:

[0009] 1. Sung, H. 等人 Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA:一份 临床肿瘤学杂志 71, 209-249, doi:10.3322 / caac.21660 (2021).

[0010] 2. Zheng, R. 等人 Cancer statistics in China, 2016].中华肿瘤 杂志[《中华肿瘤杂志》] 45, 212-220, doi:10.3760 / cma.j.cn112152-20220922-00647 (2023).

[0011] 3、Abnet, C., Arnold, M.&Wei, W. Epidemiology of Esophageal SquamousCell Carcinoma. 《胃肠病学》 154, 360-373, doi:10.1053 / j.gastro.2017.08.023(2018).

[0012] 4、Huang, J. 等人 Camrelizumab versus investigator's choice ofchemotherapy as second-line therapy for advanced or metastatic oesophagealsquamous cell carcinoma (ESCORT): a multicentre, randomised, open-label,phase 3 study. 《柳叶刀·肿瘤学》 21, 832-842, doi:10.1016 / s1470-2045(20)30110-8 (2020).

[0013] 5、Kato, K. 等人 KEYNOTE-590: Phase III study of first-linechemotherapy with or without pembrolizumab for advanced esophagealcancer. 《未来肿瘤学》(英国伦敦) 15, 1057-1066, doi:10.2217 / fon-2018-0609 (2019).

[0014] 6、Smyth, E., Gambardella, V., Cervantes, A.&Fleitas, T. Checkpointinhibitors for gastroesophageal cancers: dissecting heterogeneity to betterunderstand their role in first-line and adjuvant therapy. 《肿瘤学年鉴》: 欧洲医学肿瘤学会官方杂志 32, 590-599,doi:10.1016 / j.annonc.2021.02.004 (2021).

[0015] 7、Integrated genomic characterization of oesophagealcarcinoma. 《自然》 541, 169-175, doi:10.1038 / nature20805 (2017).

[0016] 8、Cui, Y. 等人 Whole-genome sequencing of 508 patients identifies keymolecular features associated with poor prognosis in esophageal squamous cellcarcinoma. 《细胞研究》 30, 902-913, doi:10.1038 / s41422-020-0333-6 (2020).

[0017] 9、Liu, W. 等人 Large-scale and high-resolution mass spectrometry-based proteomics profiling defines molecular subtypes of esophageal cancerfor therapeutic targeting. 《自然通讯》 12, 4961, doi:10.1038 / s41467-021-25202-5 (2021).

[0018] 10、Liu, Z. 等人Integrated multi-omics profiling yields a clinicallyrelevant molecular classification for esophageal squamous cellcarcinoma. 《癌细胞》 41, 181-195.e189, doi:10.1016 / j.ccell.2022.12.004(2023).

[0019] 11、Schmidt, D. 等人 Metabolomics in cancer research and emergingapplications in clinical oncology. CA:一份临床肿瘤学杂志 71, 333-358, doi:10.3322 / caac.21670 (2021).

[0020] 12、Martínez-Reyes, I.&Chandel, N. Cancer metabolism: lookingforward. 《自然综述·癌症》 21, 669-680, doi:10.1038 / s41568-021-00378-6(2021).

[0021] 13、Zhang, L. 等人 Genomic analyses reveal mutational signatures andfrequently altered genes in esophageal squamous cell carcinoma. 美国 《人类遗传学杂志》 96, 597-611, doi:10.1016 / j.ajhg.2015.02.017 (2015).

[0022] 14、Yang, W., Qiu, Y., Stamatatos, O., Janowitz, T.&Lukey, M.Enhancing the Efficacy of Glutamine Metabolism Inhibitors in CancerTherapy. 《癌症趋势》7, 790-804, doi:10.1016 / j.trecan.2021.04.003 (2021).

[0023] 15、Kraehenbuehl, L., Weng, C., Eghbali, S., Wolchok, J.&Merghoub, T.Enhancing immunotherapy in cancer by targeting emerging immunomodulatorypathways. 《自然综述·临床肿瘤学》 19, 37-50, doi:10.1038 / s41571-021-00552-7 (2022).

[0024] 16、Jung, K. 等人 Phase I Study of the Indoleamine 2,3-Dioxygenase 1(IDO1) Inhibitor Navoximod (GDC-0919) Administered with PD-L1 Inhibitor(Atezolizumab) in Advanced Solid Tumors. 《临床癌症研究》:美国癌症研究协会官方 杂志 25, 3220-3228, doi:10.1158 / 1078-0432.Ccr-18-2740 (2019).

[0025] 17、Geiger, R. 等人 L-Arginine Modulates T Cell Metabolism andEnhances Survival and Anti-tumor Activity. 《细胞》 167, 829-842.e813, doi:10.1016 / j.cell.2016.09.031 (2016).

[0026] 18、Wozniak, J. 等人 Mortality Risk Profiling of Staphylococcus aureusBacteremia by Multi-omic Serum Analysis Reveals Early Predictive andPathogenic Signatures. 《细胞》182, 1311-1327.e1314, doi:10.1016 / j.cell.2020.07.040 (2020).

[0027] 19. Ringel, A. 等人 Obesity Shapes Metabolism in the TumorMicroenvironment to Suppress Anti-Tumor Immunity. 《细胞》 183, 1848-1866.e1826,doi:10.1016 / j.cell.2020.11.009 (2020). Summary of the Invention

[0028] To address the lack of metabolomics-based typing features in current molecular typing methods for esophageal squamous cell carcinoma (ESCC) and the absence of a molecular typing method for ESCC based on a combination of proteomics and metabolomics, this invention provides a molecular typing biomarker for ESCC based on a combination of DIA proteomics and non-targeted metabolomics, and its application.

[0029] This invention is achieved by the following technical solution: a molecular subtyping marker for esophageal squamous cell carcinoma (ESCC) based on a combination of DIA proteomics and non-targeted metabolomics. The marker is a combination of 4 proteins and 2 metabolites. The 4 proteins are: hexokinase 3 (HK3), Sec1 family domain protein 1 (SCFD1), cytoschizoprotein 5 (SEPTIN5), and signal sequence receptor α subunit (SSR1). The 2 metabolites are creatine and 2-deoxy-D-glucose (2'DG).

[0030] The application of molecular subtyping biomarkers for esophageal squamous cell carcinoma (ESCC) based on a combination of DIA proteomics and non-targeted metabolomics, and the application of these biomarkers in esophageal squamous cell carcinoma subtyping.

[0031] The specific classification method is as follows:

[0032] (1) Collect tumor tissues and paired adjacent normal tissues for omics detection, and collect corresponding clinical information data to form an actual analysis cohort;

[0033] (2) DIA quantitative proteomics and non-targeted metabolomics detection were performed on the actual analysis cohort; the obtained data were subjected to quality control analysis, abnormal samples were removed, and the detection data of the remaining samples were re-integrated; the obtained data were standardized, and proteins and metabolites that were significantly different between tumor tissue and adjacent normal tissue were screened out for subsequent subtyping;

[0034] (3) Based on the expression levels of differentially expressed proteins and metabolites in tumor tissue obtained in step (2), patients with esophageal squamous cell carcinoma are classified into esophageal squamous cell carcinoma type I, esophageal squamous cell carcinoma type II and esophageal squamous cell carcinoma type III using the similarity network fusion method.

[0035] (4) Screen the characteristic proteins and metabolites of each subtype in step (3) and construct a classifier model that can distinguish between type I, type II and type III. The model consists of multiple proteins and metabolites.

[0036] The classifier model consists of a combination of four proteins and two metabolites. The four proteins are: hexokinase 3 (HK3), Sec1 family domain protein 1 (SCFD1), cytosolic protein 5 (SEPTIN5), and signal sequence receptor α subunit (SSR1). The two metabolites are creatine and 2-deoxy-D-glucose (2'DG). The formula for calculating the classifier model score is: score = -1.384 + 0.449 × exp(SCFD1) + 0.981 × exp(Creatine) + 0.362 × exp(SEPTIN5) - 0.608 × exp(HK3) - 0.111 × exp(SSR1) - 0.373 × exp(2'DG). When the score is ≤ 0, the sample is classified as esophageal squamous cell carcinoma type I or II; when the score is > 0, the sample is classified as esophageal squamous cell carcinoma type III.

[0037] Patients with type I esophageal squamous cell carcinoma have the longest overall survival and disease-free survival, and the best prognosis, while patients with type III esophageal squamous cell carcinoma have the shortest overall survival and disease-free survival, and the worst prognosis.

[0038] The clinical information includes: age, gender, tumor stage, degree of tumor invasion, lymph node metastasis, survival status, and overall survival.

[0039] This invention proposes a novel classification method for esophageal squamous cell carcinoma. Based on the different expression patterns of proteins and metabolites in the tumor tissues of different esophageal squamous cell carcinoma patients, the method performs molecular subtyping, classifying esophageal squamous cell carcinoma patients into type I, type II, and type III. Patients with type I esophageal squamous cell carcinoma have the longest overall survival and disease-free survival, and the best prognosis, while patients with type III esophageal squamous cell carcinoma have the shortest overall survival and disease-free survival, and the worst prognosis.

[0040] This invention fills the gap in the field of esophageal squamous cell carcinoma research by using a combination of proteomics and metabolomics for molecular typing. The proposed molecular typing model can be used as a method for the diagnosis, treatment and prognosis evaluation of esophageal squamous cell carcinoma patients. Attached Figure Description

[0041] Figure 1This is a clustering heatmap of the Similarity Network Fusion (SNF) network in Embodiment 1 of the present invention, showing the clustering situation when the number of clusters K is 2, 3, 4, 5, 6, and 7.

[0042] Figure 2 This is a heatmap of three esophageal squamous cell carcinoma subtypes clustered by SNF in Embodiment 1 of the present invention, where rows represent samples and columns represent proteins or metabolites;

[0043] Figure 3 In Example 1 of this invention, the three subtypes of esophageal squamous cell carcinoma were significantly correlated with overall survival (OS) and disease-free survival (DFS).

[0044] Figure 4 There were significant differences in the number of patients with lymph node metastasis among the three subtypes of esophageal squamous cell carcinoma.

[0045] Figure 5 The contribution of each of the six molecules in the classifier model to the model;

[0046] Figure 6 ROC curve for the predictive performance of the classifier model;

[0047] Figure 7 In cohort 2, the predicted types III, I, and II using the classifier model were also significantly correlated with overall survival (OS). Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains, and all materials publicly cited herein and cited by them are incorporated herein by reference.

[0050] Equivalent technologies of the specific embodiments described herein that are readily apparent to those skilled in the art through routine experimentation are included in this application.

[0051] Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the instruments and equipment used in the following examples are all conventional laboratory instruments and equipment; unless otherwise specified, the experimental materials used in the following examples were all purchased from conventional biochemical reagent stores.

[0052] I. Clinical Sample and Information Collection: This invention included two cohorts: Cohort 1 consisted of tumor tissues and adjacent normal control tissues collected from esophageal squamous cell carcinoma patients in Shanxi Province (126 cases in each cohort, all aliquoted into sterile, enzyme-free tubes after surgical excision and immediately placed in liquid nitrogen, then frozen at -80°C within 30 minutes of excision). Patients were actively followed up, and their clinical information was collected (including but not limited to: age, sex, tumor stage, degree of tumor invasion, lymph node metastasis, survival status, and overall survival). Cohort 2 consisted of tumor tissues and paraffin sections from 52 esophageal squamous cell carcinoma patients in Shanxi Province (no overlap with patients in Cohort 1), and clinical information was collected according to the standards of Cohort 1. All patients in both cohorts had primary esophageal squamous cell carcinoma.

[0053] II. Non-standard quantitative proteomics sequencing

[0054] 1. Protein extraction: The specific method is as follows:

[0055] (1) Weigh an appropriate amount of sample, transfer it into a 2mL centrifuge tube, add two steel balls, add an appropriate amount of 1xCocktail containing SDS L3 and EDTA at a final concentration, place it on ice for 5 minutes, and add 10mM DTT at a final concentration.

[0056] (2) The mixture was crushed and pyrolyzed using a grinder (frequency 60 Hz, time 2 min), centrifuged at 4 ℃ for 15 min at 25,000 g, and the supernatant was collected; DTT with a final concentration of 10 mM was added and the mixture was placed in a water bath at 56 ℃ for 1 h; IAM with a final concentration of 55 mM was added and the mixture was placed in a dark room for 45 min.

[0057] (3) Add cold acetone to the protein solution obtained in step (2) at a ratio of 1:5, place in a -20℃ refrigerator for 30 min, centrifuge at 25,000g and 4℃ for 15 min and discard the supernatant.

[0058] (4) Air dry the precipitate, add an appropriate amount of SDS-free L3, and use a grinder (frequency 60HZ, time 2min) to promote protein dissolution; centrifuge at 25,000g, 4℃ for 15min and take the supernatant, which is the protein solution.

[0059] 2. Protein extraction quality control:

[0060] (1) Bradford quantification: Add 0, 2, 4, 6, 8, 10, 12, 14, 16, and 18 μL of standard protein (0.2 μg / μL BSA) sequentially to positions A1 to A10 of a 96-well microplate. Then add 20, 18, 16, 14, 12, 10, 8, 6, 4, and 2 μL of pure water sequentially. Finally, add 180 μL of Coomassie Brilliant Blue G-250 quantitative working solution to each well. Measure OD595 using a microplate reader and construct a linear standard curve based on OD595 and protein concentration. Dilute the protein solution to be tested several times, add 180 μL of quantitative working solution to 20 μL of protein solution, and read OD595. Calculate the sample protein concentration based on the standard curve and the sample OD595.

[0061] (2) SDS-PAGE: Take 10 μg of protein solution for each sample, add an appropriate amount of loading buffer, mix well, heat at 95℃ for 5 min, centrifuge at 25,000g for 5 min, take the supernatant and spot it into the sample well of 12% SDS polyacrylamide gel, electrophoresis at 80V for 30 min and then electrophoresis at 120V for 120 min; after electrophoresis, put the gel into a rapid staining and destaining instrument for 10 min, and then take out the gel image for scanning.

[0062] 3. Protein hydrolysis: Take 100 μg of protein solution for each sample; add 2.5 μg of Trypsin enzyme at a protein:enzyme ratio of 40:1, and hydrolyze at 37℃ for 4 h; desalt the hydrolyzed peptides using a Strata X column and vacuum dry.

[0063] 4. High pH RP Separation: Equal amounts of peptides from all samples were mixed, diluted with mobile phase A (5% ACN, pH 9.8), and injected. A Shimadzu LC-20AD liquid chromatography system was used for sample separation with a 5μm 4.6x250mm Gemini C18 column. A gradient elution was performed at a flow rate of 1 mL / min: 5% mobile phase B (95% ACN, pH 9.8) for 10 min, 5% to 35% mobile phase B for 40 min, 35% to 95% mobile phase B for 1 min, mobile phase B for 3 min, and equilibration with 5% mobile phase B for 10 min. Elution peaks were monitored at 214 nm, and one fraction was collected every minute. The samples were combined using the chromatographic elution peaks to obtain 10 fractions, which were then freeze-dried.

[0064] 5. DDA library construction and DIA quantification (Nano-LC-MS / MS): The dried peptide sample was reconstituted with mobile phase A (2% ACN, 0.1% FA), centrifuged at 20,000g for 10 min, and the supernatant was injected. Separation was performed using a Thermo UltiMate 3000U HPLC. The sample was first enriched and desalted in a trap column, then connected in series with a self-packed C18 column (150 μm inner diameter, 1.8 μm column particle size, approximately 35 cm column length) and separated at a flow rate of 500 nL / min using the following effective gradient: 0–5 min, 5% mobile phase B (98% ACN, 0.1% FA); 5–130 min, mobile phase B linearly increased from 5% to 25%; 130–150 min, mobile phase B increased from 25% to 35%; 150–160 min, mobile phase B increased from 35% to 80%; 160–175 min, 80% mobile phase B; 175–175.5 min, mobile phase B decreased from 80% to 5%; 175.5–180 min, 5% mobile phase B. The nanoliter liquid chromatography separator was directly connected to a mass spectrometer and analyzed using the following parameters:

[0065] (1) DDA Library Preparation and Detection: Peptides separated by liquid chromatography were ionized using a nanoESI source and then detected in DDA (Data Dependent Acquisition) mode on a Fusion Lumos tandem mass spectrometer (Thermo Fisher Scientific, San Jose, CA). Key parameter settings: Ion source voltage was set to 2 kV; primary mass spectrometry scan range was 350–1,500 m / z; resolution was set to 60,000, and maximum ion implantation time (MIT) was 50 ms; secondary mass spectrometry fragmentation mode was HCD, fragmentation energy was set to 30; resolution was set to 15,000, maximum ion implantation time (MIT) was 50 ms, and dynamic exclusion time was set to 30 s. The initial m / z for secondary mass spectrometry was fixed at 100; the precursor ion selection criteria for secondary fragmentation were: ions with charges ranging from 2+ to 6+ and peak intensities exceeding 2E4, ranked in the top 30. AGC was set to: primary 1E5, secondary 2E4.

[0066] (2) DIA Mass Spectrometry Detection: The peptides separated by liquid chromatography were ionized using a nanoESI source and then detected in DIA (Data Independent Acquisition) mode on a Fusion Lumos tandem mass spectrometer (Thermo Fisher Scientific, San Jose, CA). Key parameter settings: Ion source voltage was set to 2 kV; primary mass spectrometry scan range was 400–1,500 m / z; resolution was set to 60,000; maximum ion implantation time (MIT) was 50 ms; the 400–1,500 m / z range was divided into 44 windows for continuous window fragmentation and signal acquisition. The ion fragmentation mode was HCD, the maximum ion implantation time (MIT) was 54 ms, fragment ions were detected in Orbitrap with a resolution of 30,000 and a fragmentation energy of 30; AGC was set to 5E4.

[0067] III. Non-targeted metabolomics sequencing

[0068] 1. Metabolite Extraction: After slowly thawing the sample at 4℃, 100 μL was placed in a 96-well plate, and 300 μL of extraction buffer (methanol:ACN = 2:1, v:v, pre-cooled at -20℃) + 10 μL of internal standard were added. The mixture was vortexed for 1 min, incubated at -20℃ for 2 h, and then centrifuged at 4℃, 4000g for 20 min. After centrifugation, 300 μL of the supernatant was collected, dried in a refrigerated vacuum concentrator, and then reconstituted with 150 μL of reconstitution solution (methanol:H2O = 1:1, v:v). The mixture was vortexed for 1 min and then centrifuged at 4℃, 4000 rpm. -1 Centrifuge for 30 min, and transfer the supernatant to a sample vial. Take 10 μL of the supernatant from each sample and mix them to form a QC control sample, which is used to evaluate the repeatability and stability of the LC-MS analysis process.

[0069] 2. LC-MS / MS analysis: Metabolites were separated and detected using a Waters 2D UPLC (Waters, USA) tandem Q Exactive HF high-resolution mass spectrometer (Thermo Fisher Scientific, USA).

[0070] (1) Chromatographic conditions: The chromatographic column used was a BEH C18 column (1.7 μm 2.1×100 mm, Waters, USA). The mobile phase in positive ion mode was an aqueous solution containing 0.1% formic acid (solution A) and 100% methanol containing 0.1% formic acid (solution B). The mobile phase in negative ion mode was an aqueous solution containing 10 mM formate (solution A) and 95% methanol containing 10 mM formate (solution B). Elution was performed using the following gradient: 0–1 min, 2% solution B; 1–9 min, 2%–98% solution B; 9–12 min, 98% solution B; 12–12.1 min, 98% solution B–2% solution B; 12.1–15 min, 2% solution B. The flow rate was 0.35 mL / min, the column temperature was 45℃, and the injection volume was 5 μL.

[0071] (2) Primary and secondary mass spectrometry data were acquired using a Q Exactive HF mass spectrometer (Thermo Fisher Scientific, USA). The mass-to-nucleus ratio range for mass spectrometry scanning was 70–1050, the primary resolution was 120,000, the AGC was 3e6, and the maximum injection time (IT) was 100 ms. Based on the precursor ion intensity, the top 3 fragmentation sites were selected for secondary data acquisition. The secondary resolution was 30,000, the AGC was 1e5, the maximum injection time (IT) was 50 ms, and the stepped energies were set to 20, 40, and 60 eV. Ion source (ESI) parameter settings: Sheath gas flow rate is 40, Aux gas flow rate is 10, Spray voltage (|KV|) is 3.80 for positive ion mode and 3.20 for negative ion mode, Capillary temperature is 320℃, and Aux gas heater temperature is 350℃.

[0072] During instrument testing, samples are randomly sorted to provide more reliable experimental results, thereby reducing systematic errors. One QC sample is interspersed among every 10 samples.

[0073] IV. Data Preprocessing

[0074] 1. Proteomics data preprocessing:

[0075] (1) DDA data analysis: The identification and quantitative analysis of DDA data were performed using the software MaxQuant. The identification information that meets the condition FDR<0.01 will be used to build the final spectral library for subsequent DIA analysis.

[0076] (2) DIA data analysis: The retention time of the DIA data was corrected using iRT peptides. Then, based on the Target-decoy model applicable to SWATH-MS, false positive control was performed with FDR < 0.01, thus obtaining significant quantitative results.

[0077] (3) Quality control analysis: PCA analysis was performed on the protein data identified in tumor samples, adjacent normal samples and QC samples. The degree of aggregation of QC samples was used to reflect the stability of the instrument. Outliers in tumor samples and adjacent normal samples were removed.

[0078] (4) The total peak area of ​​the quality control data was normalized, and the abundance of each protein after normalization was converted by log2 to be used as the expression level of the protein.

[0079] 2. Metabolomics data preprocessing:

[0080] (1) Raw data processing: The raw mass spectrometry data acquired by LC-MS / MS was imported into CompoundDiscoverer 3.1 software for data processing, including peak extraction, retention time correction within and between groups, adduct ion merging, missing value filling, background peak labeling, and metabolite identification. Information such as compound molecular weight, retention time, and peak area were obtained.

[0081] (2) Data normalization: Import the results of the previous step into the software metaX, and use the probability quotient normalization method (PQN) to normalize the data to obtain the relative peak area.

[0082] (3) Batch correction: Batch correction is performed using the local polynomial regression fitting signal correction method (QC-RLSC).

[0083] (4) Principal component analysis is used to check the degree of aggregation of QC samples and to determine the stability of the instrument. Metabolites with a relative peak area variation coefficient greater than 30% in the QC samples are also removed.

[0084] (5) Merging of positive and negative ion modes: The metabolite data detected in both positive and negative ion modes are merged into a single list. For metabolites identified in both modes, the metabolite that best matches the spectrum is retained. The obtained peak area data is transformed by log2 and used as the expression level of the metabolite.

[0085] V. Establishment of joint proteomics and metabolomics typing and bioinformatics analysis for esophageal squamous cell carcinoma: The Similarity Network Fusion (SNF) method was used to jointly analyze the preprocessed protein expression matrix and metabolite expression matrix. The software used was the R package SNFtool (V2.2).

[0086] The main steps are:

[0087] 1. Perform z-score standardization on the two sets of data to ensure that the expression of each protein or metabolite in each sample has a mean of 0 and a standard deviation of 1.

[0088] 2. Use the dist2 function to calculate the distance matrix at the sample levels of the two omics datasets. In this case, Euclidean distance is used for the calculation.

[0089] 3. Use the affinityMatirix function to construct a sample similarity matrix using the distance matrix. This similarity matrix is ​​equivalent to a similarity network with samples as nodes and edges representing the similarity between two samples.

[0090] 4. Use the SNF function to perform similar network fusion on the two networks, and finally obtain the fusion network of the samples in the two omics.

[0091] 5. The `displayClusters` function was used to cluster the samples based on the fusion network, and cluster heatmaps of the fusion network were plotted when the number of clusters K was 2, 3, 4, 5, 6, and 7, as shown below. Figure 1 As shown, it can be seen that the classification effect is best when the number of clusters K=3. Therefore, queue 1 is clustered into three subtypes: esophageal squamous cell carcinoma type I, esophageal squamous cell carcinoma type II and esophageal squamous cell carcinoma type III.

[0092] 6. Subsequently, according to Figure 1 The clustering situation was plotted using the R package ComplexHeatmap, as shown below. Figure 2 The composite heatmap shown is used to illustrate the clustering of proteomics and metabolomics. Figure 2 The horizontal axis represents different samples, and the vertical axis represents proteins or metabolites. Red and blue represent the relative expression levels of each protein or metabolite in different samples; darker red indicates higher expression levels, and darker blue indicates lower expression levels. From... Figure 2 It can be seen that the proteomics and metabolomics data of the three subtypes all show good clustering, and the clustering pattern is consistent with... Figure 1 The clustering heatmaps are consistent when K=3. Meanwhile... Figure 2 The sample also displayed some clinical information, such as age, gender, tumor stage, tumor grade, and lymph node metastasis.

[0093] VI. Prognostic Verification of Esophageal Squamous Cell Carcinoma Subtypes

[0094] 1. Survival analysis and Log-rank test: Survival analysis was performed using the R package survival. Survival time and status were calculated using the patient's overall survival (OS) and survival status or disease-free survival (DFS) and disease progression status. Grouping was performed using the patient's subtype information. The log-rank test was used to assess the significance of the prognostic differences among the three subtypes. A p-value less than 0.05 indicates that there is a significant difference in prognosis among the different subtypes.

[0095] 2. Visualization of survival analysis results: Kaplan-Meier curves for the three subtypes were plotted using the ggsurvplot function in the R package survminer. It was found that patients with esophageal squamous cell carcinoma type I had the longest overall survival and disease-free survival, and the best prognosis, while patients with esophageal squamous cell carcinoma type III had the shortest overall survival and disease-free survival, and the worst prognosis. Figure 3 The Log-rank test results showed significant differences in overall survival (p = 1.575e-6) and disease-free survival (p = 1.074e-5) among the three subtypes. This indicates that the esophageal squamous cell carcinoma classification of this invention can effectively predict patient prognosis.

[0096] 3. Differences in clinical features related to tumor progression among subtypes: Using the chi-square test to analyze the depth of tumor invasion, presence of lymph node metastasis, and TNM stage in patients with the three subtypes, it was found that the proportion of patients with lymph node metastasis was significantly higher in patients with esophageal squamous cell carcinoma type III, which had the worst prognosis. Figure 4 ).

[0097] VII. Feature Screening for Three Subtypes and Construction of Classifier Model: To screen for molecular features that can effectively distinguish esophageal squamous cell carcinoma type III from types I and II, which have the worst prognosis, we constructed a classifier model using the following method:

[0098] 1. Screening for proteins and metabolites with different abundance between tumor tissues and adjacent normal tissues: After the protein expression matrix and metabolite abundance matrix in cohort 1 were transformed by log2, the p-value of the difference between each protein and metabolite between tumor tissues and adjacent normal tissues was calculated using the Wilcoxon rank-sum test. The corrected p-value was obtained by the FDR correction method. A corrected p < 0.05 indicates a significant difference.

[0099] 2. Screening for proteins and metabolites with varying abundance among the three subtypes of esophageal squamous cell carcinoma: After the protein expression matrix and metabolite abundance matrix in cohort 1 were transformed by log2, the p-values ​​of the differences in protein and metabolite abundance among the three subtypes were calculated using the Kruskal-Wallis test. The corrected p-values ​​were obtained using the FDR correction method. A corrected p-value < 0.05 indicates a significant difference.

[0100] 3. Screen out proteins and metabolites that show significant differences in both steps 1 and 2;

[0101] 4. Randomly divide the samples in queue 1 into training set and validation set according to a ratio of 70% and 30%;

[0102] 5. Feature Recognition: In the training set, a supervised random forest classification model was constructed using the R package RandomForestSRC to identify the 10 differentially expressed proteins and 10 differentially expressed metabolites that contributed the most to the typing in the proteomics and metabolomics data, respectively.

[0103] 6. Constructing the Classifier Model: First, stepwise regression was used on the training set to combine the 20 molecules from the previous step to preliminarily screen for molecular combinations with good classification performance. A combination consisting of four proteins (HK3, SCFD1, SEPTIN5, SSR1) and two metabolites (Creatine, 2'DG) was found to effectively distinguish between esophageal squamous cell carcinoma types I, II, and III. Further, the `cv.glmnet` function of the R package `glmnet` was used to construct a logistic regression model for six molecules. The model with the smallest lambda value was selected as the classifier model and fitted to the training set samples. The `misClassError` function was used to calculate the prediction error rate, and the `plot.roc` function was used to plot the ROC curve of the model's predictions. The results showed that this classifier model had excellent predictive power for classification, with an AUC of 0.9086.

[0104] 7. The formula for calculating the model score from the previous step, derived using the coef function, is: score = -1.384 + 0.449 × exp(SCFD1) + 0.981 × exp(Creatine) + 0.362 × exp(SEPTIN5) - 0.608 × exp(HK3) - 0.111 × exp(SSR1) - 0.373 × exp(2'DG). When the score value ≤ 0, the sample is esophageal squamous cell carcinoma type I or II; when the score value > 0, the sample is esophageal squamous cell carcinoma type III.

[0105] 8. Use the R package vip to calculate the contribution of the 6 molecules to the classifier model. The contribution of the 6 molecules is as follows: Figure 5 As shown, Creatine contributed 34.02%, HK3 contributed 21.08%, SCFD1 contributed 15.57%, 2'DG contributed 12.93%, SEPTIN5 contributed 12.55%, and SSR1 contributed 3.85%.

[0106] 9. Validating the classifier model on the validation set: The predicted function was used to validate the model on the validation set with the same model. The misClassError function was used to calculate the error rate after prediction, and the plot.roc function was used to plot the ROC curve of the model's predictions. It was found that the classifier also had good predictive performance, with AUC = 0.9148. Figure 6 ).

[0107] VIII. External Cohort Validation of the Relationship between Subtyping Model and Prognosis: To validate whether our subtyping is applicable to other esophageal squamous cell carcinoma cohorts, we tested six molecules from the above classifier model in cohort 2 and used the classifier to predict the molecular subtyping of patients in cohort 2. The specific steps are as follows:

[0108] 1. Detection of four proteins: Immunohistochemical staining of four proteins was performed on paraffin sections of tumor tissue from patients in cohort 2. The staining steps are as follows:

[0109] (1) Dewaxing and hydration of paraffin sections: The paraffin sections were treated in an oven at 70°C for 10 minutes and dewaxed in the following order: xylene 15 minutes - xylene 10 minutes - xylene 10 minutes - anhydrous alcohol 5 minutes - anhydrous alcohol 5 minutes - 90% alcohol 2 minutes - 80% alcohol 2 minutes - 70% alcohol 2 minutes.

[0110] (2) Wash the slices with triple-distilled water for 5 minutes, then wash with PBS for 2 minutes, 3 times.

[0111] (3) Antigen retrieval: In this example, an alkaline Tris-EDTA retrieval solution with pH = 9.0 was used for antigen retrieval. The pressure cooker was preheated, and after the retrieval solution boiled, the slides were immersed in it, the lid was closed, the pressure valve was placed, and heating was continued on high heat. After the pressure valve turned and released steam, the heat was reduced to medium and the timer was set for 2.5 minutes. After the timer was set, heating was stopped, the pressure was released, and the retrieval solution tank was placed in cold water to cool to room temperature. The slides were then removed and washed with PBST buffer for 2 minutes, 3 times.

[0112] (4) Endogenous peroxidase blocking and smearing: Place the slides in 3% hydrogen peroxide solution and incubate at room temperature in the dark for 15 minutes. Then remove them and wash with PBS buffer for 2 minutes, 3 times. Use an immunohistochemical pen to circle the stained area, add an appropriate amount of normal goat serum working solution for blocking, incubate at room temperature for 10-15 minutes, and discard the serum.

[0113] (5) Primary antibody incubation: Dilute the antibody to an appropriate concentration, add it to the target area, and place the slide in a humidified chamber. Incubate at 37°C for 1 hour. Wash with PBST buffer for 5 minutes, 3 times.

[0114] (6) Secondary antibody incubation: Add secondary antibody and place in a humidified chamber, incubate at 37°C for 30 minutes, wash with PBST buffer for 5 minutes, 3 times.

[0115] (7) DAB staining: Prepare DAB working solution, add it to the staining area, observe the staining under a light microscope, and immediately immerse the slide in tap water to stop staining after staining occurs.

[0116] (8) Counterstaining, decolorization, and blueing: Stain with hematoxylin staining solution at room temperature for 2 minutes and 45 seconds, then rinse with tap water. Decolorize the sections in a hydrochloric acid-ethanol bath for 3-5 seconds, then rinse with tap water. Blue the sections in an ammonia bath and observe the staining under a microscope.

[0117] (9) Dehydrated, transparent, neutral resin sealing film.

[0118] Pathological section scanning was performed on the sections after staining the four proteins. Three to five regions of fixed size were cut from each sample. The cut regions were analyzed using the Postive Pixel Count v9 plugin of AperioImageScope software. The H-score was calculated according to the formula H-score = (weak positive %) × 1 + (moderate positive %) × 2 + (strong positive %) × 3. The average of the H-score values ​​was used as the protein expression level.

[0119] 2. Detection of two metabolites: Two metabolites were detected using spatial metabolomics. The detection steps are as follows:

[0120] (1) Tissue embedding and sectioning: The samples were embedded using OCT frozen section embedding medium (SAKURA 4583). The embedded tissues were equilibrated in a cryostat at -20°C for 1 hour and then sectioned serially at a thickness of 10µm. One section was then stained with H&E.

[0121] (2) Mass spectrometry imaging (MSI) detection: After the tissue sections were vacuum dried for 15 minutes, mass spectrometry imaging detection was performed using a Q-Exactive high-resolution mass spectrometer in positive ion mode and negative ion mode.

[0122] (3) Data extraction: Mass Imager software was used to identify mass spectrometry imaging data. The H&E staining image of each sample was overlaid with the mass spectrometry image, and six regions of the same size were selected in the tumor tissue of each sample for data extraction. In addition, three regions were extracted from the blank area of ​​each slice as quality control data.

[0123] (4) Data merging: Using MakerView software, peak alignment and normalization were performed on all extracted data in both positive and negative ion modes, and isotope peaks were removed. Simultaneously, PCA analysis was performed on all quality control data using SIMCA software to check instrument stability.

[0124] (5) Data preprocessing: In the merged data of the previous step, remove the peaks that were not detected in more than 50% of the samples, and fill the remaining missing values ​​using the KNN algorithm. PCA analysis is used to remove abnormal samples.

[0125] (6) Metabolite identification: Metabolites were matched in the HMDB database according to m / z, and matching results with Delta (ppm) < 5 were retained. The abundance of Creatine and 2'DG was finally obtained as their expression levels.

[0126] 3. Classifier model predicts patient subtype:

[0127] (1) Data processing and classification prediction: The expression levels of the six molecules were merged and the data were standardized. The classifier model mentioned above was used to predict the samples as esophageal squamous cell carcinoma type I, II and III.

[0128] (2) Survival analysis: Survival analysis was performed on patients with predicted esophageal squamous cell carcinoma types I, II, and III. The analysis steps were the same as in cohort 1, and Kaplan-Meier curves were plotted to visualize the survival analysis results. The results are as follows: Figure 7 As shown, Figure 7 The blue lines represent patients with esophageal squamous cell carcinoma type I and II predicted by the classifier model, while the red lines represent patients with esophageal squamous cell carcinoma type III predicted by the model. Analysis shows that patients with predicted esophageal squamous cell carcinoma type III have a worse prognosis, consistent with the results of cohort 1, demonstrating that the esophageal squamous cell carcinoma classification of this invention can effectively predict patient prognosis.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. The use of molecular subtyping markers for esophageal squamous cell carcinoma based on a combination of DIA proteomics and non-targeted metabolomics in the preparation of diagnostic reagents for esophageal squamous cell carcinoma, characterized in that: The biomarker is a combination of four proteins and two metabolites, wherein the four proteins are: hexokinase 3 (HK3), Sec1 family domain-containing protein 1 (SCFD1), and cytoschizokinin 5 (…). septin 5 SEPTIN5), signal sequence receptor α subunit ( Signal sequence receptor subunit 1, SSR1); the two metabolites are creatine and 2-deoxy-D-glucose (2'DG). The markers were used to detect esophageal squamous cell carcinoma ( Esophageal squamous cell carcinoma The specific method for ESCC classification is as follows: (1) Collect tumor tissues and paired adjacent normal tissues for omics detection, and collect corresponding clinical information data to form an actual analysis cohort; (2) DIA quantitative proteomics and non-targeted metabolomics detection were performed on the actual analysis cohort; the obtained data were subjected to quality control analysis, abnormal samples were removed, and the detection data of the remaining samples were re-integrated; the obtained data were standardized, and proteins and metabolites that were significantly different between tumor tissue and adjacent normal tissue were screened out for subsequent subtyping; (3) Based on the expression levels of differentially expressed proteins and metabolites in tumor tissues obtained in step (2), ESCC patients are classified into ESCC type I, ESCC type II and ESCC type III using the similarity network fusion method; (4) Screen the characteristic proteins and metabolites of each subtype in step (3) and construct a classifier model that can distinguish between type I, type II and type III; The classifier model consists of a combination of four proteins and two metabolites. The four proteins are HK3, SCFD1, SEPTIN5, and SSR1; the two metabolites are Creatine and 2'DG. The classifier model score is calculated using the following formula: score = -1.384 + 0.449 × exp(SCFD1) + 0.981 × exp(Creatine) + 0.362 × exp(SEPTIN5) - 0.608 × exp(HK3) - 0.111 × exp(SSR1) - 0.373 × exp(2'DG). When the score is ≤ 0, the sample is classified as ESCC type I or II; when the score is > 0, the sample is classified as ESCC type III. Patients with ESCC type I have the longest overall survival and disease-free survival, and the best prognosis, while patients with ESCC type III have the shortest overall survival and disease-free survival, and the worst prognosis.

2. The use according to claim 1, characterized in that: The clinical information includes: age, gender, tumor stage, degree of tumor invasion, lymph node metastasis, survival status, and overall survival.