Esophageal squamous cell carcinoma diagnosis marker combination and screening method and application thereof
By employing a tissue-serum combined screening strategy and a random forest diagnostic model, 36 serum metabolites were screened for the early diagnosis of esophageal squamous cell carcinoma, solving the problems of insufficient sensitivity and specificity in existing technologies and achieving a highly efficient non-invasive early diagnostic effect.
Patent Information
- Application Number
- CN202610652856.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies cannot provide a highly sensitive and specific non-invasive diagnostic method for the early diagnosis of esophageal squamous cell carcinoma (ESCC), and existing serum markers have low sensitivity and specificity, especially limited ability to detect early lesions.
A tissue-serum joint screening strategy was adopted, using high-resolution, high-sensitivity ultra-high performance liquid chromatography-quadrupole time-of-flight mass spectrometry (UHPLC-QTOF/MS) to simultaneously analyze cancerous tissue, adjacent tissue and serum samples of ESCC patients, screened out 36 serum metabolites as diagnostic markers, and constructed a random forest diagnostic model.
It achieves accurate, economical, and non-invasive early diagnosis of ESCC, with high sensitivity and specificity. In particular, the sensitivity for early ESCC is 97.5% and the specificity is 98.9%, effectively improving the early detection rate of esophageal cancer.
Smart Images

Figure CN122631808A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical detection technology, specifically to a combination of serum metabolic diagnostic markers for the early diagnosis of esophageal squamous cell carcinoma (ESCC), its screening method, a diagnostic model based on the combination of diagnostic markers, and its construction method. Background Technology
[0002] Esophageal cancer is one of the most common malignant tumors worldwide. China has the highest incidence and mortality rates of esophageal cancer, with over 90% of cases being histologically classified as esophageal squamous cell carcinoma (ESCC). ESCC has an insidious onset, with atypical early symptoms, and most patients are diagnosed at an intermediate or advanced stage, resulting in a 5-year survival rate of less than 25%. In contrast, the 5-year survival rate for patients diagnosed at an early stage (TNM stage 0-II) can be increased to 47%-83%. Therefore, early diagnosis is crucial for improving patient prognosis.
[0003] Currently, commonly used clinical methods for esophageal cancer diagnosis, such as barium swallow X-ray and CT scans, are not sensitive to early lesions. While endoscopy combined with biopsy is the gold standard, it is an invasive procedure with poor patient compliance, complex operation, and high cost, making large-scale implementation difficult. Existing serum protein markers, such as carcinoembryonic antigen (CEA) and cytokeratin 19 fragment (CYFRA21-1), are widely used, but their sensitivity and specificity are not high, especially their ability to detect early esophageal cancer is limited.
[0004] Metabolomics, by monitoring changes in endogenous small-molecule metabolites after disease disturbance, can directly reflect the physiological and pathological state of the body, providing a powerful tool for discovering novel biomarkers. However, existing metabolomics studies based on single serum samples are easily affected by confounding factors such as diet, lifestyle, and environment, resulting in high false-positive rates and poor reproducibility between different studies, thus limiting their clinical application. Although tissue metabolomics can directly reflect metabolic changes in the tumor microenvironment, tissue sampling is invasive and not suitable for screening. Therefore, how to integrate the advantages of both types of samples to screen for serum biomarkers that can truly reflect tumor metabolic abnormalities and are suitable for non-invasive diagnosis is a pressing technical challenge in this field. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a diagnostic marker combination for esophageal squamous cell carcinoma. This diagnostic marker combination consists of serum metabolic markers and is obtained through an innovative tissue and body fluid combined screening strategy. It can reflect the pathological state of esophageal squamous cell carcinoma with high specificity and high sensitivity, and is especially suitable for the early non-invasive diagnosis of esophageal cancer.
[0006] This invention is the first to employ a combined tissue-serum screening strategy. First, we utilize high-resolution, high-sensitivity ultra-high performance liquid chromatography-quadrupole time-of-flight mass spectrometry (UHPLC-QTOF / MS) to simultaneously analyze paired cancerous and adjacent normal tissues from ESCC patients, as well as serum samples from ESCC patients and healthy volunteers. By comparing the two sets of data, we screened for metabolites that were significantly perturbed in cancerous tissues and showed a consistent trend of change in the peripheral blood of patients. This strategy eliminates a large number of "noise" metabolites caused by individual differences and environmental factors at the source, retaining biomarkers that truly reflect the pathological process of tumors, greatly improving the reliability and clinical translational potential of diagnostic markers.
[0007] Based on the above strategy, this invention ultimately identified a combination of diagnostic markers for esophageal squamous cell carcinoma consisting of 36 serum metabolites. This combination covers multiple metabolic pathways that are crucial in tumor development and progression, including amino acid metabolism, purine metabolism, and central carbon metabolism.
[0008] Furthermore, the diagnostic marker combination for esophageal squamous cell carcinoma of the present invention consists of the following 36 serum metabolic markers: (S)-2-aminobutyric acid (Aminobutanoate), betaine, dopamine, D-erythrulose, iminoglycine, hypoxanthine, 4-aminobutyraldehyde, L-glutamate, 4-aminocatechol, xanthine, L-aspartate, and L-galacton-1,4-lactone. o-1,4-lactone), Glycerone, 3-(4-hydroxyphenyl)lactate, 3-Sulfopyruvate, N-Benzyloxycarbonylglycine, 3-Deoxy-D-manno-2-octyl ketone, (R)-3-Ureidoisobutyrate, cis-4,5-dihydroxycyclohexa-1(6),2-diene-1,2-dicarboxylic acid (cis-4,5-Dihydroxycyclohexa-1(6),2-diene-1,2-dicarboxylic acid),2-dicarboxylate), N5-Phenyl-L-glutamine, N8-Acetylspermidine, Carboxynorspermidine, N(ω)-Hydroxyarginine, Dehydroascorbate, Indolelactate, 6-Acetamido-2-oxohexanoate, 2-Amino-2-deoxy-D-gluconate, Nopaline, (5-L-Glutamyl)-L-glutamate, Thiamine, Deoxyguanidinoproclavaminic acid The following compounds are listed: isoniazid alpha-ketoglutaric acid, N-acetylneuraminate, portulacaxanthin II, diethanolamine, and 1-aminocyclopropane-1-carboxylate.
[0009] This invention also provides a method for screening the above-mentioned combination of diagnostic markers for esophageal squamous cell carcinoma. This method comprehensively utilizes tissue metabolomics and serum metabolomics for paired screening, and specifically includes the following steps: (1) Collect paired cancer tissue and adjacent normal tissue samples from patients with esophageal squamous cell carcinoma, as well as serum samples from patients and healthy volunteers; (2) Untargeted metabolomics detection was performed on tissue samples and serum samples using ultra-high performance liquid chromatography-quadrupole time-of-flight mass spectrometry (UHPLC-QTOF / MS); (3) Through multivariate statistical analysis, differential metabolites in cancer tissue compared with adjacent tissue and in patient serum compared with healthy serum were screened out. (4) Based on the mass-to-charge ratio and retention time of the metabolites, a total of 36 metabolites that were present in both tissue and serum samples and showed a consistent dysregulation trend were screened out. The combination of these 36 metabolites is the diagnostic marker combination for esophageal squamous cell carcinoma.
[0010] The present invention also provides the application of the above-mentioned combination of diagnostic markers for esophageal squamous cell carcinoma in the preparation of early diagnostic products for esophageal squamous cell carcinoma.
[0011] The present invention also provides an early diagnostic kit for esophageal squamous cell carcinoma, which includes the above-mentioned combination of diagnostic markers for esophageal squamous cell carcinoma.
[0012] This invention also provides a method for constructing an early diagnostic model of esophageal squamous cell carcinoma and the obtained early diagnostic model of esophageal squamous cell carcinoma. The method includes the following steps: (1) Serum samples were collected from patients with esophageal squamous cell carcinoma and healthy individuals as analytical samples. The metabolic fingerprint profiles of the samples were obtained by detecting and analyzing the samples using liquid chromatography-mass spectrometry (LC-MS). (2) Perform data preprocessing on the metabolic fingerprint spectrum, extract the abundance information of the above 36 serum metabolic markers, and form a diagnostic marker data matrix; (3) Based on the above two-dimensional matrix of diagnostic markers, a random forest model was constructed using the randomForest package in R language to obtain the diagnostic model of esophageal squamous cell carcinoma.
[0013] Furthermore, in step (1), the analysis samples are derived from patients with esophageal squamous cell carcinoma diagnosed by histopathology and healthy volunteers, wherein the healthy volunteers are individuals who have been confirmed by endoscopy to have no upper gastrointestinal lesions.
[0014] Furthermore, when using this diagnostic model, serum samples of the individual to be tested are obtained, and metabolic fingerprints are obtained using liquid chromatography-mass spectrometry (LC-MS). The metabolic fingerprints are preprocessed to extract the abundance information of 36 serum metabolic markers, forming a diagnostic marker data matrix. This diagnostic marker data matrix is then input into the constructed esophageal squamous cell carcinoma diagnostic model to obtain the predicted probability values output by the model. These predicted probability values are compared with a diagnostic threshold. If the predicted probability value is greater than or equal to the diagnostic threshold, the probability of having esophageal squamous cell carcinoma is relatively high; if it is less than the diagnostic threshold, the probability of not having esophageal squamous cell carcinoma is relatively high.
[0015] The diagnostic marker combination, model, and screening method provided by this invention offer a precise, economical, and non-invasive new early diagnosis solution for esophageal squamous cell carcinoma. The random forest diagnostic model constructed based on this combination demonstrates excellent diagnostic efficacy: its area under the receiver operating characteristic (AUC) is as high as 0.982 (95% confidence interval: 0.957-0.998). Crucially, the model achieves a sensitivity of 97.5% and a specificity of 98.9% for early-stage ESCC (TNM stage 0-II), and a sensitivity of 92.7% and a specificity of 93.8% for advanced-stage patients (TNM stage III-IV), effectively addressing the core weakness of existing technologies in early esophageal cancer diagnosis. This diagnostic model is expected to be widely applied to high-risk population screening and clinical auxiliary diagnosis, and has significant clinical and social implications for improving the early detection rate, prognosis, and mortality rate of esophageal cancer patients. Attached Figure Description
[0016] Figure 1 Cluster analysis heatmap of 36 serum metabolic biomarkers between ESCC patients and healthy volunteers.
[0017] Figure 2 ROC curve of a random forest diagnostic model based on 36 serum metabolic biomarkers for diagnosing ESCC. Detailed Implementation
[0018] The technical solution and effects of the present invention will be further explained below with reference to specific embodiments, but the present invention is not limited thereto.
[0019] Example 1: Screening and Identification of Diagnostic Markers 1. Study subject enrollment and sample collection Relying on the Shandong Province Upper Gastrointestinal Cancer Screening Platform, we recruited 81 patients with newly diagnosed ESCC who were given surgical treatment as their first choice, as well as 81 healthy volunteers.
[0020] Inclusion criteria: (1) ESCC patient group: diagnosed with esophageal squamous cell carcinoma by histopathology, and not having received any anti-tumor treatment; no history of other malignant tumors; no metabolic diseases such as diabetes and hyperthyroidism, or neurological diseases affecting metabolism. (2) Healthy volunteer group: confirmed by upper gastrointestinal endoscopic ultrasound screening to have no upper gastrointestinal lesions; no history of malignant tumors or metabolic diseases. Both groups excluded patients with a history of surgery within 6 months. Among the ESCC patients, 43 patients were in the early stage (TNM stage 0-II), accounting for 53.1%; and 38 patients were in the middle and late stages (TNM stage III-IV), accounting for 46.9%.
[0021] Sample collection: (1) Serum samples: Fast for 8-12 hours before blood collection. Collect peripheral blood using a 5mL vacuum blood collection tube. Centrifuge at 3000rpm for 15min at 4℃ within 2 hours to separate the serum. Aliquot the serum into cryovials (400μL per tube) and store at -80℃. (2) Tissue samples: Collect surgically removed tissue specimens within 30min of ex vivo. For cancerous tissue (Tumor), collect samples from the lesion site and remove superficial inflammatory cells and necrotic tissue. For adjacent normal tissue (NAT), collect normal esophageal tissue more than 2cm from the outer edge of the lesion. After sampling, rinse with physiological saline, divide into 20-50mg portions, and store at -80℃.
[0022] 2. Sample pretreatment and LC-MS detection Serum samples were thawed at 4°C, and 50 μL was added to a 96-well plate. 150 μL of pre-cooled methanol was added, and the plate was vortexed for 30 s and incubated at -20°C for 2 h to precipitate proteins. The samples were then centrifuged at 4000 rpm for 20 min at 4°C, and the supernatant was transferred to an LC-MS vial and stored at -80°C for analysis. For tissue samples, 10 mg was homogenized with 200 μL of a ceramic microbead aqueous solution (homogenized at 6000 rpm for 20 s, three times with 5 s intervals, kept at low temperature by liquid nitrogen). Then, 800 μL of a 1:1 methanol-acetonitrile extraction solvent was added, and the sample was subjected to three cycles of vortexing-liquid nitrogen freeze-thaw-ultrasonic lysis. The sample was incubated at -20°C for 1 h to precipitate proteins, centrifuged at 13000 rpm for 15 min at 4°C, and the supernatant was vacuum dried. The supernatant was then reconstituted with 100 μL of a 1:1 acetonitrile-water solution, sonicated for 10 min, and centrifuged to collect the supernatant for analysis.
[0023] A quality control (QC) sample is inserted for every eight test samples to monitor the stability of data acquisition. The QC serum sample is prepared by mixing a small amount of each of the serum samples, and the same applies to the QC tissue sample.
[0024] Metabolomics analysis was performed using an ultra-high performance liquid chromatography (UHPLC) system (1290 UHPLC, Agilent) tandem triple quadrupole time-of-flight mass spectrometer (TripleTOF 6600, AB SCIEX). An amide column (1.7 μm; 2.1 mm × 100 mm) was used for chromatographic separation at a column temperature of 25 °C, a flow rate of 0.5 mL / min, and an injection volume of 2 μL. Mobile phase A was 25 mM ammonium hydroxide aqueous solution in positive ion mode (ESI+) and 25 mM ammonium acetate aqueous solution in negative ion mode (ESI-); mobile phase B was acetonitrile. The linear gradient elution program was: 0–0.5 min 95% B, 0.5–7 min 95% B, 7–8 min 95%→65% B, 8–9 min 65%→40% B, 9–9.1 min 40%→95% B, 9.1–12 min 95% B. Mass spectrometry scanning range: m / z 50~1200 Da. Ion source parameters: GAS 160psi, GAS2 60psi, CUR 30psi, TEM 600℃, ISVF +5000V / -4500V.
[0025] 3. Data analysis and differential metabolite screening Raw mass spectrometry data were converted to mzXML format using ProteoWizard, and peak identification, retention time correction, and peak alignment were performed using the R package XCMS. Peak annotation was performed using the R package CAMERA. Peaks with a detection rate below 50% in QC samples were excluded, and the signals were normalized using MetNormizer. Missing values were estimated using both minimum and MissForest regression models.
[0026] Partial least squares discriminant analysis (PLS-DA) models were constructed for cancerous tissue vs. adjacent normal tissue and serum from ESCC patients vs. serum from healthy volunteers, respectively. The tissue model had R²Y = 0.95 and Q²Y = 0.599; the serum model had R²Y = 0.958 and Q²Y = 0.743. The reliability of the models was verified through 200 permutation tests.
[0027] The screening criteria for differential metabolites were: variable projection importance (VIP) > 1 in the PLS-DA model and p-value < 0.05 after correction for Benjamini-Hochberg false discovery rate (FDR).
[0028] 4. Combined screening to obtain diagnostic marker combinations Based on the polarity, mass-to-charge ratio (M / Z ±30ppm), and retention time (RT ±30s) of metabolites, tissue-differential metabolites were matched with serum-differential metabolites, retaining only metabolites that showed significant changes in cancer tissue and exhibited a consistent direction of change in the serum of ESCC patients. Thirty-six candidate serum biomarkers were ultimately identified, namely: (S)-2-aminobutyric acid (Aminobutanoate), betaine, dopamine, D-erythrulose, iminoglycine, hypoxanthine, 4-aminobutyraldehyde, L-glutamate, 4-aminocatechol, xanthine, L-aspartate, and L-galactono-1,4-l Actone, Glycerone, 3-(4-Hydroxyphenyl)lactate, 3-Sulfopyruvate, N-Benzyloxycarbonylglycine, 3-Deoxy-D-manno-2-octyl ketone, (R)-3-Ureidoisobutyrate, cis-4,5-Dihydroxycyclohexa-1(6),2-diene-1,2-dicarboxylic acid2-dicarboxylate), N5-Phenyl-L-glutamine, N8-Acetylspermidine, Carboxynorspermidine, N(ω)-Hydroxyarginine, Dehydroascorbate, Indolelactate, 6-Acetamido-2-oxohexanoate, 2-Amino-2-deoxy-D-gluconate, Nopaline, (5-L-Glutamyl)-L-glutamate, Thiamine, Deoxyguanidinoproclavaminic acid The following compounds are listed: isoniazid alpha-ketoglutaric acid, N-acetylneuraminate, portulacaxanthin II, diethanolamine, and 1-aminocyclopropane-1-carboxylate.
[0029] Example 2: Construction and Performance Verification of the Diagnostic Model 1. Study sample allocation and baseline characteristics A total of 162 serum samples (81 ESCC patients + 81 healthy volunteers) were included in this study. Baseline information for ESCC patients and healthy volunteers is shown in Table 1-1: There were no significant differences between the two groups in terms of sex, smoking, and alcohol consumption (p>0.05), but there were statistically significant differences in age and BMI distribution, which were adjusted for during the analysis.
[0030] Table 1-1 Baseline information and clinical characteristics of ESCC patients and healthy volunteers 2. Construction of the diagnostic model The 162 serum samples mentioned above were used as analytical samples. The metabolic fingerprints of the samples were obtained by liquid chromatography-mass spectrometry (LC-MS) under the same conditions as in Example 1.
[0031] The metabolic fingerprint was preprocessed according to the method in Example 1 to extract the abundance information of 36 serum metabolic markers and form a diagnostic marker data matrix. Using the diagnostic marker data matrix as input variables, a random forest diagnostic model was constructed using the randomForest package in R. The model construction process is as follows: (1) Multiple bootstrap sample sets were randomly extracted with replacement from the training set using the bootstrap method to construct multiple classification trees; (2) A portion of variables were randomly extracted from each node of each tree, and the variable with the strongest classification ability was selected for node partitioning; (3) Each tree was allowed to grow to its maximum extent without pruning; (4) The voting results of all classification trees were combined to make a classification judgment. The entire modeling process adopted the 5-fold cross-validation method for internal rotation training and validation, which effectively avoided model overfitting and objectively evaluated the model's stability and generalization ability.
[0032] When the constructed diagnostic model is used for discrimination, the abundance data of 36 metabolites in the serum sample to be tested are input, and the model outputs a probability prediction value between 0 and 1. If the value is greater than or equal to the optimal diagnostic cutoff value (threshold) determined by ROC analysis, it is judged as ESCC positive; if it is less than the cutoff value, it is judged as ESCC negative.
[0033] 3. Performance evaluation of the diagnostic model The diagnostic efficacy of the model was evaluated using receiver operating characteristic (ROC) curves. The results showed that the area under the curve (AUC) of the random forest diagnostic model constructed based on these 36 serum metabolic biomarkers was 0.982 (95% CI: 0.957–0.998), indicating that the model has extremely high discriminative power.
[0034] Further subgroup analysis by tumor stage was conducted to evaluate the model's ability to identify early-stage and advanced-stage ESCC: For patients with early-stage ESCC (TNM stage 0-II) (n=43), the diagnostic sensitivity was 97.5% and the specificity was 98.9%. For patients with advanced-stage ESCC (TNM stage III-IV) (n=38), the diagnostic sensitivity was 92.7% and the specificity was 93.8%.
[0035] This result demonstrates that the diagnostic marker combination and diagnostic model of the present invention not only have excellent detection capabilities for intermediate and advanced ESCC, but more importantly, they also maintain extremely high sensitivity (97.5%) and specificity (98.9%) for early esophageal squamous cell carcinoma (TNM 0-II stage), which is originally difficult to detect, effectively solving the technical problem of lacking sensitive and reliable serum markers for early diagnosis of esophageal cancer in the prior art.
[0036] The scope of protection of this invention is not limited to the specific embodiments described above. Any non-creative modifications or variations made based on the core idea of this invention—that is, using a combined tissue and serum screening strategy to find diagnostic markers in the blood that track abnormal tumor metabolism, and combining and modeling them—are within the scope of protection of this invention.
Claims
1. A combination of diagnostic markers for esophageal squamous cell carcinoma, characterized in that, This combination consists of the following 36 serum metabolic markers: (S)-2-aminobutyric acid (Aminobutanoate), betaine, dopamine, D-erythrulose, iminoglycine, hypoxanthine, 4-aminobutyraldehyde, L-glutamate, 4-aminocatechol, xanthine, L-aspartate, and L-galactono-1,4-lactone. tone), glycerone, 3-(4-hydroxyphenyl)lactate, 3-sulfopyruvate, N-benzyloxycarbonylglycine, 3-deoxy-D-manno-2-octyl ketone, (R)-3-ureidoisobutyrate, cis-4,5-dihydroxycyclohexa-1(6), 2-diene-1,2-dicarboxylic acid,2-dicarboxylate), N5-Phenyl-L-glutamine, N8-Acetylspermidine, Carboxynorspermidine, N(ω)-Hydroxyarginine, Dehydroascorbate, Indolelactate, 6-Acetamido-2-oxohexanoate, 2-Amino-2-deoxy-D-gluconate, Nopaline, (5-L-Glutamyl)-L-glutamate, Thiamine, Deoxyguanidinoproclavaminic acid The following compounds are listed: isoniazid alpha-ketoglutaric acid, N-acetylneuraminate, portulacaxanthin II, diethanolamine, and 1-aminocyclopropane-1-carboxylate.
2. A method for screening the combination of diagnostic markers for esophageal squamous cell carcinoma as described in claim 1, characterized in that, Includes the following steps: (1) Collect paired cancer tissue and adjacent normal tissue samples from patients with esophageal squamous cell carcinoma, as well as serum samples from patients and healthy volunteers; (2) Untargeted metabolomics detection was performed on tissue samples and serum samples using ultra-high performance liquid chromatography-quadrupole time-of-flight mass spectrometry; (3) Through multivariate statistical analysis, differential metabolites in cancer tissue compared with adjacent tissue and in patient serum compared with healthy serum were screened out. (4) Based on the mass-to-charge ratio and retention time of the metabolites, a total of 36 metabolites that were present in both tissue and serum samples and showed a consistent dysregulation trend were screened out. The combination of these 36 metabolites is the diagnostic marker combination for esophageal squamous cell carcinoma.
3. The use of the combination of diagnostic markers for esophageal squamous cell carcinoma as described in claim 1 in the preparation of early diagnostic products for esophageal squamous cell carcinoma.
4. A diagnostic kit for early esophageal squamous cell carcinoma, characterized in that, Includes the combination of diagnostic markers for esophageal squamous cell carcinoma as described in claim 1.
5. A method for constructing an early diagnostic model for esophageal squamous cell carcinoma, characterized in that, Includes the following steps: (1) Serum samples were collected from patients with esophageal squamous cell carcinoma and healthy individuals as analytical samples. The samples were analyzed by liquid chromatography-mass spectrometry to obtain metabolic fingerprint profiles. (2) Perform data preprocessing on the metabolic fingerprint profile, extract the abundance information of the 36 serum metabolic markers in claim 1, and form a diagnostic marker data matrix; (3) Based on the above two-dimensional matrix of diagnostic markers, a random forest model was constructed using the randomForest package in R language to obtain the diagnostic model of esophageal squamous cell carcinoma.
6. The esophageal squamous cell carcinoma diagnostic model constructed according to the construction method described in claim 5.