A protein marker combination, model and application thereof for early screening of esophageal squamous cell carcinoma

By combining protein biomarkers GPLD1, CBL, DGKG, TUBB6, and a machine learning model, the invasiveness and accuracy issues of early screening for esophageal squamous cell carcinoma were resolved, achieving a highly stable and repeatable early screening effect.

CN122631889APending Publication Date: 2026-08-25SHANGHAI JIAOTONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610803809.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies for early screening of esophageal squamous cell carcinoma (ESCC) are highly invasive, have poor patient compliance, and blood-based proteomics analysis is easily affected by individual genetic background and lifestyle, resulting in poor accuracy and reproducibility of biomarkers.

Method used

A naive Bayes model was constructed using a combination of protein biomarkers GPLD1, CBL, DGKG, and TUBB6, along with serum sample detection and machine learning algorithms, to distinguish between the early and precancerous stages of esophageal squamous cell carcinoma. The longitudinal design eliminated the influence of individual differences.

Benefits of technology

It achieves high stability and reproducibility in early screening, with an AUROC of 0.943 in the training set and 0.840 in the independent validation set, making it suitable for large-scale screening with minimal invasiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122631889A_ABST
    Figure CN122631889A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of biomedical detection, and discloses a protein marker combination, a model and application thereof for early screening of esophageal squamous cell carcinoma. The marker combination comprises at least two of GPLD1, CBL, DGKG and TUBB6, can accurately distinguish the early stage of esophageal squamous cell carcinoma from the precancerous lesion stage, and the area under the receiver operating characteristic curve (AUC) is not less than 0.799. The application further discloses a kit based on the above marker, a model construction method and a device, provides a non-invasive and high-sensitivity detection scheme for early screening of esophageal squamous cell carcinoma, and has important clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical detection technology, specifically relating to a combination of protein biomarkers, a model, and their applications for early screening of esophageal squamous cell carcinoma. Background Technology

[0002] Esophageal squamous cell carcinoma (ESCC) is one of the most common malignant tumors of the upper gastrointestinal tract, and its prognosis is closely related to the timing of screening. The 5-year survival rate for early-stage ESCC can reach over 80%, while the 5-year survival rate for late-stage patients is less than 20%. However, because early-stage ESCC lacks specific symptoms, the vast majority of patients are diagnosed at an intermediate or advanced stage, missing the optimal treatment window.

[0003] Currently, early screening for ESCC mainly relies on endoscopic examination combined with histopathological analysis. However, this method has problems such as high invasiveness, poor patient compliance, and high consumption of medical resources, making it difficult to use as a large-scale screening method. Blood-based proteomics analysis provides an effective approach for developing non-invasive screening biomarkers, but existing inventions mostly adopt cross-sectional designs, which are easily affected by confounding factors such as individual genetic background and lifestyle, resulting in poor accuracy and reproducibility of biomarkers.

[0004] Therefore, there is an urgent need in this field to develop an accurate, stable, and non-invasive biomarker and method for early screening of ESCC. Summary of the Invention

[0005] In view of this, the present invention provides a combination of protein biomarkers, a model, and their applications for early screening of esophageal squamous cell carcinoma.

[0006] The objective of this invention is achieved through the following technical solution: In a first aspect, the present invention provides a protein biomarker combination for early screening of esophageal squamous cell carcinoma, the protein biomarker combination being composed of at least two proteins selected from the following: GPLD1, CBL, DGKG, and TUBB6.

[0007] The protein biomarker combination includes GPLD1, CBL, DGKG, and TUBB6.

[0008] In samples from the early stages of esophageal squamous cell carcinoma, compared to samples from precancerous esophageal lesions: The expression level of GPLD1 was downregulated; The expression levels of CBL, DGKG, and TUBB6 were upregulated.

[0009] Preferably, the protein biomarker combination consists of GPLD1, CBL, DGKG, and TUBB6.

[0010] The protein biomarker combination is used to distinguish between the early stage of esophageal squamous cell carcinoma and the precancerous stage of esophageal squamous cell carcinoma. The early phase of esophageal squamous cell carcinoma (eESCC) includes high-grade intraepithelial neoplasia and esophageal squamous cell carcinoma, while the precancerous phase of esophageal squamous cell carcinoma (preESCC) includes esophagitis and low-grade intraepithelial neoplasia. In the serum of patients with early-stage esophageal squamous cell carcinoma, the expression level of GPLD1 was significantly lower than that of patients with precancerous esophageal squamous cell carcinoma; while the expression levels of CBL, DGKG, and TUBB6 were significantly higher than those of patients with precancerous esophageal squamous cell carcinoma.

[0011] In a second aspect, a kit for early screening of esophageal squamous cell carcinoma includes reagents for detecting the expression levels of a combination of protein markers as described in the first aspect in a biological sample.

[0012] The biological samples are serum, plasma, exosomes, tissue homogenates, or pleural effusions.

[0013] Preferably, the biological sample is serum.

[0014] The reagents include specific antibodies, antibody fragments, nucleic acid aptamers, or affinity ligands against GPLD1, CBL, DGKG, and TUBB6.

[0015] The reagents are suitable for enzyme-linked immunosorbent assay (ELISA), chemiluminescence immunoassay, electrochemiluminescence immunoassay, mass spectrometry multiple reaction monitoring (MRM), or proximity extension assay.

[0016] The reagents used for mass spectrometry detection include one or more of the following substances: magnetic beads, washing buffer, lysis buffer, termination buffer, elution buffer, lysine endopeptidase solution, trypsin solution, C18 peptide desalting column, and formic acid solution.

[0017] Thirdly, a model-building device for early screening of esophageal squamous cell carcinoma, including... Data input module: used to acquire expression level data of the combination of protein biomarkers as described in the first aspect in the subject's biological samples; The calculation module is equipped with a classification model built based on the expression levels of the protein biomarker combination, which is used to generate early screening results for esophageal squamous cell carcinoma based on the expression level data. Output module: Used to output the screening results.

[0018] The classification model is either a Naive Bayes model, a Support Vector Machine model, a Random Forest model, a Logistic Regression model, or an XGBoost model.

[0019] The classification model is a Naive Bayes model, and the area under the receiver operating characteristic curve of the Naive Bayes model is not less than 0.799 when distinguishing between the early stage of esophageal squamous cell carcinoma and esophageal precancerous lesions.

[0020] <Fourth Aspect> A method for screening protein biomarkers for early screening of esophageal squamous cell carcinoma, comprising the following steps: a) Collect longitudinal serum samples from patients with precancerous esophageal lesions, including baseline T1 samples and follow-up T2 samples from the same patient; b) For patients who progressed to the early stage of esophageal squamous cell carcinoma during the follow-up period (progression group) and patients who did not progress after the follow-up period (non-progression group), proteomics technology was used to detect the protein expression levels in their serum at baseline and during the follow-up period. c) Within-group paired tests were performed in patients in the progression group and the non-progression group to screen for differentially expressed proteins that showed significant changes in T2 at the follow-up period relative to T1 at the baseline. The screening criteria were a paired test p value <0.05 and a fold change in expression FC >1.20 or FC <0.83. d) Differentially expressed proteins selected in the progression group were compared with those selected in the no-progression group. Proteins that also showed significant changes in the no-progression group were excluded, and the remaining proteins were identified as potential early screening biomarkers for esophageal squamous cell carcinoma. e) Use machine learning algorithms to perform feature selection on the potential biomarkers obtained in step d) to obtain the candidate protein biomarkers; f) After evaluating the statistical robustness of the differences in expression levels of the candidate protein biomarkers obtained in step e) between preESCC and eESCC, the combination of protein biomarkers is obtained.

[0021] The proteomics technology mentioned is a data-independent acquisition mass spectrometry technique.

[0022] The data-independent acquisition mass spectrometry technique employs a parallel accumulation-continuous fragmentation combined with independent data acquisition mode.

[0023] The precancerous lesions of the esophagus include esophagitis and / or low-grade intraepithelial neoplasia; the early stages of the esophageal squamous cell carcinoma include high-grade intraepithelial neoplasia and / or esophageal squamous cell carcinoma.

[0024] The machine learning algorithm described in step e) is the Support Vector Machine-Recursive Feature Elimination Algorithm.

[0025] <Fifth Aspect> A method for constructing an early screening model for esophageal squamous cell carcinoma, characterized by comprising the following steps: (a) Obtain expression level data of the protein biomarker combination obtained by the screening method described above in the training samples; the protein biomarker combination consists of at least two proteins selected from the following: GPLD1, CBL, DGKG, TUBB6; (b) Using the expression level of the protein biomarker combination as input features, a classification model is trained using a machine learning algorithm, wherein the machine learning algorithm is selected from Naive Bayes model, support vector machine, random forest, logistic regression, XGBoost, LightGBM, gradient boosting machine or deep neural network; (c) The model performance was evaluated through cross-validation to obtain an early screening model for esophageal squamous cell carcinoma; The area under the receiver operating characteristic curve (AUC) of the classification model in the independent validation set for distinguishing between early-stage esophageal squamous cell carcinoma and precancerous lesions is no less than 0.799.

[0026] Furthermore, the machine learning algorithm is selected from Naïve Bayes, Support Vector Machine (SVM), Random Forest, Logistic Regression, XGBoost, LightGBM, Gradient Boosting Machine (GBM), or Deep Neural Network (DNN); preferably, the machine learning algorithm is Naïve Bayes.

[0027] Furthermore, the protein biomarker combination consists of GPLD1, CBL, DGKG, and TUBB6.

[0028] As one embodiment of the present invention, a method for constructing an early screening model for esophageal squamous cell carcinoma includes the following steps: Step 1: Screening and Feature Selection of Proteins Related to Esophageal Squamous Cell Carcinoma (1) Obtaining esophageal squamous cell carcinoma-related proteins: Differential expression screening of serum proteomics data at baseline and follow-up in the progression group and the non-progression group respectively. Then, the differentially expressed proteins obtained from the two groups were matched, and differentially expressed proteins common to both groups were excluded. The remaining specific differentially expressed proteins in the progression group were identified as esophageal squamous cell carcinoma-related proteins. The screening criteria for differential expression were P<0.05 and fold change >1.20 or <0.83. The difference between baseline and follow-up was compared using a paired test. (2) Obtaining candidate proteins: The esophageal squamous cell carcinoma-related proteins obtained in step (1) are divided into two categories: significantly upregulated and significantly downregulated. The intersection of these proteins with the differentially expressed proteins between the progression group and the non-progression group during the follow-up period is taken to obtain potential proteins. (3) Feature selection: The potential proteins obtained in step (2) are selected by using the support vector machine-recursive feature elimination algorithm based on feature importance ranking, and the protein with the most informative features is selected as a candidate early screening biomarker; preferably, the selected candidate early screening biomarker includes at least two of GPLD1, CBL, DGKG and TUBB6; more preferably, it is composed of GPLD1, CBL, DGKG and TUBB6.

[0029] Step 2: Construction and Validation of the Prediction Model (a) Obtain the expression level data of the protein biomarker combination screened by the method described in step one in the training samples; (b) Using the expression level of the protein biomarker combination as input features, a classification model is trained using a machine learning algorithm; the machine learning algorithm is selected from Naïve Bayes, Support Vector Machine (SVM), Random Forest, Logistic Regression, XGBoost, LightGBM, Gradient Boosting Machine (GBM) or Deep Neural Network (DNN), preferably Naïve Bayes; (c) The model performance was evaluated through cross-validation to obtain an early screening model for esophageal squamous cell carcinoma; preferably, the model was constructed and validated by repeating 10-fold cross-validation 10 times. (d) Model performance requirements: The area under the receiver operating characteristic (AUROC) curve of the classification model in distinguishing between early stage of esophageal squamous cell carcinoma and precancerous lesions in an independent validation set shall not be less than 0.799.

[0030] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention adopts a vertical design and screens biomarkers by pairing and comparing the same individual, which effectively eliminates the influence of confounding factors such as individual genetic background and lifestyle, and the obtained biomarkers have higher stability and reproducibility.

[0031] (2) The prediction model constructed in this invention achieves an AUROC of 0.943 in the training set, 0.840 in the independent validation set, and 0.799 in the T2 dataset, demonstrating excellent screening performance.

[0032] (3) The present invention uses serum samples for detection, which has the advantages of convenient sampling, minimal trauma and high patient compliance, and is suitable as a large-scale screening method. Attached Figure Description

[0033] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This includes follow-up information for research subjects and sample size; Figure 2 To achieve in-depth identification and quantification of serum proteins in the discovery cohort based on DIA proteomics; A represents the number of proteins quantified in each serum sample at each stage of esophageal squamous cell carcinoma and precancerous lesions; B represents the total number and average number of proteins quantified in serum samples at each stage of esophageal squamous cell carcinoma and precancerous lesions; C represents the quantitative dynamic range of serum proteomics data; D represents the integrity curve of serum proteomics data; E represents the Venn plot showing the overlap between the HPA database and blood proteins in this study; F represents secreted proteins annotated based on the HPA database. Figure 3 A schematic diagram illustrating the identification of candidate protein biomarkers for early screening; Figure 4 The results of expression profiling and functional enrichment analysis of candidate protein biomarkers; Figure 5 To construct protein-protein interaction networks and perform functional annotation based on candidate protein biomarkers; Figure 6 To identify potential protein features for building machine learning models based on support vector machine-recursive feature elimination; Figure 7 Box plots were used to show the differences in expression levels of four early screening protein biomarkers in preESCC and eESCC; Figure 8 To construct a machine learning-based model to distinguish between eESCC and preESCC patients. A shows the ROC curves for distinguishing between eESCC and preESCC patients based on the early diagnosis model in the training set; B shows the ROC curves for distinguishing between eESCC and preESCC patients based on the early diagnosis model in the validation set; C shows the ROC curves for distinguishing between eESCC and preESCC patients based on the early diagnosis model in the T2 dataset. Figure 9 To compare the expression levels of four early screening protein biomarkers between the preESCC and eESCC phases; Figure 10 To validate the differences in expression levels of four early screening protein biomarkers between normal controls and eESCC in an external independent validation cohort based on ELISA; . Detailed Implementation

[0034] The present invention will be described in detail below with reference to embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several adjustments and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0035] Example 1 1. Research Subjects This study included 91 patients diagnosed with precancerous ESCC at baseline (T1) and collected serum samples (n=91), including 6 cases of esophagitis (ESO) and 85 cases of low-grade intraepithelial neoplasia (LGIN). After a median follow-up of 38.6 months (T2), serum samples were collected again (n=91). Among them, 26 patients were diagnosed with high-grade intraepithelial neoplasia (HGIN), 6 patients were diagnosed with ESCC, and the remaining 59 patients were still in the precancerous stage of ESCC. Figure 1 All of the above diagnoses were independently determined by two trained pathologists in accordance with the "Guidelines for Esophageal Cancer Screening and Early Diagnosis and Treatment in China". In case of inconsistent results, the diagnosis was determined through discussion.

[0036] Since HGIN patients typically require endoscopic mucosal resection and other clinical interventions, while LGIN patients generally only require follow-up monitoring, HGIN and ESCC are classified as the early phase of ESCC (eESCC), while ESO and LGIN are combined into the precancerous phase of ESCC (preESCC). Based on whether patients in the preESCC phase progressed to the eESCC phase during follow-up, the 91 patients were divided into a progression group (32 cases) and a no-progression group (59 cases). Furthermore, plasma samples were collected from 73 subjects in an independent external validation cohort for external validation of the biomarkers, including 22 normal controls, 27 HGIN cases, and 24 ESCC cases. This invention aims to obtain serum protein biomarkers for early screening and diagnosis of ESCC through this longitudinal design, and the invention has been approved by the Public Health and Nursing Research Ethics Committee of Shanghai Jiao Tong University School of Medicine.

[0037] 2. Proteomics Data Acquisition and Processing 2.1 Sample Pretreatment Serum sample proteomics pretreatment used a low-abundance protein enrichment and pretreatment kit (OSFP0002-96, Shanghai Omicsolution Co., Ltd.). The process mainly included protein enrichment, protein denaturation, reductive alkylation, enzymatic digestion, and peptide desalting. The specific process is as follows: (1) Materials and reagents: Magnetic beads: included with the kit Washing buffer: Included in the kit Lysis buffer: Included in the kit Termination buffer: Included in the kit Elution buffer: 90% acetonitrile solution containing 0.2% trifluoroacetic acid. Lysyl endopeptidase and trypsin solution: Included in the kit C18 peptide desalting column: included with the kit 500 μg of magnetic beads were dissolved in 50 μL of washing buffer, followed by 50 μL of serum sample. The mixture was incubated at 37°C and 1,000 rpm for 1 h. After incubation, the magnetic beads were collected using a magnetic separation device and washed three times with 150 μL of washing buffer. The precipitate was then resuspended in 20 μL of lysis buffer and heated at 95°C and 1,000 rpm with shaking for 5 min. After cooling to room temperature, lysyl endopeptidase and trypsin solution (enzyme to blood sample volume ratio: 3:20) were added and incubated at 37°C and 1,000 rpm with shaking for 2 h. After the reaction was complete, stop buffer was added, and the supernatant was collected after centrifugation at 20,000 × g for 1 min. The peptide sample was then purified using a C18 peptide desalting column. Clean peptides were obtained by elution twice with 50 μL of elution buffer and centrifugation at 700 × g for 1 min. The sample was then concentrated using a vacuum refrigerated centrifuge and 10 μL of lysin was added. The peptide fragments were resuspended in μL of 0.1% formic acid solution and then stored in an ultra-low temperature freezer at -80℃ for instrumental testing.

[0038] 2. LC-MS / MS analysis Peptide samples were used to acquire proteomics data using a liquid chromatography-mass spectrometry system consisting of the Evosep One (Evosep Biosystems) liquid chromatography system and the timsTOF Pro 2 (Bruker Daltonics) trapping ion mobility coupled time-of-flight mass spectrometer.

[0039] Peptides were analyzed using an Evosep Endurance column (15 cm × 150 μm, 1.9 μm ReproSil-PurC18 beads by Maish) at a 30 SPD (30 samples per day, 44 min) method, with the column temperature maintained at 50 °C. Mobile phases A and B were an aqueous solution containing 0.1% formic acid and an acetonitrile solution containing 0.1% formic acid, respectively. The timsTOF Pro 2 was operated in parallel accumulation-serial fragmentation combined with data-independent acquisition (diaPASEF) mode. Basic parameters are shown in Table 1. Collision energy was linearly related to ion mobility, 1 / K0 = 1.60 Vs / cm. 2 At a voltage of 59.0 eV, 1 / K0 = 0.60 Vs / cm 2 The time is 20.0 eV. The diaPASEF windowing method is optimized using the py_diAID algorithm. The optimized diaPASEF method contains 25 diaPASEF scans, each containing two ion mobility windows (as shown in Table 2).

[0040] Table 1. Basic parameters of mass spectrometry for proteomics data acquisition

[0041] Table 2. Optimized diaPASEF separation window parameters after serum sample proteomics data acquisition.

[0042] 2.3 Database Retrieval Raw mass spectrometry data files were retrieved using Spectronaut (version 19, Biognosys AG) software for database searching. FASTA files were downloaded from the UniProt database of the human proteome. Based on sample files and mixed sample fractionation files, a project-specific mixed spectral library containing 74,424 precursor ions, 54,169 peptides, and 6,852 proteins was built using the Pulsar engine in Spectronaut software for DIA library search analysis. Specific search parameters were as follows: maximum missed cleavage sites were set to 2; peptide lengths ranged from 7 to 52 amino acids; variable modifications selected were methionine oxidation and protein N-terminal acetylation; fixed modifications selected were cysteine ​​iodine acetylation; and the false discovery rate for both peptide and protein levels was set to 1%.

[0043] 3. Proteomics data analysis Protein expression matrices were exported from Spectronaut software for subsequent analysis. Protein data were standardized using the median and transformed with log2. Proteins with quantitative information in at least one-third of the samples were retained for further analysis. Secreted proteins were annotated using the HPA (Human Protein Atlas, version 25.0) database. Protein functional annotation or enrichment was performed using Metascape (version 3.5). Based on whether the protein expression data conformed to a normal distribution, t-tests or Wilcoxon rank-sum tests were used to screen differentially expressed proteins between two groups. Paired tests were used for differentially expressed proteins between the baseline (T1) and follow-up (T2) groups, while unpaired tests were used for other groups. The screening criteria were as follows: P <0.05 and fold change (FC) >1.20 or FC <0.83. The screening process for ESCC-related proteins in this invention is as follows: First, differentially expressed proteins in the T1 and T2 phases are screened in the progression group and the non-progression group, respectively; then, the differentially expressed proteins obtained from the two groups are matched, and the differentially expressed proteins remaining in the progression group after excluding the common differentially expressed proteins are the ESCC-related proteins.

[0044] 4. Biomarker screening and machine learning model construction All machine learning models involved in this invention were constructed in R software using the "mlr3" package (version 0.22.1), and missing values ​​were imputed using the "missForest" package (version 1.5). For ESCC early screening biomarkers, protein feature screening mainly includes the following two steps: First, ESCC-related proteins are classified into significantly upregulated and downregulated categories, and their intersections with differentially expressed proteins in the T2 stage progressive group (eESCC) and the non-progressive group (preESCC) are taken to obtain potential proteins; then, a Support Vector Machine-Recursive Feature Elimination (SVM-RFE) algorithm based on feature importance ranking is used for feature selection, screening out the proteins with the most informative features as candidate early screening biomarkers; finally, the statistical robustness of the obtained candidate protein biomarkers in terms of expression level differences between preESCC and eESCC is evaluated to obtain the final protein biomarker combination.

[0045] This invention uses serum proteomics data from patients in the T1 (preESCC) and T2 (eESCC) stages of the progression group as the discovery dataset, and randomly divides it into a training set (70%) and a validation set (30%). Based on the aforementioned protein features, a Naive Bayes model with 10-fold cross-validation and 10 repetitions is constructed in the training set to distinguish eESCC patients, and the model's performance is evaluated in the validation set. Simultaneously, serum proteomics data from patients in the progression-free group (preESCC) and the progression group (eESCC) in the T2 dataset are used as the validation dataset to evaluate the model's performance.

[0046] 5. Validation of Enzyme-Linked Immunosorbent Assay (ELISA) Biomarkers were validated using ELISA kits in plasma from an independent external validation cohort. Specifically, GPLD1 concentration was determined using a human GPLD1 ELISA kit (FineTest, #EH1534). CBL, DGKG, and TUBB6 concentrations were determined using a human CBL ELISA kit (EIAab, #E15462h), a human DGKG ELISA kit (EIAab, #E11333h), and a human TUBB6 ELISA kit (EIAab, #E11224h), respectively. Absorbance was read at 450 nm and quantification was performed using a BioTek Synergy H1 microplate reader.

[0047] (1) Achieving in-depth identification and accurate quantification of serum proteins based on DIA proteomics This invention, based on DIA proteomics technology, achieves in-depth identification and accurate quantification of serum proteins. A total of 4,854 proteins were quantified from 182 serum samples, averaging 3,444 proteins per sample. This protein dataset exhibits high identification depth, wide dynamic range, and high data integrity, while also containing abundant secreted proteins, laying a solid data foundation for subsequent exploration of early screening protein biomarkers. Figure 2 To enable in-depth identification and quantification of serum proteins in the discovery cohort based on DIA proteomics; Figure 2 A represents the number of proteins quantified in each serum sample at each stage of esophageal squamous cell carcinoma and precancerous lesions; Figure 2 B represents the total number and average number of proteins quantified in serum samples from various stages of esophageal squamous cell carcinoma and precancerous lesions. Figure 2 C represents the quantitative dynamic range of serum proteomics data; Figure 2 D represents the integrity curve of serum proteomics data; Figure 2 E is a Venn diagram showing the overlap between the HPA database and blood proteins in this study; Figure 2 F represents a secreted protein based on annotations from the HPA database.

[0048] (2) Identification of early screening biomarkers for ESCC based on machine learning models This invention identified 144 potential early screening protein biomarkers through matching analysis. Significantly upregulated proteins are mainly involved in biological processes such as oxidative phosphorylation, aerobic respiration, fatty acid metabolism, and lipid biosynthesis, while significantly downregulated proteins are closely related to cell-cell adhesion and the extracellular matrix. Figure 3 , Figure 4 , Figure 5 Based on machine learning models, protein feature screening identified GPLD1, WBP11, CBL, DGKG, TUBB6, and ITPKB as candidate early screening biomarkers for ESCC. Figure 6 To identify candidate protein features for building machine learning models based on support vector machine-recursive feature elimination, and after further evaluating the differences in expression levels of the aforementioned candidate protein features between the preESCC and eESCC groups, robust GPLD1, CBL, DGKG, and TUBB6 were selected as early screening biomarkers. Figure 7The expression levels of four early screening protein biomarkers differed between preESCC and eESCC. Subsequently, a Naive Bayes machine learning model was constructed based on this to distinguish between eESCC and preESCC patients. The area under the receiver operating characteristic curve (AUROC) in the training set reached 0.943. Figure 8 To build an early diagnosis model based on machine learning to distinguish between eESCC and preESCC patients; Figure 8 A represents the ROC curve used in the training set to distinguish between eESCC and preESCC patients based on the early diagnosis model. Meanwhile, the AUROC reached 0.840 when validating the model's performance on the validation set. Figure 8 B represents the ROC curve distinguishing between eESCC and preESCC patients based on the early diagnosis model in the validation set. Further validation of the early diagnosis model on the T2 dataset yielded an AUROC of 0.799, indicating excellent performance. Figure 8 C represents the ROC curve distinguishing between eESCC and preESCC patients based on the early diagnosis model on the T2 dataset. The expression levels of protein biomarkers differed significantly between preESCC and eESCC patients. Specifically, GPLD1 expression was significantly decreased in the serum of eESCC patients, while CBL, DGKG, and TUBB6 were significantly upregulated in serum samples from eESCC patients. Figure 9 This study compared the expression levels of four early screening protein biomarkers between the preESCC and eESCC stages. Finally, to further confirm the universality and effectiveness of the biomarkers, the expression levels of these biomarkers were verified using ELISA in plasma samples from an external validation cohort. Compared to normal controls, GPLD1 expression was significantly reduced in the eESCC group, while CBL, DGKG, and TUBB6 expression was significantly increased. Figure 10 To validate the differences in expression levels of four early screening protein biomarkers between normal controls and eESCC in an external cohort based on ELISA, ).

[0049] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that these are merely illustrative examples, and any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A combination of protein biomarkers for early screening of esophageal squamous cell carcinoma, characterized in that, The protein biomarker combination consists of at least two proteins selected from the following: GPLD1, CBL, DGKG, and TUBB6.

2. The protein biomarker combination according to claim 1, characterized in that, The protein biomarker combination includes GPLD1, CBL, DGKG, and TUBB6.

3. The protein biomarker combination according to claim 1, characterized in that, In samples from the early stages of esophageal squamous cell carcinoma, compared to samples from precancerous esophageal lesions: The expression level of GPLD1 was downregulated; The expression levels of CBL, DGKG, and TUBB6 were upregulated.

4. The protein biomarker combination according to claim 1, characterized in that, The protein biomarker combination is used to distinguish between the early stage of esophageal squamous cell carcinoma and the precancerous stage of esophageal squamous cell carcinoma. The early stage of esophageal squamous cell carcinoma includes high-grade intraepithelial neoplasia and esophageal squamous cell carcinoma, and the precancerous stage of esophageal squamous cell carcinoma includes esophagitis and low-grade intraepithelial neoplasia. In the serum of patients with early-stage esophageal squamous cell carcinoma, the expression level of GPLD1 was significantly lower than that of patients with precancerous esophageal squamous cell carcinoma; while the expression levels of CBL, DGKG, and TUBB6 were significantly higher than those of patients with precancerous esophageal squamous cell carcinoma.

5. A kit for early screening of esophageal squamous cell carcinoma, characterized in that, The kit contains reagents for detecting the expression levels of the combination of protein biomarkers of any one of claims 1-4 in a biological sample.

6. The reagent kit according to claim 5, characterized in that, The biological sample is serum; the reagents include specific antibodies, antibody fragments, nucleic acid aptamers or affinity ligands against GPLD1, CBL, DGKG and TUBB6; the reagents are suitable for enzyme-linked immunosorbent assay (ELISA), chemiluminescent immunoassay, electrochemiluminescent immunoassay, mass spectrometry multiple reaction monitoring or proximity extension assay.

7. A model construction device for early screening of esophageal squamous cell carcinoma, characterized in that, include: Data input module: used to acquire expression level data of the combination of protein biomarkers as described in any one of claims 1-4 in the biological samples of the subjects; The calculation module is equipped with a classification model built based on the expression levels of the protein biomarker combination, which is used to generate early screening results for esophageal squamous cell carcinoma based on the expression level data. Output module: Used to output the screening results; The classification model is either a Naive Bayes model, a Support Vector Machine model, a Random Forest model, a Logistic Regression model, or an XGBoost model.

8. The model building apparatus according to claim 7, characterized in that, The classification model is a Naive Bayes model, and the area under the receiver operating characteristic curve of the Naive Bayes model is not less than 0.799 when distinguishing between the early stage of esophageal squamous cell carcinoma and esophageal precancerous lesions.

9. A method for screening protein biomarkers for early screening of esophageal squamous cell carcinoma, characterized in that, Includes the following steps: a) Collect longitudinal serum samples from patients with precancerous esophageal lesions, including baseline T1 samples and follow-up T2 samples from the same patient; b) For patients who progressed to the early stage of esophageal squamous cell carcinoma during the follow-up period (progression group) and patients who did not progress after the follow-up period (non-progression group), proteomics technology was used to detect the protein expression levels in their serum at baseline and during the follow-up period. c) Within-group paired tests were performed in patients in the progression group and the non-progression group to screen for differentially expressed proteins that showed significant changes in T2 at the follow-up period relative to T1 at the baseline. The screening criteria were a paired test p value <0.05 and a fold change in expression FC >1.20 or FC <0.

83. d) Differentially expressed proteins selected in the progression group were compared with those selected in the no-progression group. Proteins that also showed significant changes in the no-progression group were excluded, and the remaining proteins were identified as potential early screening biomarkers for esophageal squamous cell carcinoma. e) Use machine learning algorithms to perform feature selection on the potential biomarkers obtained in step d) to obtain the candidate protein biomarkers; f) After evaluating the statistical robustness of the differences in expression levels of the candidate protein biomarkers obtained in step e) between preESCC and eESCC, the combination of protein biomarkers is obtained.

10. A method for constructing an early screening model for esophageal squamous cell carcinoma, characterized in that, Includes the following steps: (a) Obtain expression level data of the protein biomarker combination obtained by the screening method of claim 9 in training samples; the protein biomarker combination consists of at least two proteins selected from the following: GPLD1, CBL, DGKG, TUBB6; (b) Using the expression level of the protein biomarker combination as input features, a classification model is trained using a machine learning algorithm, wherein the machine learning algorithm is selected from Naive Bayes, Support Vector Machine, Random Forest, Logistic Regression, XGBoost, LightGBM, Gradient Boosting Machine or Deep Neural Network; (c) The model performance was evaluated through cross-validation to obtain an early screening model for esophageal squamous cell carcinoma; The area under the receiver operating characteristic curve (AUC) of the classification model in the independent validation set for distinguishing between early-stage esophageal squamous cell carcinoma and precancerous lesions is no less than 0.799.