Diagnostic model for esophageal cancer or esophageal intraepithelial neoplasia and application thereof

By detecting the expression levels of biomarkers FN1, VWF, FBXO34, and ITGA2B, and combining this with an ensemble learning approach to construct a diagnostic model, the problem of early diagnosis of esophageal cancer and esophageal intraepithelial neoplasia was solved, enabling efficient early screening and treatment and improving diagnostic efficacy.

CN120818599AActive Publication Date: 2025-10-21CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510995052.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-21
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Current technologies have low sensitivity and detection rate in the detection of esophageal cancer and esophageal intraepithelial neoplasia, especially for low-grade esophageal intraepithelial neoplasia.

Method used

Four biomarkers, FN1, VWF, FBXO34, and ITGA2B, were used. The expression levels of these biomarkers were detected by techniques such as nucleic acid sequencing, nucleic acid hybridization, chromatography, mass spectrometry, digital imaging, and protein immunoassay. An integrated learning method was used to construct a diagnostic prediction model for the early diagnosis of esophageal cancer and to differentiate esophageal intraepithelial neoplasia.

Benefits of technology

It improves the diagnostic efficacy of esophageal squamous cell carcinoma and esophageal intraepithelial neoplasia, with high sensitivity and specificity, making it suitable for early diagnosis and promoting improved cure rates and prolonged survival rates for patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120818599A_ABST
    Figure CN120818599A_ABST
Patent Text Reader

Abstract

The invention discloses an esophageal cancer or esophageal intraepithelial neoplasia diagnosis model and application thereof, the diagnosis model comprises the following biomarker combination: FN1, VWF, FBXO34 and ITGA2B, and according to the model, whether a subject suffers from esophageal squamous cell carcinoma or esophageal intraepithelial neoplasia can be accurately diagnosed and predicted. The model has excellent diagnosis efficiency and relatively high sensitivity, and is beneficial to further improving a screening method for early diagnosis of esophageal squamous cell carcinoma, so that the improvement of the cure rate and the prolonging of the survival rate of a patient are promoted; effective technical support is provided for early diagnosis, early treatment, prognosis improvement and death rate reduction of esophageal squamous cell carcinoma.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biotechnology, and in particular, relates to a diagnostic model for esophageal cancer or esophageal intraepithelial neoplasia and an application thereof. Background Art

[0002] Esophageal cancer is the eighth most common cancer type worldwide, with the sixth highest mortality rate among all cancer types, and a 5-year survival rate of <20%. Esophageal adenocarcinoma (EAC) and esophageal squamous cell carcinoma (ESCC) are the two main subtypes of esophageal malignancies. ESCC develops due to the malignant transformation of esophageal epithelial cells, making early diagnosis and treatment of ESCC crucial.

[0003] Esophageal intraepithelial neoplasia (low-grade) refers to the morphological manifestation of mild structural disorder of the esophageal mucosa. It is an abnormal proliferation of the esophagus under stimulation such as inflammation. It is temporarily classified as a benign lesion and a precancerous lesion. By detecting the methylation level of esophageal intraepithelial neoplasia (low-grade), the detection rate of esophageal precancerous lesions can be effectively improved. The prior art discloses some reagents for DNA methylation detection and esophageal cancer detection kits, which have good sensitivity and specificity for the detection of esophageal cancer and non-cancerous conditions. However, for precancerous lesions of esophageal cancer, including esophageal intraepithelial neoplasia (high-grade) and esophageal intraepithelial neoplasia (low-grade), especially esophageal intraepithelial neoplasia (low-grade), the detection sensitivity and detection rate are relatively low.

[0004] Therefore, it is extremely important to provide a diagnostic marker and diagnostic model that can diagnose esophageal squamous cell carcinoma and esophageal intraepithelial neoplasia. Summary of the Invention

[0005] In order to make up for the deficiencies of the prior art, the present invention aims to provide a novel diagnostic model for esophageal squamous cell carcinoma or esophageal intraepithelial neoplasia and its application.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A first aspect of the present invention provides a biomarker for early diagnosis of esophageal cancer, diagnosis of esophageal intraepithelial neoplasia, or differentiation between esophageal cancer and esophageal intraepithelial neoplasia.

[0008] Furthermore, the biomarkers include one or more of FN1, VWF, FBXO34, and ITGA2B.

[0009] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0010] In the present invention, the biomarkers include genes and proteins encoded by them and their homologs, mutations, and isoforms. The term encompasses full-length, unprocessed markers, as well as any form of markers derived from processing in cells. The term encompasses naturally occurring variants of the marker (e.g., splice variants or allelic variants).

[0011] The second aspect of the present invention provides the use of a reagent for detecting the expression level of the biomarker described in the first aspect of the present invention in the preparation of a product for diagnosing esophageal cancer, diagnosing esophageal intraepithelial neoplasia, or distinguishing esophageal cancer from esophageal intraepithelial neoplasia.

[0012] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0013] Furthermore, the reagents include reagents for detecting the expression amount of the biomarkers in a sample through nucleic acid sequencing technology, nucleic acid hybridization technology, chromatography technology, mass spectrometry technology, digital imaging technology, protein immunoassay technology, dye technology, and / or second-generation sequencing technology.

[0014] Furthermore, the reagent includes a reagent for detecting the expression level of the biomarker mRNA and / or protein.

[0015] Furthermore, the reagent for detecting the expression level of the biomarker mRNA is a reagent for detecting the level of cDNA complementary to the mRNA transcribed from the biomarker.

[0016] Preferably, the reagent for detecting the expression level of the biomarker mRNA is a primer or a probe.

[0017] Furthermore, the reagent for detecting the expression level of the biomarker protein is a reagent for detecting the level of the polypeptide or protein encoded by the biomarker.

[0018] Preferably, the reagent for detecting the expression amount of the biomarker protein is an antibody, an antibody fragment or an affinity protein.

[0019] Furthermore, the sample of the biomarker is from a tissue sample, primary or cultured cells or cell lines, cell supernatant, cell lysate, platelets, serum, plasma, vitreous humor, lymph fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysate, tissue culture fluid, or tissue extract.

[0020] Furthermore, the sample of the biomarker is derived from a tissue sample, platelets, serum, plasma, whole blood, tissue culture fluid, or tissue extract. In a specific embodiment of the present invention, the sample of the biomarker is derived from plasma.

[0021] Furthermore, the sample of the biomarker is from a human or a non-human mammal.

[0022] Furthermore, the sample of the biomarker is from a human.

[0023] In some specific embodiments, suitable mammals falling within the scope of the present invention include vertebrates, specifically including, but not limited to, any member of the subphylum Chordata, including primates, and including monkey species, rodents, Leporidae, cattle, sheep, goats, pigs, horses, dogs, cats, birds (e.g., chickens, turkeys, ducks, geese, companion birds (such as canaries, budgies, etc.)), marine mammals, reptiles, and fish.

[0024] Furthermore, the primers described in the present invention can be prepared by chemical synthesis methods well known to those skilled in the art, and can be appropriately designed by referring to known information using methods well known to those skilled in the art, and prepared by chemical synthesis.

[0025] Furthermore, the probe described in the present invention can be prepared by chemical synthesis, appropriately designed with reference to known information using methods known to those skilled in the art, and prepared by chemical synthesis, or can be prepared by preparing a gene containing a desired nucleic acid sequence from a biological material and amplifying it using primers designed to amplify the desired nucleic acid sequence.

[0026] In a specific embodiment of the present invention, the reagent includes any reagent that can be used to detect the expression level of any one or more of the biomarkers FN1, VWF, FBXO34 and / or ITGA2B, including but not limited to antibodies, antibody functional fragments, conjugated antibodies or affinity proteins that specifically bind to the protein markers.

[0027] As used herein, "antibody" refers to an immunoglobulin molecule capable of binding to an epitope present on an antigen. The term is intended to encompass not only intact immunoglobulin molecules, such as monoclonal and polyclonal antibodies, but also bispecific antibodies, humanized antibodies, chimeric antibodies, anti-idiopathic (anti-ID) antibodies, single-chain antibodies, Fab fragments, F(ab') fragments, fusion proteins, and any modified forms of the foregoing that contain an antigen recognition site with the desired specificity.

[0028] A third aspect of the present invention provides a product for early diagnosis of esophageal cancer, diagnosis of esophageal intraepithelial neoplasia, or differentiation between esophageal cancer and esophageal intraepithelial neoplasia.

[0029] Furthermore, the product comprises a reagent for detecting the expression level of the biomarker described in the first aspect of the present invention.

[0030] Furthermore, the product is an in vitro diagnostic product.

[0031] Preferably, the in vitro diagnostic product is an in vitro diagnostic kit.

[0032] Furthermore, the reagents include reagents for detecting the expression level of the biomarker mRNA and / or reagents for detecting the expression level of the biomarker protein.

[0033] Preferably, the reagent is a primer, a probe, an antibody, an antibody fragment, and / or an affinity protein.

[0034] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0035] A fourth aspect of the present invention provides a diagnostic prediction model for diagnosing / predicting esophageal cancer or esophageal intraepithelial neoplasia.

[0036] Furthermore, the diagnostic prediction model includes the following biomarker combination: FN1, VWF, FBXO34 and ITGA2B.

[0037] Furthermore, the diagnosis prediction model is constructed using an integrated learning method.

[0038] Furthermore, the ensemble learning method includes linear regression algorithm, support vector machine algorithm, nearest neighbor / k-nearest neighbor algorithm, logistic regression algorithm, decision tree algorithm, k-means algorithm, random forest algorithm, naive Bayes algorithm, dimensionality reduction algorithm, and gradient boosting algorithm.

[0039] Furthermore, the ensemble learning method is a logistic regression algorithm.

[0040] Preferably, the ensemble learning method is an ordered logistic regression algorithm.

[0041] Furthermore, the diagnostic prediction model uses the following formula to diagnose / predict whether a sample has esophageal cancer or esophageal intraepithelial neoplasia or whether it has the risk of esophageal cancer or esophageal intraepithelial neoplasia:

[0042] ; Where n is the number of proteins used for diagnosis prediction, Expi is the expression level of each protein, and Coefi is the regression coefficient of each protein.

[0043] In a specific embodiment of the present invention, the regression coefficient of FN1 is 1.95, the regression coefficient of VWF is 2.06, the regression coefficient of FBXO34 is 0.49, and the regression coefficient of ITGA2B is 0.72.

[0044] Furthermore, in a specific embodiment of the present invention, the diagnostic prediction model uses the following formula to diagnose / predict whether a sample has esophageal cancer or esophageal intraepithelial neoplasia or whether it has the risk of esophageal cancer or esophageal intraepithelial neoplasia:

[0045] .

[0046] Furthermore, in a specific embodiment of the present invention, the diagnostic prediction model obtains diagnostic prediction results according to the following criteria:

[0047] like , the prediction category is esophageal intraepithelial neoplasia; if , the predicted category is esophageal cancer.

[0048] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0049] A fifth aspect of the present invention provides a system or device for diagnosing esophageal cancer or esophageal intraepithelial neoplasia.

[0050] Furthermore, the system or device includes:

[0051] (1) a data acquisition module for acquiring expression profile data of a biomarker combination in a sample of a test subject, wherein the biomarker combination is FN1, VWF, FBXO34, and ITGA2B;

[0052] (2) a diagnostic prediction module, configured to provide the expression profile data of the biomarker combination obtained by the data acquisition module as input data to a trained diagnostic prediction model, wherein the diagnostic prediction model is trained to make predictions for the subject based on the expression profile data of the biomarker combination of the subject;

[0053] (3) A prediction result acquisition module, which is used to obtain the output result of the diagnosis prediction model in the diagnosis prediction module to obtain the prediction result of the subject.

[0054] Preferably, the diagnostic prediction model is the diagnostic prediction model described in the fourth aspect of the present invention.

[0055] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0056] A sixth aspect of the present invention provides a computer device.

[0057] Furthermore, the computer device includes a memory and a processor, the memory stores a program, and the processor implements the following method when executing the program: obtaining biomarker combination expression profile data in a sample of a test subject, wherein the biomarker combination is FN1, VWF, FBXO34, and ITGA2B;

[0058] Providing the biomarker combination expression profile data as input data to a trained diagnostic prediction model;

[0059] Output the diagnostic prediction results of the subject to be tested.

[0060] Preferably, the diagnostic prediction model is the diagnostic prediction model described in the fourth aspect of the present invention.

[0061] Furthermore, the computer devices of the present invention include (but are not limited to) any terminal such as a personal computer, server, or the like that can interact with a user via a keyboard, touchpad, or voice control device. The computing devices herein may also include mobile terminals, which include (but are not limited to) any electronic device that can interact with a user via a keyboard, touchpad, or voice control device, such as tablet computers, smartphones, personal digital assistants (PDAs), smart wearable devices, and the like. The network in which the computing device resides includes (but is not limited to) the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), and the like.

[0062] Furthermore, the memory of the present invention includes non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Operating system code is stored thereon. For example, the memory also stores code or instructions, and by running these codes or instructions, the model for diagnosing and predicting esophageal cancer or precancerous lesions provided by the embodiments disclosed herein can be implemented. Volatile memory may include random access memory (RAM) or external cache memory.

[0063] Furthermore, the computer device of the present invention may include a processor, memory, external interfaces, a display, and an input device connected via a system bus. The processor is used to provide computing and control capabilities. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen layer covering the display, or may be, for example, a keypad, trackball, or touchpad provided on the housing of the computing device, or may be an external keyboard, touchpad, or mouse.

[0064] The processor may include one or more microprocessors or digital processors. The processor may call program code stored in a memory to execute related functions. The processor, also known as a central processing unit (CPU), may be a very large-scale integrated circuit (VLSI) that serves as both the computing core and the control unit.

[0065] A seventh aspect of the present invention provides a computer-readable storage medium.

[0066] Furthermore, the computer-readable storage medium includes a stored computer program.

[0067] Wherein, when the computer program is running, the computer-readable storage medium is controlled to implement the method described in the sixth aspect of the present invention.

[0068] An eighth aspect of the present invention provides the use of a biomarker combination of FN1, VWF, FBXO34, and ITGA2B in constructing a diagnostic prediction model for diagnosing esophageal cancer or esophageal intraepithelial neoplasia.

[0069] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0070] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0071] Compared with the existing technology, the present invention uses FN1, VWF, FBXO34 and ITGA2B for the diagnosis of esophageal squamous cell carcinoma or esophageal intraepithelial neoplasia for the first time, and it has been verified that it has excellent diagnostic efficacy, high sensitivity and specificity, and is particularly suitable for the early diagnosis of esophageal squamous cell carcinoma or esophageal intraepithelial neoplasia. It helps to further improve the screening method for early diagnosis of esophageal squamous cell carcinoma, thereby promoting the improvement of patient cure rate and extension of survival rate, and providing effective technical support for the early diagnosis and early treatment of esophageal squamous cell carcinoma, improving prognosis and reducing mortality. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 Figure 2 is a diagram of research design and proteomic feature analysis; Figure 1 A is an overview of the study cohort and the overall study flow chart; Figure 1 B shows the number of peptides identified in the proteomic data, the number of specific peptides, the total number of proteins, and the number of proteins with a missing value ratio of less than 25%; Figure 1 C is a boxplot of the number of identified proteins in the three groups of samples: HC, EIN, and ESCC; Figure 1 D is the protein data integrity distribution curve; Figure 1 E is the growth curve of the cumulative number of identified proteins as the number of samples increases; Figure 1 F is the ranking of the median intensity of proteins in each group of samples, marking the top ten most abundant proteins in each group and showing their relative contribution to the total protein intensity of the group; Figure 1 G is a bar graph showing the number of identified non-secreted proteins and secreted proteins, as well as a pie chart showing the proportion of secreted proteins classified by function and subcellular localization; Figure 1 H is the classification of the 2120 identified proteins into five categories based on the transcript specificity of esophageal tissue and other tissues, and their distribution is shown; Figure 1I is a histogram of the Pearson correlation coefficient distribution among 150 samples based on the preprocessed proteomics data;

[0073] Figure 2 Figure 2 is a functional protein module diagram related to EIN and ESCC; Figure 2 A: Nine functional protein modules were identified in ESCC samples and 12 functional modules were identified in EIN samples through WGCNA analysis; Figure 2 B shows the association between the enrichment scores of the 21 protein modules and the clinical characteristics of ESCC; Figure 2 C is a visualization showing the distribution of proteins in the nine ESCC-related protein modules; Figure 2 D is a visualization showing the distribution of each protein in the 12 EIN-related protein modules; Figure 2 E is the normalized enrichment score histogram of 21 disease-related modules in the three groups of samples: HC, EIN, and ESCC;

[0074] Figure 3 The results of plasma protein biomarker screening for ESCC diagnosis; Figure 3 A is a two-dimensional visualization of 1,323 proteins in 50 HC, 50 EIN, and 50 ESCC samples using the UMAP method, with each dot representing one sample. Figure 3 B. The left figure is a Venn diagram of differentially expressed proteins between the ESCC vs. HC and EIN vs. HC groups. The middle figure is a scatter plot of the fold difference between the two groups. The right figure shows the average abundance of the 12 differentially expressed proteins in the HC, EIN, and ESCC groups. Figure 3 C is the expression heat map of 12 differentially expressed proteins in the three groups of samples; Figure 3 D is the top ten GO functions (including biological processes, cellular components and molecular functions) and KEGG pathways significantly enriched in the 12 differentially expressed proteins; Figure 3 E The left figure is a Venn diagram of differentially expressed proteins between the ESCC vs. HC and ESCC vs. EIN groups. The middle figure is a scatter plot of the fold difference between the two groups. The right figure shows the average abundance of 12 differentially expressed proteins in the HC, EIN, and ESCC groups. Figure 3 F is the expression heat map of 15 differentially expressed proteins in the three groups of samples; Figure 3 G represents the top ten GO functions (including biological processes, cellular components, and molecular functions) and KEGG pathways that were significantly enriched in the 15 differentially expressed proteins; Figure 3 H is the log2 (concentration) results of four candidate proteins detected by ELISA in CAMS cohort 1;

[0075] Figure 4To evaluate the diagnostic performance of PROSECP for EIN and ESCC in the discovery and validation cohorts; Figure 4 A is the ROC curve of the diagnostic performance of PROSECP in CAMS cohort 1; Figure 4 B is the confusion matrix of PROSECP classification results in CAMS cohort 1; Figure 4 C is a heat map of the expression abundance of four protein biomarkers in CAMS cohort 1; Figure 4 D is the ROC curve of the diagnostic performance of PROSECP in CAMS cohort 2; Figure 4 E is the confusion matrix of PROSECP classification results in CAMS cohort 2; Figure 4 F is the heat map of the expression abundance of four protein biomarkers in CAMS cohort 2; Figure 4 G is a boxplot of the levels of four protein biomarkers in three groups (HC, EIN, and ESCC) of samples in CAMS cohort 2; Figure 4 H is the ROC curve of the diagnostic performance of PROSECP in the FAHZU cohort; Figure 4 I is the confusion matrix of PROSECP classification results in the FAHZU cohort;

[0076] Figure 4 J is the heat map of the expression abundance of four protein biomarkers in CAMS cohort 1; Figure 4 K is the boxplot of the levels of four protein biomarkers in the three groups (HC, EIN, and ESCC) of samples in the FAHZU cohort;

[0077] Figure 5 The diagnostic performance of individual protein biomarkers of FN1, VWF, FBXO34, and ITGA2B; Figure 5 A is the diagnostic efficacy evaluation of these four protein biomarkers for ESCC in CAMS cohort 1, CAMS cohort 2, and FAHZU cohort (ROC curve); Figure 5 B is the diagnostic efficacy evaluation of four protein biomarkers for EIN in CAMS cohort 1, CAMS cohort 2, and FAHZU cohort (ROC curve); Figure 5 C shows the diagnostic efficacy evaluation (ROC curve) of the four protein biomarkers for ESCC and EIN in CAMS cohort 1, CAMS cohort 2, and FAHZU cohort. DETAILED DESCRIPTION

[0078] It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of this application can be combined with each other. In addition, the terms used herein are only used to describe specific embodiments and should not limit the scope of the present invention, as the scope of the present invention is limited to the scope limited by the appended claims. The present invention will be described in detail below with reference to the examples. The experimental methods in the examples where specific conditions are not specified are generally based on conventional conditions or the conditions recommended by the manufacturer.

[0079] The materials and methods used in the examples of the present invention are as follows:

[0080] 1. Study population and clinical sample collection:

[0081] Plasma samples were collected from 153 healthy controls (HC), 116 patients with esophageal intraepithelial neoplasia (EIN), and 206 patients with esophageal squamous cell carcinoma (ESCC) at two medical centers in China. One hundred HC patients, 104 EIN patients, and 106 ESCC patients were enrolled at the Cancer Hospital of the Chinese Academy of Medical Sciences (CAMS) (hereafter referred to as CAMS Cohort 2); and 53 HC patients, 12 EIN patients, and 100 ESCC patients were enrolled at the First Affiliated Hospital of Zhengzhou University (FAHZU) (hereafter referred to as FAHZU Cohort 2). All ESCC patients underwent surgical or endoscopic resection. The final diagnosis of ESCC or EIN was pathologically confirmed. Plasma samples were collected at the time of diagnosis and before surgical intervention or treatment. Healthy controls with no history of malignancy or other major diseases were also included in the study, and plasma samples were collected during routine clinical visits.

[0082] 2. Blood sampling and plasma separation:

[0083] Before any treatment, peripheral blood samples (10 ml per subject) were collected from all subjects using EDTA-K2 tubes (BD, 367525). Plasma was then separated by centrifugation at 800 g for 15 minutes at 4°C within two hours of collection. The supernatant plasma was transferred to a new centrifuge tube and then centrifuged at 12,000 g for 10 minutes at 4°C to remove cellular debris. The resulting plasma was aliquoted and stored at -80°C until further analysis by proteomic analysis or ELISA.

[0084] 3. DIA-MS proteomics workflow:

[0085] DIA-MS proteomic analysis was performed by PTM-BIO (Jingjie Biotechnology Co., Ltd.). A total of 150 plasma samples from the discovery cohort (CAMS Cohort 1) underwent DIA-MS proteomic analysis. High-abundance proteins were removed using the Pierce™ Top 14 High Abundance Protein Removal Kit (Thermo Fisher, A36371), and protein concentrations were determined using the BCA Protein Assay Kit (Beyotime, P0009). Protein solutions were reduced with 5 mM dithiothreitol at 56°C for 30 minutes and then alkylated with 11 mM iodoacetamide for 15 minutes at room temperature in the dark. The alkylated samples were then transferred to ultrafiltration tubes for FASP digestion. Samples were first digested three times with 8 M urea at 12,000 g for 20 minutes at room temperature, followed by three digestions with 200 mM TEAB. Trypsin was added at a trypsin-to-protein ratio of 1:50 for overnight digestion. Peptides were recovered by centrifugation at 12,000 g for 10 minutes at room temperature, repeated twice. The peptides were then desalted using a C18 ZipTip column. Tryptic peptides were dissolved in solvent A and loaded directly onto a homemade reversed-phase analytical column (25 cm length, 100 μm ID). The mobile phase consisted of solvent A (0.1% formic acid, 2% acetonitrile / water) and solvent B (0.1% formic acid, 90% acetonitrile / water). The peptides were separated using the following gradient: 0–1.6 min, 4%–22.5% B; 1.6–2.0 min, 22.5%–35% B; 2.0–2.6 min, 35%–55% B; 2.6–2.7 min, 55%–99% B; 2.7–6.8 min, 99% B; and 6.8–7.6 min, 99% B. All gradients were run on a Vanquish Neo UPLC system (Thermo Fisher Scientific) at a constant flow rate of 300 nL / min. Separated peptides were analyzed using an Orbitrap Astral equipped with a nanoelectrospray ionization source. The electrospray voltage was 1900 V. Precursors were analyzed using the Orbitrap detector, and fragments were analyzed using the Astral detector. Full MS scans were performed at a resolution of 240,000 over a scan range of 380–980 m / z. The first fixed mass of the MS / MS scan was 150.0 m / z at a resolution of 80,000. HCD fragmentation analysis was performed at a normalized collision energy (NCE) of 25%. The automatic gain control (AGC) target was set to 500%, and the maximum injection time was 3 ms.

[0086] 4. Peptide identification and protein quantification:

[0087] DIA data were processed using DIA-NN (v.1.8). Tandem mass spectra were searched against the Homo_sapiens_9606_SP_20230103 database (20,389 entries) linked to a reverse decoy database. Trypsin / P was designated as the cleavage enzyme, with a maximum of one missed cleavage allowed. N-terminal methionine removal and cysteine ​​carbamidomethylation were designated as fixed modifications. The false discovery rate (FDR) was adjusted to less than 1%. The raw LC-MS dataset was first searched against the database and then converted to a matrix containing protein-normalized intensities (raw intensities corrected for sample / batch effects). Normalized intensities (I) were then centrally converted to relative quantification (R). The formula is as follows, where i represents sample and j represents protein: Rij = Iij / Mean(Ij).

[0088] 5. Preprocessing of plasma proteomics data:

[0089] First, proteins with more than 25% missing values ​​were excluded from further analysis. Second, median centering was used to correct for systematic biases introduced by variations in sample preparation or measurement techniques. The remaining missing values ​​were imputed using the k-nearest neighbor (KNN) method with a parameter of k = 10. Protein intensities were then log-2 transformed for subsequent statistical analysis. Finally, the ComBat tool in the R package sva (v3.50.0) was used to mitigate batch effects caused by varying sampling times.

[0090] 6. Weighted gene co-expression network analysis:

[0091] Weighted gene co-expression network analysis (WGCNA) was used to identify key modules in ESCC and EIN samples. Preprocessed protein intensities from 50 ESCC and 50 EIN samples were used for module identification. Soft threshold power was determined by selecting the minimum threshold that resulted in a scale-free R² fit value of 0.85. The parameters for constructing the protein co-expression network were as follows: soft threshold power, minimum module size of 20, and merge cut height of 0.15. The topological overlap matrix (TOM) was calculated using the Pearson correlation function. Pathway enrichment analysis was then performed to functionally annotate the identified modules. Signature genes for each module were used to assess the association between the module and clinical information.

[0092] 7. Stability assessment of protein modules:

[0093] To assess the stability of identified protein modules, we performed data downsampling and reshuffling using the WGCNA process using proteomics data from 50 ESCC and 50 EIN samples. In each iteration, we randomly selected 80% of the ESCC / EIN samples to form a subset for protein module identification. The same WGCNA procedure was used to identify protein modules, including determining soft threshold power and other functional parameters. The stability of protein modules was assessed by calculating their accuracy, defined as the consistency of module clustering across different downsampled datasets. The downsampling and module identification process was repeated 20 times, each using a different random seed to ensure that the results were not affected by a specific data split. The average accuracy for each protein module was calculated across 20 iterations.

[0094] 8. Functional annotation and enrichment analysis:

[0095] Proteins were annotated as secretory and esophageal-specific based on the Human Protein Atlas (HPA, version 19). The atlas predicted 2,520 genes encoding secretory proteins (representing 12% of all human protein-coding genes) and categorized them into ten groups based on literature, bioinformatics, and experimental data from various sources, including the HPA, UniProt, GTEx, and FANTOM. Based on transcriptome analysis of the esophagus and other tissues, the HPA further categorized genes into five groups: "elevated in the esophagus," "elevated in other tissues but expressed in the esophagus," "low tissue specificity but expressed in the esophagus," "not detected in the esophagus," and "not detected in any tissue." GO and KEGG enrichment analysis was performed using the R package clusterProfiler (v4.0.5). GO or KEGG pathways with an FDR-adjusted p-value of less than 0.05 were considered statistically significant.

[0096] 9. ELISA test:

[0097] ELISA assays are performed according to the manufacturer's instructions. Samples and standards are appropriately diluted and added to a 96-well plate coated with antibodies, followed by incubation to allow binding. A biotinylated detection antibody is then added to bind to the target. After incubation and washing, streptavidin-conjugated horseradish peroxidase (HRP) is added to bind to the biotin on the detection antibody. A substrate is introduced, and an enzymatic reaction produces a color change. The reaction is stopped when a clear color gradient appears in the standard curve wells, indicating that the different concentrations of the standard have reacted appropriately. The absorbance is measured using a microplate reader to quantify the concentration of the target molecule based on the standard curve.

[0098] 10. Abundance difference analysis:

[0099] The Kruskal-Wallis test and the Nemenyi nonparametric test were used to identify proteins with differential abundance among the ESCC, EIN, and HC groups. P values ​​were adjusted for multiple comparisons using the FDR method. Proteins with an FDR-adjusted P value less than 0.05 and a change greater than 1.5-fold were considered to have significant differential abundance.

[0100] 11. Identify biomarkers for early diagnosis of ESCC:

[0101] The learning vector quantization (LVQ) model was used to evaluate the ability of individual proteins to distinguish ESCC / EIN patients based on their abundance, and candidate proteins with an accuracy greater than 0.8 were selected for further validation by ELISA. After ELISA validation, proteins that were consistent with the proteomic results were selected as early diagnostic markers for ESCC.

[0102] 12. Model development and performance evaluation:

[0103] An ordinal logistic regression model was developed using ELISA data from CAMS cohort 1 (including 28 HC, 30 EIN, and 30 ESCC samples) to estimate the probability of malignancy for ESCC. The model was externally validated using data from CAMS cohort 2 and the FAHZU cohort. Model performance was evaluated using the area under the receiver operating characteristic (ROC) curve (AUC), area under the precision-recall curve (AUPRC), sensitivity (SEN), specificity (SPE), and confusion matrix. ROC curves and AUCs were visualized and calculated using the R package pROC (v1.18.5). A decision curve analysis (DCA) was also performed to assess the clinical utility of the model.

[0104] Example 1 Study Design and Cohort Characteristics

[0105] The overall workflow of this study and detailed plasma specimen cohort information are as follows Figure 1Figure A. The primary objective of this study was to systematically investigate plasma proteomic alterations and identify potential plasma protein biomarkers for the early diagnosis of ESCC. Therefore, we conducted a multicenter discovery-validation study. We obtained plasma samples from 475 patients from two medical centers in China (CAMS and FAHZU). During the discovery phase, CAMS Cohort 1 comprised 50 HC patients from CAMS, 50 EIN patients, and 50 ESCC patients. Plasma samples from this cohort were subjected to quantitative proteomic analysis using DIA-MS. Validation was performed using two independent cohorts: CAMS Cohort 2 (n = 160, comprising 50 HC patients from CAMS, 54 EIN patients, and 56 ESCC patients) and the FAHZU cohort (n = 165, comprising 53 HC patients from FAHZU, 12 EIN patients, and 100 ESCC patients). Candidate protein biomarkers were screened by proteomic analysis and validated using ELISA in CAMS Cohort 1. A plasma protein-based ESCC risk prediction screening model, named PROSECP, was developed in CAMS cohort 1 and validated using ELISA in two independent cohorts.

[0106] Example 2 Plasma proteome profiles of EIN and ESCC

[0107] We performed plasma proteomic analysis of all individuals in CAMS cohort 1 using the "Blood + DIA-MS" proteomics strategy. We identified a total of 13,703 peptides, 13,344 unique peptides, and 2,120 proteins. An average of 1,525 proteins were identified in the HC population, 1,505 proteins were identified in EIN patients, and 1,525 proteins were identified in ESCC patients ( Figure 1 B). There was no significant difference in overall proteomic coverage among the three groups ( Figure 1 C). Among the proteins identified by DIA-MS, there were 1240 proteins with 100% integrity, 1482 proteins with 75% integrity, and 1516 proteins with 50% integrity ( Figure 1 D). As the sample size increases, the number of detected proteins gradually stabilizes, indicating that the protein detection coverage is deep and the stability is good ( Figure 1 E). Furthermore, the number of identified proteins was not affected by factors such as age, sex, or sampling time, nor was protein abundance affected by these variables. Quantitative protein intensities across all samples spanned eight orders of magnitude, with the top 10 most abundant proteins accounting for approximately 40% of the total plasma protein abundance ( Figure 1F). Based on the annotations of the Human Protein Atlas, 861 of the 2,120 proteins identified were classified as "secreted." Of these, 40.77% were secreted into the blood, 23.58% were secreted into cells and on cell membranes, 10.57% were secreted into the extracellular matrix, and the rest entered other tissues and systems ( Figure 1 G). Based on the esophageal-specific proteome annotation from the Human Protein Atlas, 5.7% of proteins were found to be elevated in the esophagus, 40.4% of proteins were elevated in other tissues but also expressed in the esophagus, and 37.6% of proteins were expressed in the esophagus but had low tissue specificity ( Figure 1 H). To minimize the impact of missing values, we excluded 797 proteins with more than 25% missing values, leaving 1323 proteins for further analysis. The missing values ​​were estimated and classified using the KNN method. The average Pearson correlation coefficient between plasma samples was 0.968 ( Figure 1 I), indicating that the plasma samples have high reproducibility and the mass spectrometry platform has good stability.

[0108] Example 3 Clinically relevant functional protein modules

[0109] To identify clinically relevant protein modules associated with EIN and ESCC, we performed weighted gene co-expression network analysis (WGCNA) on plasma proteomics data of ESCC and EIN samples, respectively. A total of 9 functional modules were identified in ESCC, with module sizes ranging from 46 to 135 proteins ( Figure 2 A, 2C), EIN identified a total of 12 functional modules, with module sizes ranging from 27 to 271 proteins ( Figure 2 A, 2D). Robustness analysis confirmed the stability of these protein modules, as demonstrated by self-validation in CAMS cohort 1. By superimposing protein abundance onto the resulting WGCNA network, we identified six modules—ME05, ME06, ME07, ME11, ME12, and ME19—that were significantly upregulated in ESCC samples ( Figure 2 E). In addition, ME11 and ME21 were significantly downregulated in EIN samples. ME13 and ME16 were significantly upregulated in both EIN and ESCC samples, while ME15 was specifically upregulated in EIN samples ( Figure 2 E).

[0110] We further explored the relationships between these modules and clinical characteristics, including demographics, biochemical features, and tumor biomarkers ( Figure 2B). Specifically, ME05 is upregulated in ESCC and significantly correlated with lymph node metastasis. Proteins in ME05 are primarily enriched in biological processes related to cytoskeletal dynamics (actin remodeling), cell adhesion and migration, and platelet function. These processes are crucial for the ability of tumor cells to shed from their primary site, invade the extracellular matrix, traverse the blood and lymphatic vessels, and ultimately colonize lymph nodes. Another module upregulated in ESCC, ME12, is associated with HER2 receptor positivity, which may indicate increased invasiveness and proliferation. Two modules upregulated in both EIN and ESCC samples, ME13 and ME16, are associated with c-MET receptor positivity. These modules are enriched for functions critical to tumor growth and metastasis, including signal transduction, cell growth, ECM-receptor interactions, and the PI3K-Akt signaling pathway.

[0111] Example 4 Application of plasma proteomic biomarkers in the early diagnosis of ESCC

[0112] Uniform manifold approximation and projection (UMAP) analysis of 1323 proteins showed significant differences among the three groups (HC, EIN, and ESCC), indicating that they had different plasma proteomic profiles ( Figure 3 A). To identify disease-associated changes in the plasma proteome, we performed differential protein abundance analysis comparing plasma samples from patients with ESCC or EIN with those from a HC population. This analysis identified 65 significantly differentially expressed proteins (DEPs), 11 of which were upregulated and 1 was downregulated in ESCC and EIN plasma samples compared with HC samples ( Figure 3 B). The expression profiles of these 12 dysregulated proteins effectively distinguished ESCC and EIN samples from HC samples ( Figure 3 C). Functional enrichment analysis of these dysregulated proteins revealed their involvement in key biological processes, including enzyme activity regulation, lipid metabolism, reactive oxygen species response, oxidative stress, and cell adhesion. In addition, these proteins were associated with molecular functions related to protein and lipid binding ( Figure 3 D). We further compared the plasma proteome profiles of ESCC patients with those of EIN patients and HC patients and identified 15 ESCC-specific proteins ( Figure 3 E). The expression profiles of these 15 ESCC-specific proteins effectively distinguished ESCC in EIN and HC samples ( Figure 3 F). Functional enrichment analysis showed that these ESCC-specific proteins were mainly involved in biological processes such as hemostasis, coagulation, and cell adhesion. In addition, they were also associated with receptor activation, molecular functions related to various binding interactions, and known cancer-related pathways ( Figure 3 G).

[0113] To identify potential protein biomarkers for detecting EIN and ESCC, we used a learning vector quantization (LVQ) model to evaluate the diagnostic performance of 12 disease-related proteins and 15 ESCC-specific proteins. By evaluating the accuracy of each protein in distinguishing EIN, ESCC, and HC, we selected 11 proteins with an accuracy higher than 0.8 as candidate biomarkers. To verify the stability and reproducibility of these 11 candidate biomarkers, we used ELISA to detect the expression levels of these proteins in a randomly selected discovery cohort. Among them, the expression levels of 4 proteins (FN1, VWF, FBXO34, and ITGA2B) detected by ELISA showed good correlation with the results of DIA-MS proteomics detection. Compared with HC, the plasma levels of FN1, VWF, FBXO34, and ITGA2B gradually increased from EIN to ESCC ( Figure 3 H), indicating that elevated their plasma levels were associated with a higher risk of disease, thus validating their predictive value and potential as biomarkers for disease risk prediction.

[0114] Example 5 Establishment and independent validation of a plasma protein-based ESCC risk prediction model

[0115] To evaluate the clinical potential of the identified biomarkers, we developed a plasma protein-based early ESCC risk prediction screening model, PROSECP, which was established based on the expression levels of four protein biomarkers measured by ELISA in the discovery cohort. PROSECP was developed using an ordinal regression method and showed high diagnostic performance in distinguishing EIN and ESCC patients from HC, with an AUC of 0.981 (95% CI 0.958-1.000), a specificity of 92.9%, and a sensitivity of 96.7%. Specifically, PROSECP was able to accurately distinguish EIN from HC with an AUC of 0.935 (95% CI 0.855-1.000), a specificity of 92.9%, and a sensitivity of 93.3%. The method was also able to distinguish ESCC from HC with an AUC of 0.999 (95% CI 0.996-1.000), a specificity of 100%, and a sensitivity of 96.7% ( Figure 4 AC).

[0116] To validate the abnormal elevation of these four protein biomarkers found in the discovery cohort, we performed ELISA testing on plasma samples from two independent validation cohorts (CAMS cohort 2 and FAHZU cohort). Consistent with the proteomic results of the discovery cohort, the expression levels of the four biomarkers were significantly elevated in EIN and ESCC samples compared with HC. In addition, their expression levels were significantly increased in EIN and ESCC samples, highlighting their potential and robustness as biomarkers of disease progression ( Figure 4 F, G, J and K).

[0117] We applied PROSECP to two independent validation cohorts to further assess its robustness and generalizability. In CAMS cohort 2, the AUC of PROSECP was 0.989 (95% CI 0.978-1.000) ( Figure 4 D), in the FAHZU cohort, the AUC was 0.979 (95% CI 0.957-1.000) ( Figure 4 H), indicating that PROSECP maintained robust and high diagnostic performance in distinguishing ESCC and EIN from HC. Notably, PROSECP easily distinguished patients with ESCC or EIN from HC, with AUCs of 0.997 (95% CI 0.992-1.000) and 0.970 (95% CI 0.940-1.000) in CAMS cohort 2, respectively ( Figure 4 D, E), the AUCs of the FAHZU cohort were 0.985 (95% CI 0.966-1.000) and 0.904 (95% CI 0.821-0.987), respectively ( Figure 4 H, I).

[0118] We further evaluated the diagnostic efficacy of these four protein individual biomarkers, and the results are as follows: Figure 5 shown. Figure 5 A is the receiver operating characteristic (ROC) curves of the four protein biomarkers for ESCC in CAMS cohort 1, CAMS cohort 2, and FAHZU cohort. The results showed that the ROC of each biomarker was greater than 0.7, indicating that each biomarker could distinguish healthy subjects from ESCC. Figure 5 B. Figure 5 C shows the receiver operating characteristic (ROC) curves of the four protein biomarkers for EIN and for ESCC and EIN in CAMS cohort 1, CAMS cohort 2, and FAHZU cohort, respectively. Similarly, the ROC of each biomarker was greater than 0.7, indicating that each biomarker could distinguish healthy subjects from EIN and could also significantly distinguish ESCC from EIN.

[0119] Furthermore, decision curve analysis (DCA) showed that PROSECP provided greater net benefit across the threshold probability range for distinguishing patients with ESCC or EIN from HC in all three cohorts, indicating its good clinical utility.

[0120] The above embodiments are only provided for understanding the method and core concept of the present invention. It should be noted that, without departing from the principles of the present invention, a number of improvements and modifications may be made to the present invention by a person skilled in the art, and such improvements and modifications shall fall within the scope of protection of the claims of the present invention.

Claims

1. A biomarker for early diagnosis of esophageal cancer, diagnosis of esophageal intraepithelial neoplasia, or differentiation between esophageal cancer and esophageal intraepithelial neoplasia, characterized in that: The biomarkers include one or more of FN1, VWF, FBXO34, and ITGA2B; Preferably, the esophageal cancer is esophageal squamous cell carcinoma.

2. Use of a reagent for detecting the expression level of the biomarker according to claim 1 in the preparation of a product for diagnosing esophageal cancer, diagnosing esophageal intraepithelial neoplasia, or distinguishing esophageal cancer from esophageal intraepithelial neoplasia.

3. The use according to claim 2, characterized in that The esophageal cancer is esophageal squamous cell carcinoma; Preferably, the reagents include reagents for detecting the expression amount of the biomarker in the sample through nucleic acid sequencing technology, nucleic acid hybridization technology, chromatography technology, mass spectrometry technology, digital imaging technology, protein immunoassay technology, dye technology, and / or second-generation sequencing technology.

4. The use according to claim 3, characterized in that The reagents include reagents for detecting the expression levels of the biomarker mRNA and / or protein; Preferably, the reagent for detecting the expression level of the biomarker mRNA is a reagent for detecting the level of cDNA complementary to the mRNA transcribed from the biomarker; Preferably, the reagent for detecting the expression level of the biomarker mRNA is a primer or a probe; Preferably, the reagent for detecting the expression level of the biomarker protein is a reagent for detecting the level of the polypeptide or protein encoded by the biomarker; Preferably, the reagent for detecting the expression amount of the biomarker protein is an antibody, an antibody fragment or an affinity protein.

5. A product for early diagnosis of esophageal cancer, diagnosis of esophageal intraepithelial neoplasia, or differentiation between esophageal cancer and esophageal intraepithelial neoplasia, characterized in that: The product comprises a reagent for detecting the expression amount of the biomarker according to claim 1; Preferably, the product is an in vitro diagnostic product; More preferably, the in vitro diagnostic product is an in vitro diagnostic kit; Preferably, the reagent includes a reagent for detecting the expression level of the biomarker mRNA and / or a reagent for detecting the expression level of the biomarker protein; More preferably, the reagent is a primer, a probe, an antibody, an antibody fragment, and / or an affinity protein; Preferably, the esophageal cancer is esophageal squamous cell carcinoma.

6. A diagnostic prediction model for diagnosing / predicting esophageal cancer or esophageal intraepithelial neoplasia, characterized in that: The diagnostic prediction model includes the following biomarker combination: FN1, VWF, FBXO34, and ITGA2B; Preferably, the diagnostic prediction model is constructed using an ensemble learning method; Preferably, the ensemble learning method includes linear regression algorithm, support vector machine algorithm, nearest neighbor / k-nearest neighbor algorithm, logistic regression algorithm, decision tree algorithm, k-means algorithm, random forest algorithm, naive Bayes algorithm, dimensionality reduction algorithm, gradient boosting algorithm; Preferably, the ensemble learning method is a logistic regression algorithm; More preferably, the ensemble learning method is an ordered logistic regression algorithm; Preferably, the diagnostic prediction model uses the following formula to diagnose / predict whether a sample is esophageal cancer or esophageal intraepithelial neoplasia or whether it has the risk of esophageal cancer or esophageal intraepithelial neoplasia: ; Where n is the number of proteins used for diagnosis prediction, Expi is the expression level of each protein, and Coefi is the regression coefficient of each protein; Preferably, the esophageal cancer is esophageal squamous cell carcinoma.

7. A system or device for diagnosing esophageal cancer or esophageal intraepithelial neoplasia, characterized in that: The system or device comprises: (1) a data acquisition module for acquiring expression profile data of a biomarker combination in a sample of a test subject, wherein the biomarker combination is FN1, VWF, FBXO34, and ITGA2B; (2) a diagnostic prediction module, configured to provide the expression profile data of the biomarker combination obtained by the data acquisition module as input data to a trained diagnostic prediction model, wherein the diagnostic prediction model is trained to make predictions for the subject based on the expression profile data of the biomarker combination of the subject; (3) a prediction result acquisition module, used to obtain the output result of the diagnosis prediction model in the diagnosis prediction module to obtain the prediction result of the subject; Preferably, the diagnostic prediction model is the diagnostic prediction model according to claim 6; Preferably, the esophageal cancer is esophageal squamous cell carcinoma.

8. A computer device, characterized in that: The computer device includes a memory and a processor, wherein the memory stores a program, and when the processor executes the program, the following method is implemented: obtaining biomarker combination expression profile data in a sample of a test subject, wherein the biomarker combination is FN1, VWF, FBXO34, and ITGA2B; Providing the biomarker combination expression profile data as input data to a trained diagnostic prediction model; Outputting the diagnostic prediction results of the subject to be tested; Preferably, the diagnostic prediction model is the diagnostic prediction model according to claim 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program; Wherein, when the computer program is running, the computer-readable storage medium is controlled to implement the method described in claim 8.

10. Application of a biomarker combination of FN1, VWF, FBXO34, and ITGA2B in developing a diagnostic prediction model for esophageal cancer or esophageal intraepithelial neoplasia; Preferably, the esophageal cancer is esophageal squamous cell carcinoma.

Citation Information

Patent Citations

  • Esophageal cancer molecular markers and application thereof

    CN116287225A

  • Application of platelets and markers thereof in treatment and diagnosis of liver cancer

    CN116445617A

  • Biomarker panels for barrett's esophagus and esophageal adenocarcinoma

    WO2010115077A2