A diagnostic model for esophageal cancer or esophageal intraepithelial neoplasia and application thereof

CN120818599BActive Publication Date: 2026-08-21CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510995052.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2026-08-21
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

现有技术公开了一些用于DNA甲基化检测的试剂及食管癌检测试剂盒,对食管癌与非癌的检测具有良好的灵敏性和特异性,但是对于食管癌的瘤前病变,包括食管上皮内瘤变(高级别)和食管上皮内瘤变(低级别),尤其是对食管上皮内瘤变(低级别)的检测灵敏度和检出率偏低

Benefits of technology

[0071]与现有技术相比,本发明首次采用FN1、VWF、FBXO34和ITGA2B用于E食管鳞状细胞癌或食管上皮内瘤变的诊断,且经验证发现其具有优异的诊断效能,具有较高的敏感性和特异性,尤其适用于食管鳞状细胞癌或食管上皮内瘤变的早期诊断,有助于进一步改进食管鳞状细胞癌早期诊断的筛查方法,进而促进患者治愈率的提高与生存率的延长,为食管鳞状细胞癌的早期诊断和早期治疗、预后的改善和死亡率的降低提供了有效的技术支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120818599B_ABST
    Figure CN120818599B_ABST
Patent Text Reader

Abstract

The application discloses an esophageal cancer or esophageal intraepithelial neoplasia diagnosis model and application thereof, wherein the diagnosis model comprises the following biomarker combination: FN1, VWF, FBXO34 and ITGA2B, and the model can be used for accurately diagnosing and predicting whether a subject is suffering from esophageal squamous cell carcinoma or esophageal intraepithelial neoplasia. The model has excellent diagnosis efficiency and high sensitivity, is helpful to further improve the screening method of early diagnosis of esophageal squamous cell carcinoma, and further promotes the improvement of the cure rate and the prolongation of the survival rate of patients, and provides effective technical support for the early diagnosis and early treatment of esophageal squamous cell carcinoma, the improvement of prognosis and the reduction of mortality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology. Specifically, this invention relates to a diagnostic model for esophageal cancer or esophageal intraepithelial neoplasia and its application. Background Technology

[0002] Esophageal cancer is the eighth most common cancer worldwide, ranking sixth in mortality among all cancer types, with a 5-year survival rate of less than 20%. Esophageal adenocarcinoma (EAC) and esophageal squamous cell carcinoma (ESCC) are two major subtypes of esophageal malignancies. ESCC arises from the malignant transformation of esophageal epithelial cells, making early diagnosis and treatment of esophageal squamous cell carcinoma crucial.

[0003] Esophageal intraepithelial neoplasia (low-grade) refers to a morphological manifestation of mild structural disorder of the esophageal mucosa. It is an abnormal proliferation of the esophagus under the stimulation of inflammation and other factors, and is temporarily considered a benign lesion, but is classified as a precancerous lesion. Detecting the methylation level of esophageal intraepithelial neoplasia (low-grade) can effectively improve the detection rate of precancerous lesions of esophageal cancer. Existing technologies disclose some reagents for DNA methylation detection and esophageal cancer detection kits, which have good sensitivity and specificity for detecting esophageal cancer and non-cancerous lesions. However, for precancerous lesions of esophageal cancer, including esophageal intraepithelial neoplasia (high-grade) and esophageal intraepithelial neoplasia (low-grade), especially for esophageal intraepithelial neoplasia (low-grade), the detection sensitivity and detection rate are relatively low.

[0004] Therefore, it is extremely important to provide a diagnostic biomarker and model that can diagnose esophageal squamous cell carcinoma and esophageal intraepithelial neoplasia. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a novel diagnostic model for esophageal squamous cell carcinoma or esophageal intraepithelial neoplasia and its application.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] The first aspect of the present invention provides a biomarker for the early diagnosis of esophageal cancer, the diagnosis of esophageal intraepithelial neoplasia, or the differentiation between esophageal cancer and esophageal intraepithelial neoplasia.

[0008] Furthermore, the biomarkers include one or more of FN1, VWF, FBXO34, and ITGA2B.

[0009] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0010] In this invention, the biomarkers include genes and the proteins they encode, as well as their homologs, mutations, and isotypes. The term encompasses full-length, unprocessed biomarkers, and any form of biomarker derived from cell processing. The term also encompasses naturally occurring variants of the biomarker (e.g., splice variants or allelic variants).

[0011] The second aspect of the present invention provides the use of reagents for detecting the expression levels of the biomarkers described in the first aspect of the present invention in the preparation of products for diagnosing esophageal cancer, diagnosing esophageal intraepithelial neoplasia, or differentiating between esophageal cancer and esophageal intraepithelial neoplasia.

[0012] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0013] Furthermore, the reagents include those for detecting the expression levels of the biomarkers in a sample using nucleic acid sequencing, nucleic acid hybridization, chromatography, mass spectrometry, digital imaging, protein immunoassay, dye technology, and / or next-generation sequencing.

[0014] Furthermore, the reagents include reagents for detecting the expression levels of the biomarker mRNA and / or protein.

[0015] Furthermore, the reagent for detecting the expression level of the biomarker mRNA is a reagent for detecting the level of cDNA complementary to the mRNA transcribed from the biomarker.

[0016] Preferably, the reagent used to detect the expression level of the biomarker mRNA is a primer or a probe.

[0017] Furthermore, the reagent used to detect the expression level of the biomarker protein is a reagent used to detect the level of the polypeptide or protein encoded by the biomarker.

[0018] Preferably, the reagent used to detect the expression level of the biomarker protein is an antibody, antibody fragment, or affinity protein.

[0019] Furthermore, the biomarker samples are derived from tissue samples, primary or cultured cells or cell lines, cell supernatants, cell lysates, platelets, serum, plasma, vitreous fluid, lymph, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysates, tissue culture fluid, or tissue extracts.

[0020] Furthermore, the biomarker samples are derived from tissue samples, platelets, serum, plasma, whole blood, tissue culture medium, or tissue extracts. In a specific embodiment of the present invention, the biomarker samples are derived from plasma.

[0021] Furthermore, the samples of the biomarkers are derived from humans or non-human mammals.

[0022] Furthermore, the samples of the biomarkers were derived from humans.

[0023] In some specific embodiments, suitable mammals falling within the scope of this invention include vertebrates, specifically including but not limited to any member of the subphylum Chordata, including primates, as well as monkey species, rodents, rabbits, cattle, sheep, goats, pigs, horses, dogs, felines, birds (e.g., chickens, turkeys, ducks, geese, companion birds (e.g., canaries, budgerigars, etc.)), marine mammals, reptiles, and fish.

[0024] Furthermore, the primers described in this invention can be prepared by chemical synthesis methods well known to those skilled in the art, appropriately designed using methods well known to those skilled in the art and with reference to known information, and prepared by chemical synthesis.

[0025] Furthermore, the probes described in this invention can be prepared by chemical synthesis, by appropriately designing them with reference to known information using methods known to those skilled in the art, and by preparing them by chemical synthesis, or by preparing a gene containing the desired nucleic acid sequence from biological material and amplifying it using primers designed for amplifying the desired nucleic acid sequence.

[0026] In a specific embodiment of the present invention, the reagent includes any reagent capable of detecting the expression level of any one or more proteins among the biomarkers FN1, VWF, FBXO34 and / or ITGA2B, including but not limited to antibodies, antibody functional fragments, conjugated antibodies or affinity proteins that specifically bind to the protein biomarkers.

[0027] The term "antibody" as used in this invention refers to an immunoglobulin molecule capable of binding to an epitope present on an antigen. The term is intended to encompass not only intact immunoglobulin molecules, such as monoclonal and polyclonal antibodies, but also bispecific antibodies, humanized antibodies, chimeric antibodies, anti-idiopathic (anti-ID) antibodies, single-chain antibodies, Fab fragments, F(ab') fragments, fusion proteins, and any modified forms of the foregoing containing an antigen recognition site with desired specificity.

[0028] A third aspect of the present invention provides a product for the early diagnosis of esophageal cancer, the diagnosis of esophageal intraepithelial neoplasia, or the differentiation between esophageal cancer and esophageal intraepithelial neoplasia.

[0029] Furthermore, the product includes reagents for detecting the expression levels of the biomarkers described in the first aspect of the present invention.

[0030] Furthermore, the product in question is an in vitro diagnostic product.

[0031] Preferably, the in vitro diagnostic product is an in vitro diagnostic kit.

[0032] Furthermore, the reagents include reagents for detecting the mRNA expression level of the biomarker and / or reagents for detecting the protein expression level of the biomarker.

[0033] Preferably, the reagent is a primer, probe, antibody, antibody fragment, and / or affinity protein.

[0034] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0035] The fourth aspect of the present invention provides a diagnostic prediction model for diagnosing / predicting esophageal cancer or esophageal intraepithelial neoplasia.

[0036] Furthermore, the diagnostic prediction model includes the following combination of biomarkers: FN1, VWF, FBXO34, and ITGA2B.

[0037] Furthermore, the diagnostic prediction model is constructed using an ensemble learning method.

[0038] Furthermore, the ensemble learning methods include linear regression, support vector machine, nearest neighbor / k-nearest neighbor, logistic regression, decision tree, k-average, random forest, naive Bayes, dimensionality reduction, and gradient boosting.

[0039] Furthermore, the ensemble learning method is the logistic regression algorithm.

[0040] Preferably, the ensemble learning method is an ordered logistic regression algorithm.

[0041] Furthermore, the diagnostic prediction model uses the following formula to diagnose / predict whether a sample has esophageal cancer, esophageal intraepithelial neoplasia, or a risk of esophageal cancer or esophageal intraepithelial neoplasia:

[0042] Where n is the number of proteins used for diagnostic prediction, Expi is the expression level of each protein, and Coefi is the regression coefficient of each protein.

[0043] In a specific embodiment of the present invention, the regression coefficient of FN1 is 1.95, the regression coefficient of VWF is 2.06, the regression coefficient of FBXO34 is 0.49, and the regression coefficient of ITGA2B is 0.72.

[0044] Furthermore, in a specific embodiment of the present invention, the diagnostic prediction model uses the following formula to diagnose / predict whether a sample is esophageal cancer, esophageal intraepithelial neoplasia, or whether it has a risk of esophageal cancer or esophageal intraepithelial neoplasia:

[0045] .

[0046] Furthermore, in a specific embodiment of the present invention, the diagnostic prediction model obtains diagnostic prediction results through the following criteria:

[0047] like The predicted category is esophageal intraepithelial neoplasia; if The predicted category is esophageal cancer.

[0048] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0049] The fifth aspect of the present invention provides a system or apparatus for diagnosing esophageal cancer or esophageal intraepithelial neoplasia.

[0050] Furthermore, the system or apparatus includes:

[0051] (1) Data acquisition module, used to acquire expression profile data of biomarker combination in the sample of the subject to be tested, wherein the biomarker combination is FN1, VWF, FBXO34 and ITGA2B;

[0052] (2) Diagnostic prediction module, used to provide the expression profile data of the biomarker combination obtained by the data acquisition module as input data to the trained diagnostic prediction model, wherein the diagnostic prediction model is trained to predict the subject based on the expression profile data of the subject's biomarker combination.

[0053] (3) Prediction result acquisition module, used to acquire the output result of the diagnostic prediction model in the diagnostic prediction module and obtain the prediction result of the subject.

[0054] Preferably, the diagnostic prediction model is the diagnostic prediction model described in the fourth aspect of the present invention.

[0055] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0056] The sixth aspect of the present invention provides a computer device.

[0057] Furthermore, the computer device includes a memory and a processor. The memory stores a program, and when the processor executes the program, it implements the following method: acquiring biomarker combination expression profile data in a sample of a subject to be tested, wherein the biomarker combination is FN1, VWF, FBXO34 and ITGA2B;

[0058] The biomarker combination expression profile data is used as input data to the trained diagnostic prediction model;

[0059] Output the diagnostic prediction results for the test subjects.

[0060] Preferably, the diagnostic prediction model is the diagnostic prediction model described in the fourth aspect of the present invention.

[0061] Furthermore, the computer device of the present invention includes (but is not limited to) any terminal such as a personal computer or server capable of human-computer interaction with a user via a keyboard, touchpad, or voice control device. The computing device described herein may also include a mobile terminal, which includes (but is not limited to) any electronic device capable of human-computer interaction with a user via a keyboard, touchpad, or voice control device, such as a tablet computer, smartphone, personal digital assistant (PDA), smart wearable device, etc. The network in which the computing device operates includes (but is not limited to) the Internet, wide area network (WAN), metropolitan area network (MAN), local area network (LAN), virtual private network (VPN), etc.

[0062] Furthermore, the memory of the present invention includes non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Operating system code is stored thereon. For example, code or instructions are also stored on the memory, by which the esophageal cancer or precancerous lesion diagnostic prediction model provided in the embodiments disclosed herein can be implemented. Volatile memory may include random access memory (RAM) or external cache memory.

[0063] Furthermore, the computer device of the present invention may include a processor, a memory, an external interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touch layer covering the display screen, or, for example, buttons, a trackball, or a touchpad mounted on the casing of the computing device, or an external keyboard, touchpad, or mouse.

[0064] The processor may include one or more microprocessors or digital processors. The processor can call program code stored in memory to execute related functions. The processor, also known as a central processing unit (CPU), can be a very large-scale integrated circuit, serving as both the computational core and the control unit.

[0065] A seventh aspect of the present invention provides a computer-readable storage medium.

[0066] Furthermore, the computer-readable storage medium includes a stored computer program.

[0067] The computer program controls the computer-readable storage medium to implement the method described in the sixth aspect of the present invention during runtime.

[0068] The eighth aspect of the present invention provides the application of a combination of biomarkers FN1, VWF, FBXO34 and ITGA2B in constructing a diagnostic prediction model for esophageal cancer or esophageal intraepithelial neoplasia.

[0069] Furthermore, the esophageal cancer is esophageal squamous cell carcinoma.

[0070] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0071] Compared with existing technologies, this invention is the first to use FN1, VWF, FBXO34, and ITGA2B for the diagnosis of esophageal squamous cell carcinoma or esophageal intraepithelial neoplasia. Verification has shown that it has excellent diagnostic efficacy, with high sensitivity and specificity, and is particularly suitable for the early diagnosis of esophageal squamous cell carcinoma or esophageal intraepithelial neoplasia. This helps to further improve screening methods for the early diagnosis of esophageal squamous cell carcinoma, thereby promoting increased cure rates and prolonged survival. It provides effective technical support for the early diagnosis and treatment of esophageal squamous cell carcinoma, improved prognosis, and reduced mortality. Attached Figure Description

[0072] Figure 1 To study the design and proteomics characterization analysis diagrams; among which, Figure 1 A is an overview of the research cohort and a flowchart of the overall research process; Figure 1 B shows the number of peptides identified in the proteomic data, the number of specific peptides, the total number of proteins, and the number of proteins with a missing value ratio of less than 25%. Figure 1 C is a box plot showing the number of proteins identified in the three groups of samples: HC, EIN, and ESCC. Figure 1 D represents the protein data integrity distribution curve; Figure 1 E represents the growth curve of the cumulative number of proteins identified as the number of samples increases; Figure 1 F represents the ranking of median protein intensity in each group of samples, labeling the top ten most abundant proteins in each group and showing their relative contribution to the total protein intensity in that group. Figure 1 G is a bar chart showing the number of identified non-secretory proteins and secretory proteins, as well as a pie chart showing the proportion of secretory proteins classified by function and subcellular location. Figure 1 H represents the classification of 2120 identified proteins into five categories based on transcript specificity in esophageal and other tissues, and their distribution is shown. Figure 1I is a histogram of the distribution of Pearson correlation coefficients among 150 samples based on preprocessed proteomics data;

[0073] Figure 2 This is a diagram of functional protein modules related to EIN and ESCC; among them, Figure 2 A represents the identification of 9 functional protein modules in the ESCC sample and 12 functional modules in the EIN sample by WGCNA analysis. Figure 2 B represents the association between the enrichment scores of the 21 protein modules and the clinical characteristics of ESCC. Figure 2 C is a visualization showing the distribution of each protein in the nine ESCC-related protein modules; Figure 2 D is a visualization showing the distribution of each protein in the 12 EIN-related protein modules; Figure 2 E is a bar chart of the normalized enrichment scores of 21 disease-related modules in the three groups of samples: HC, EIN and ESCC.

[0074] Figure 3 The results of plasma protein biomarker screening for ESCC diagnosis; among which, Figure 3 A is a two-dimensional visualization of 1,323 proteins from 50 HC, 50 EIN, and 50 ESCC samples using the UMAP method, with each point representing one sample; Figure 3 The left figure is a Venn plot of differentially expressed proteins between the ESCC vs. HC and EIN vs. HC groups; the middle figure is a scatter plot of the fold differences between the two groups; and the right figure shows the average abundance of 12 differentially expressed proteins in the HC, EIN and ESCC groups. Figure 3 C is a heatmap of the expression of 12 differentially expressed proteins in the three groups of samples; Figure 3 D represents the top ten GO functions (including biological processes, cellular components, and molecular functions) and the KEGG pathway that were significantly enriched among the 12 differentially expressed proteins. Figure 3 The left figure is a Venn plot of differentially expressed proteins between the ESCC vs. HC and ESCC vs. EIN groups; the middle figure is a scatter plot of the fold differences between the two groups; and the right figure shows the average abundance of 12 differentially expressed proteins in the HC, EIN and ESCC groups. Figure 3 F is a heatmap of the expression of 15 differentially expressed proteins in the three groups of samples; Figure 3 G represents the top ten GO functions (including biological processes, cellular components, and molecular functions) and the KEGG pathway, which were significantly enriched among 15 differentially expressed proteins. Figure 3 H is the log2 (concentration) result of detecting four candidate proteins using the ELISA method in CAMS cohort 1;

[0075] Figure 4To evaluate the diagnostic performance of PROSECP for EIN and ESCC in the discovery and verification queues; where, Figure 4 A is the ROC curve of the diagnostic performance of PROSECP in CAMS cohort 1; Figure 4 B is the confusion matrix of the PROSECP classification results in CAMS cohort 1; Figure 4 C is a heatmap of the expression abundance of four protein biomarkers in CAMS cohort 1; Figure 4 D is the ROC curve of the diagnostic performance of PROSECP in CAMS cohort 2; Figure 4 E is the confusion matrix of the PROSECP classification results in CAMS cohort 2; Figure 4 F is a heatmap of the expression abundance of four protein biomarkers in CAMS cohort 2; Figure 4 G is a box plot of the levels of four protein biomarkers in three groups (HC, EIN, ESCC) of samples in CAMS cohort 2. Figure 4 H represents the ROC curve of PROSECP diagnostic performance in the FAHZU queue; Figure 4 I is the confusion matrix of the PROSECP classification results in the FAHZU queue;

[0076] Figure 4 J is a heatmap of the expression abundance of four protein biomarkers in CAMS cohort 1; Figure 4 K is a box plot of the levels of four protein biomarkers in three groups (HC, EIN, ESCC) of samples in the FAHZU cohort;

[0077] Figure 5 The diagnostic performance of individual protein biomarkers FN1, VWF, FBXO34, and ITGA2B; among them, Figure 5 A represents the diagnostic efficacy (ROC curve) of these four protein biomarkers for ESCC in the CAMS cohort 1, CAMS cohort 2, and FAHZU cohorts. Figure 5 B represents the diagnostic efficacy (ROC curve) of four protein biomarkers for EIN in the CAMS cohort 1, CAMS cohort 2, and FAHZU cohorts. Figure 5 C represents the diagnostic efficacy (ROC curve) of four protein biomarkers for ESCC and EIN in the CAMScohort 1, CAMScohort 2, and FAHZU cohorts. Detailed Implementation

[0078] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. Furthermore, the terminology used herein is for the purpose of describing specific implementations only and should not limit the scope of the invention, as the scope of the invention is limited only to the scope defined by the appended claims. The invention will now be described in detail with reference to embodiments. Experimental methods not specifically described in the embodiments are generally performed under conventional conditions or as recommended by the manufacturer.

[0079] The materials and methods used in the embodiments of this invention are as follows:

[0080] 1. Study population and clinical sample collection:

[0081] Plasma samples were collected from 153 healthy controls (HC) patients, 116 patients with esophageal intraepithelial neoplasia (EIN), and 206 patients with esophageal squamous cell carcinoma (ESCC) at two medical centers in China. The Chinese Academy of Medical Sciences (CAMS) cohort included 100 HC patients, 104 EIN patients, and 106 ESCC patients (CAMS cohort 2); the First Affiliated Hospital of Zhengzhou University (FAHZU) cohort included 53 HC patients, 12 EIN patients, and 100 ESCC patients (FAHZU cohort 2). All ESCC patients underwent surgical or endoscopic resection. The final diagnosis of ESCC or EIN was pathologically confirmed. Plasma samples were collected at the time of diagnosis and before surgical intervention or treatment. Healthy controls with no history of malignant tumors or other major diseases were also included in the study, and plasma samples were collected during routine clinical visits.

[0082] 2. Blood sampling and plasma separation:

[0083] Prior to any treatment, peripheral blood samples (10 mL per subject) were collected from all subjects using EDTA-K2 tubes (BD, 367525). Plasma was then separated by centrifugation at 800 g for 15 minutes at 4°C within two hours of collection. The supernatant plasma was transferred to a new centrifuge tube and centrifuged at 12,000 g for 10 minutes at 4°C to remove cell debris. The resulting plasma was aliquoted and stored at -80°C until further use for proteomic analysis or ELISA.

[0084] 3. DIA-MS proteomics workflow:

[0085] DIA-MS proteomics analysis was performed by PTM-BIO (Jingjie Biotechnology Co., Ltd.). In this analysis, 150 plasma samples from the discovery cohort (CAMS cohort 1) underwent DIA-MS proteomics analysis. High-abundance proteins were removed using the Pierce™ Top 14 High Abundance Protein Removal Kit (Thermo Fisher, A36371), and protein concentrations were determined using the BCA Protein Detection Kit (Beyotime, P0009). Protein solutions were reduced with 5 mM dithiothreitol at 56°C for 30 min, followed by alkylation with 11 mM iodoacetamide at room temperature in the dark for 15 min. The alkylated samples were then transferred to ultrafiltration tubes for FASP digestion. Samples were first digested three times at 12,000 g for 20 min each time with 8 M urea at room temperature, followed by three digestions with 200 mM TEAB. Trypsin was added at a trypsin-to-protein mass ratio of 1:50 for overnight digestion. Peptides were recovered by centrifugation at 12,000 g for 10 min at room temperature, repeated twice. Desalting was then performed using a C18 ZipTip column. The trypsin peptide was dissolved in solvent A and directly packed into a self-made reversed-phase analytical column (25 cm long, 100 μm id). The mobile phase consisted of solvent A (0.1% formic acid, 2% acetonitrile / water) and solvent B (0.1% formic acid, 90% acetonitrile / water). The peptide separation gradients were as follows: 0–1.6 min, 4%–22.5% B; 1.6–2.0 min, 22.5%–35% B; 2.0–2.6 min, 35%–55% B; 2.6–2.7 min, 55%–99% B; 2.7–6.8 min, 99% B; 6.8–7.6 min, 99% B. All gradients were performed at a constant flow rate of 300 nl / min on a Vanquish Neo UPLC system (Thermo Fisher Scientific). The isolated peptides were analyzed using an Orbitrap Astral equipped with a nanoelectrospray ionization source. The electrospray voltage was 1900 V. The precursor was analyzed using an Orbitrap detector, and the fragments were analyzed using an Astral detector. The full MS scan resolution was set to 240,000, with a scan range of 380–980 m / z. The first fixed mass for the MS / MS scan was 150.0 m / z, with a resolution of 80,000. HCD fragmentation analysis was performed at a normalized collision energy (NCE) of 25%. The automatic gain control (AGC) target was set to 500%, and the maximum injection time was 3 ms.

[0086] 4. Peptide identification and protein quantification:

[0087] DIA data was processed using DIA-NN (v.1.8). Tandem mass spectrometry was searched against the Homo_sapiens_9606_SP_20230103 database (20,389 entries) linked to the reverse decoy database. Trypsin / P was designated as the lysin, allowing a maximum of one missed cut. N-terminal methionine removal and cysteine ​​aminomethylation were designated as fixed modifications. The false detection rate (FDR) was adjusted to less than 1%. The raw LC-MS dataset was first searched against the control database and then converted into a matrix containing normalized protein intensities (raw intensities after correcting for sample / batch effects). The normalized intensities (I) were then converted to relative quantifications (R) after centralization. The calculation formula is as follows, where i represents the sample and j represents the protein: Rij = Iij / Mean(Ij).

[0088] 5. Preprocessing of plasma proteomics data:

[0089] First, proteins with missing values ​​exceeding 25% were excluded in further analysis. Second, the median centering method was used to correct for systematic biases that might arise from variations in sample preparation or measurement techniques. The remaining missing values ​​were estimated using the k-nearest neighbor (KNN) method with parameter set to k=10. Then, the protein intensities were log-2 transformed for subsequent statistical analysis. Finally, the ComBat tool in the R package sva (v3.50.0) was used to mitigate batch effects caused by different sampling times.

[0090] 6. Weighted gene co-expression network analysis:

[0091] Weighted Gene Co-expression Network (WGCNA) analysis was used to identify key modules in ESCC and EIN samples. Preprocessed protein intensities from 50 ESCC and 50 EIN samples were used for module identification. Soft threshold power was determined by selecting the minimum threshold that allowed for a scale-free R-squared fit of 0.85. The parameters for constructing the protein co-expression network were as follows: soft threshold power, minimum module size of 20, and merge cut height of 0.15. The topological overlap matrix (TOM) was calculated using the Pearson correlation function. Pathway enrichment analysis was then performed, and the identified modules were functionally annotated. Characteristic genes for each module were used to assess the association between the module and clinical information.

[0092] 7. Stability assessment of the protein module:

[0093] To assess the stability of the identified protein modules, we downsampled and recombined the WGCNA process using proteomics data from 50 ESCC and 50 EIN samples. In each iteration, we randomly selected 80% of the ESCC / EIN samples to form a subset for protein module identification. The same WGCNA procedure was used to identify the protein modules, including determining soft-threshold power and other functional parameters. The stability of the protein modules was assessed by calculating their accuracy, defined as the consistency of module clustering across different downsampled datasets. The downsampling and module identification process was repeated 20 times, each time using a different random seed to ensure that the results were not affected by specific data splits. The mean accuracy for each protein module was calculated over 20 iterations.

[0094] 8. Functional annotation and enrichment analysis:

[0095] Proteins were labeled as secretory and esophageal specific based on the Human Protein Atlas (HPA, version 19), which predicted 2520 genes encoding secretory proteins (representing 12% of all human protein-coding genes) and categorized into ten classes based on literature, bioinformatics, and experimental data from various sources, including HPA, UniProt, GTEx, and FANTOM. Based on transcriptomic analysis of the esophagus and other tissues, HPA further categorized genes into five groups: "elevated in the esophagus," "elevated in other tissues but expressed in the esophagus," "low tissue specificity but expressed in the esophagus," "not detected in the esophagus," and "not detected in any tissue." GO and KEGG enrichment analyses were performed using the R package clusterProfiler (v4.0.5). GO or KEGG pathways with an FDR-adjusted p-value less than 0.05 were considered statistically significant.

[0096] 9. ELISA test:

[0097] ELISA assays were performed according to the manufacturer's instructions. Samples and standards, appropriately diluted, were added to antibody-coated 96-well plates and incubated to allow binding. Biotin-labeled detection antibodies were then added to bind the target analyte. After incubation and washing, streptavidin-conjugated horseradish peroxidase (HRP) was added, which binds to the biotin on the detection antibody. The substrate was introduced, resulting in a color change via an enzymatic reaction. The reaction stopped when a clear color gradient was observed in the wells of the standard curve, indicating that different concentrations of the standard reacted appropriately. Absorbance was measured using a microplate reader, and the concentration of the target molecule was quantified according to the standard curve.

[0098] 10. Abundance Difference Analysis:

[0099] The Kruskal-Wallis test and Nemenyi nonparametric test were used to identify proteins with varying abundances among the ESCC, EIN, and HC groups. Multiple comparison adjustment (FDR) was performed on p-values. Proteins with FDR-adjusted p-values ​​less than 0.05 and changes greater than 1.5-fold were considered to have significant differential abundance.

[0100] 11. Identify biomarkers for early diagnosis of ESCC:

[0101] The ability of individual proteins to differentiate between ESCC / EIN patients based on varying abundance was evaluated using a learned vector quantization (LVQ) model. Candidate proteins with an accuracy greater than 0.8 were selected for further validation via ELISA. After ELISA validation, proteins consistent with the proteomic results were selected as early diagnostic biomarkers for ESCC.

[0102] 12. Model Development and Performance Evaluation:

[0103] Using ELISA data from CAMS cohort 1 (including 28 HC, 30 EIN, and 30 ESCC samples), an ordered logistic regression model was built to estimate the malignancy risk of ESCC. The model was externally validated using data from CAMS cohort 2 and the FAHZU cohort. The model's performance was evaluated using the area under the receiver operating characteristic (ROC) curve (AUC), precision-recall curve (AUPRC), sensitivity (SEN), specificity (SPE), and confusion matrix. ROC curves and AUC were visualized and calculated using the R package pROC (v1.18.5). Decision curve analysis (DCA) was also performed to assess the model's clinical applicability.

[0104] Example 1: Study Design and Queue Characteristics

[0105] The overall workflow and detailed plasma sample cohort information of this study are as follows: Figure 1As shown in Figure A. The main objective of this study was to systematically investigate plasma proteomic alterations and identify potential protein biomarkers in plasma for the early diagnosis of ESCC. Therefore, we conducted a multicenter discovery-validation study. We obtained plasma samples from 475 patients from two medical centers in China (CAMS and FAHZU). In the discovery phase, CAMS cohort 1 included 50 patients with HC, 50 patients with EIN, and 50 patients with ESCC from CAMS. Plasma samples from this cohort underwent quantitative proteomic analysis using DIA-MS. Validation was performed using two independent cohorts: CAMS cohort 2 (n=160, including 50 patients with HC, 54 patients with EIN, and 56 patients with ESCC from CAMS) and the FAHZU cohort (n=165, including 53 patients with HC, 12 patients with EIN, and 100 patients with ESCC from FAHZU). Candidate protein biomarkers were screened by proteomic analysis and validated using ELISA in CAMS cohort 1. A plasma protein-based ESCC risk prediction screening model, named PROSECP, was established in CAMS cohort 1 and validated using ELISA in two separate cohorts.

[0106] Example 2 Plasma proteomic mapping of EIN and ESCC

[0107] We performed plasma proteomics analysis on all individuals in CAMS cohort 1 using a "Blood+DIA-MS" proteomics strategy. We identified a total of 13,703 peptides, 13,344 unique peptides, and 2,120 proteins. The average number of proteins identified was 1,525 in the HC population, 1,505 in EIN patients, and 1,525 in ESCC patients. Figure 1 B). There was no significant difference in overall proteomics coverage among the three groups. Figure 1 C). Of the proteins identified by DIA-MS, 1240 were 100% intact, 1482 were 75% intact, and 1516 were 50% intact. Figure 1 D). As the sample size increases, the number of detected proteins gradually stabilizes, indicating that the protein detection coverage is deep and the stability is good. Figure 1 E). Furthermore, the number of proteins identified was not affected by factors such as age, sex, or sampling time, nor was protein abundance affected by these variables. The quantitative protein intensity ranged across all samples by eight orders of magnitude, with the top 10 most abundant proteins accounting for approximately 40% of the total plasma protein abundance (E). Figure 1F). According to the annotation of the Human Protein Atlas, 861 of the 2120 identified proteins were classified as "secretory." Of these, 40.77% were secreted into the bloodstream, 23.58% into cells and on the cell membrane, 10.57% into the extracellular matrix, and the remainder entered other tissues and systems. Figure 1 G). Based on the esophageal-specific proteome annotation from the Human Protein Atlas, 5.7% of the proteins were elevated in the esophagus, 40.4% were elevated in other tissues but also expressed in the esophagus, and 37.6% were expressed in the esophagus but with low tissue specificity. Figure 1 H). To minimize the impact of missing values, we excluded 797 proteins with missing values ​​exceeding 25%, leaving 1323 proteins for further analysis. Missing values ​​were estimated and categorized using the KNN method. The average Pearson correlation coefficient among plasma samples was 0.968 (H). Figure 1 I) indicates that the plasma samples have high repeatability and the mass spectrometry platform has good stability.

[0108] Example 3: Clinically Relevant Functional Protein Modules

[0109] To identify clinically relevant protein modules associated with EIN and ESCC, we performed weighted gene co-expression network (WGCNA) analysis on plasma proteomics data from ESCC and EIN samples. Nine functional modules were identified in ESCC, ranging in size from 46 to 135 proteins. Figure 2 A, 2C), EIN identified a total of 12 functional modules, with module sizes ranging from 27 to 271 proteins. Figure 2 A, 2D). Robustness analysis confirmed the stability of these protein modules, as evidenced by the self-validation of CAMS cohort 1. By superimposing protein abundance onto the resulting WGCNA network, we identified six modules—ME05, ME06, ME07, ME11, ME12, and ME19—that were significantly upregulated in the ESCC samples ( Figure 2 E). Furthermore, ME11 and ME21 were significantly downregulated in the EIN sample. ME13 and ME16 were significantly upregulated in both the EIN and ESCC samples, while ME15 was specifically upregulated in the EIN sample (E). Figure 2 E).

[0110] We further explored the relationship between these modules and clinical characteristics, including demographics, biochemical features, and tumor biomarkers. Figure 2(B) Specifically, ME05 expression was upregulated in ESCC and significantly associated with lymph node metastasis. Proteins in ME05 are primarily enriched in biological processes related to cytoskeleton dynamics (actin remodeling), cell adhesion and migration, and platelet function. These processes are crucial for the ability of tumor cells to detach from their primary site, invade the extracellular matrix, cross blood and lymphatic vessels, and ultimately colonize lymph nodes. Another upregulated module in ESCC, ME12, is associated with HER2 receptor positivity, which may indicate higher invasiveness and proliferative capacity. Two modules, ME13 and ME16, which were upregulated in both EIN and ESCC samples, are associated with C-MET receptor positivity. These modules are rich in functions crucial for tumor growth and metastasis, including signal transduction, cell growth, ECM-receptor interaction, and the PI3K-Akt signaling pathway.

[0111] Example 4: Application of plasma proteomic biomarkers in early diagnosis of ESCC

[0112] Uniform manifold approximation and projection (UMAP) analysis of 1323 proteins revealed significant differences among the three groups (HC, EIN, and ESCC), indicating distinct plasma proteomic characteristics. Figure 3 A). To identify disease-related changes in the plasma proteome, we performed a differential protein abundance analysis comparing plasma samples from patients with ESCC or EIN with those from the HC population. This analysis identified 65 significantly differentially expressed proteins (DEPs), of which 11 were upregulated and 1 was downregulated in ESCC and EIN plasma samples compared to HC samples. Figure 3 B). The expression profiles of these 12 dysregulated proteins effectively distinguished the ESCC and EIN samples from the HC samples. Figure 3 C). Functional enrichment analysis of these dysregulated proteins revealed their involvement in key biological processes, including enzyme activity regulation, lipid metabolism, reactive oxygen species responses, oxidative stress, and cell adhesion. Furthermore, these proteins are associated with molecular functions related to protein and lipid binding. Figure 3 D). We further compared the plasma proteomic profiles of ESCC patients with those of EIN and HC patients, identifying 15 ESCC-specific proteins (D). Figure 3 E). The expression profiles of these 15 ESCC-specific proteins effectively distinguished ESCCs in EIN and HC samples. Figure 3 F). Functional enrichment analysis showed that these ESCC-specific proteins are primarily involved in biological processes such as hemostasis, coagulation, and cell adhesion. Furthermore, they are also associated with receptor activation, molecular functions related to various binding interactions, and known cancer-related pathways. Figure 3 G).

[0113] To identify potential protein biomarkers for detecting EIN and ESCC, we evaluated the diagnostic performance of 12 disease-related proteins and 15 ESCC-specific proteins using a learned vector quantization (LVQ) model. By evaluating the accuracy of each protein in distinguishing between EIN, ESCC, and HC, we selected 11 proteins with an accuracy higher than 0.8 as candidate biomarkers. To validate the stability and reproducibility of these 11 candidate biomarkers, we used ELISA to detect the expression levels of these proteins in a randomly selected discovery cohort. The expression levels of four proteins detected by ELISA (FN1, VWF, FBXO34, and ITGA2B) showed good correlation with the DIA-MS proteomics results. Compared to HC, the plasma levels of FN1, VWF, FBXO34, and ITGA2B gradually increased from EIN to ESCC. Figure 3 The presence of H indicates that elevated plasma levels of these markers are associated with a higher risk of disease, thus validating their predictive value and potential as biomarkers for predicting disease risk.

[0114] Example 5: Establishment and Independent Validation of an ESCC Risk Prediction Model Based on Plasma Proteins

[0115] To evaluate the clinical potential of the identified biomarkers, we developed PROSECP, a plasma protein-based early ESCC risk prediction screening model. This model was built upon the expression levels of four protein biomarkers measured by ELISA in the discovery cohort. PROSECP, constructed using an ordinal regression approach, demonstrated high diagnostic performance in distinguishing EIN and ESCC patients from HC, with an AUC of 0.981 (95% CI 0.958–1.000), specificity of 92.9%, and sensitivity of 96.7%. Specifically, PROSECP accurately distinguished between EIN and HC, with an AUC of 0.935 (95% CI 0.855–1.000), specificity of 92.9%, and sensitivity of 93.3%. This method also differentiated ESCC from HC, with an AUC of 0.999 (95% CI 0.996–1.000), specificity of 100%, and sensitivity of 96.7%. Figure 4 AC).

[0116] To validate the anomalous elevation of these four protein biomarkers identified in the discovery cohort, we performed ELISA assays on plasma samples from two independent validation cohorts (CAMS cohort 2 and FAHZU cohort). Consistent with the proteomics results from the discovery cohort, the expression levels of all four biomarkers were significantly elevated in EIN and ESCC samples compared to HC. Furthermore, their expression levels were significantly increased in both EIN and ESCC samples, highlighting their potential and robustness as biomarkers of disease progression. Figure 4 F, G, J, and K).

[0117] We applied PROSECP to two independent validation cohorts to further evaluate its robustness and generality. In CAMS cohort 2, the AUC of PROSECP was 0.989 (95% CI 0.978–1.000). Figure 4 D), in the FAHZU cohort, the AUC was 0.979 (95% CI 0.957-1.000). Figure 4 (H), indicating that PROSECP maintained robust and high diagnostic performance in differentiating ESCC and EIN from HC. Notably, PROSECP easily distinguished ESCC or EIN patients from HC, with AUCs of 0.997 (95% CI 0.992–1.000) and 0.970 (95% CI 0.940–1.000) in CAMS cohort 2, respectively. Figure 4 The AUCs for the D and E cohorts of the FAHZU cohort were 0.985 (95% CI 0.966–1.000) and 0.904 (95% CI 0.821–0.987), respectively. Figure 4 H, I).

[0118] Furthermore, we evaluated the diagnostic efficacy of these four individual protein biomarkers, and the results are as follows: Figure 5 As shown. Figure 5 A shows the ROC curves of four protein biomarkers for ESCC in the CAMS cohort 1, CAMS cohort 2, and FAHZU cohorts. The results show that the ROC of each biomarker is greater than 0.7, indicating that each biomarker can distinguish between healthy individuals and those with ESCC. Figure 5 B Figure 5 C represents the ROC curves of four protein biomarkers against EIN and against ESCC and EIN in the CAMS cohort 1, CAMScohort 2, and FAHZU cohorts, respectively. Similarly, the ROC of each biomarker is greater than 0.7, indicating that each biomarker can distinguish between healthy individuals and EIN, and can also significantly distinguish between ESCC and EIN.

[0119] Furthermore, decision curve analysis (DCA) showed that PROSECP provided a greater net benefit across all three cohorts in the threshold probability range for distinguishing ESCC or EIN patients from HC, indicating its good clinical applicability.

[0120] The above description of the embodiments is only for understanding the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from the principles of the invention, and these improvements and modifications will also fall within the protection scope of the claims of the present invention.

Claims

1. The application of a reagent for detecting the expression level of a biomarker in the preparation of products for diagnosing esophageal cancer or esophageal intraepithelial neoplasia, characterized in that, The biomarkers are a combination of FN1, VWF, FBXO34 and ITGA2B, and the esophageal cancer is esophageal squamous cell carcinoma.

2. The application according to claim 1, characterized in that, The reagents include those used to detect the expression levels of the biomarkers in a sample using chromatographic, mass spectrometric, and / or protein immunoassay techniques.

3. A method for constructing a diagnostic prediction model for diagnosing / predicting esophageal cancer or esophageal intraepithelial neoplasia, characterized in that, The diagnostic prediction model includes the following combination of biomarkers: FN1, VWF, FBXO34, and ITGA2B. The diagnostic prediction model is trained to predict the subject based on the expression profile data of the subject's biomarker combination. The diagnostic prediction model is constructed using an ensemble learning method. The esophageal cancer is esophageal squamous cell carcinoma.

4. The construction method according to claim 3, characterized in that, The ensemble learning methods include linear regression, support vector machine, nearest neighbor / k-nearest neighbor, logistic regression, decision tree, k-means, random forest, naive Bayes, dimensionality reduction, and gradient boosting.

5. The construction method according to claim 3, characterized in that, The ensemble learning method is the logistic regression algorithm.

6. The construction method according to claim 3, characterized in that, The ensemble learning method is the ordered logistic regression algorithm.

7. The construction method according to claim 3, characterized in that, The diagnostic prediction model uses the following formula to diagnose whether a sample is esophageal cancer or esophageal intraepithelial neoplasia, or to predict whether there is a risk of esophageal cancer or esophageal intraepithelial neoplasia: Where n is the number of proteins used for diagnostic prediction. Expi For the expression level of each protein, Coefi The regression coefficient for each protein.

8. A device for diagnosing esophageal cancer or esophageal intraepithelial neoplasia, characterized in that, The device includes: (1) Data acquisition module, used to acquire expression profile data of biomarker combination in the sample of the subject to be tested, wherein the biomarker combination is FN1, VWF, FBXO34 and ITGA2B; (2) Diagnostic prediction module, used to provide the expression profile data of the biomarker combination obtained by the data acquisition module as input data to the trained diagnostic prediction model, wherein the diagnostic prediction model is trained to predict the subject based on the expression profile data of the subject's biomarker combination. (3) Prediction result acquisition module, used to acquire the output result of the diagnostic prediction model in the diagnostic prediction module, and obtain the prediction result of the subject; The diagnostic prediction model is a diagnostic prediction model constructed by the construction method according to any one of claims 3-7, and the esophageal cancer is esophageal squamous cell carcinoma.

9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a program, and when the processor executes the program, it implements the following method: acquiring biomarker combination expression profile data in a sample of a subject to be tested, wherein the biomarker combination is FN1, VWF, FBXO34 and ITGA2B; The biomarker combination expression profile data is used as input data to the trained diagnostic prediction model; Output the diagnostic prediction results for the test subject; The diagnostic prediction model is a diagnostic prediction model constructed by the construction method described in any one of claims 3-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; The computer program controls the computer-readable storage medium to implement the method of claim 9 during runtime.

Citation Information

Patent Citations

  • Application of platelets and markers thereof in treatment and diagnosis of liver cancer

    CN116445617A