Methods and applications of an auxiliary diagnostic model for esophageal cancer based on N-glycosyl fragments of serum immunoglobulin G Fc.

By establishing an esophageal cancer auxiliary diagnostic model based on the N-glycosyl fragment of immunoglobulin G Fc in serum, the problem of early esophageal cancer screening in existing technologies has been solved, achieving high-throughput, low-cost, and accurate esophageal cancer diagnosis. It is suitable for large-scale population screening and avoids the shortcomings of traditional examinations.

CN115524502BActive Publication Date: 2026-04-28SHANDONG FIRST MEDICAL UNIV & SHANDONG ACADEMY OF MEDICAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG FIRST MEDICAL UNIV & SHANDONG ACADEMY OF MEDICAL SCI
Filing Date
2022-10-21
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies lack high-throughput, low-cost early screening methods for esophageal cancer. Gastroscopy is difficult to perform, expensive, and has low acceptance among the population. Tumor markers and upper gastrointestinal imaging have low sensitivity and specificity, making them difficult to use for population screening.

Method used

An auxiliary diagnostic model for esophageal cancer based on the N-glycosyl fragment of immunoglobulin G Fc in serum was established. The content of biomarkers was detected by UPLC. Logistic regression was used to determine the model cutoff value, and Lasso regression was used to control the model complexity and simplify the sample processing steps. The percentage area of ​​the N-glycosyl peaks of 24 immunoglobulin G was detected.

Benefits of technology

This method provides a high-throughput, low-cost, highly accurate, and sensitive diagnostic method for esophageal cancer, suitable for large-scale population screening. It avoids the radiation damage of imaging examinations and the invasiveness of gastroscopy, improves patient compliance, and has a diagnostic accuracy far higher than conventional cancer markers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115524502B_ABST
    Figure CN115524502B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of clinical medicine, and particularly relates to a method for establishing an esophageal cancer auxiliary diagnosis model based on immunoglobulin G Fc segment N-glycosyl in serum and application thereof. The biomarker on which the esophageal cancer auxiliary diagnosis model provided by the present application is based is immunoglobulin G Fc segment N-glycosyl GP10, GP11, GP14, GP17 and GP23. The diagnosis model provided by the present application is constructed based on the level of five initial immunoglobulin G N-glycosyl in serum, and has the advantages of simplicity, convenient calculation and easy judgment. The model has high accuracy, high sensitivity and strong specificity for esophageal cancer diagnosis, and is much higher than the diagnosis of the conventional cancer markers CEA and CA199. The esophageal cancer diagnosis model constructed by the present application provides an effective, reliable and convenient method for the clinical diagnosis of gastrointestinal health, and has good auxiliary diagnosis value for gastrointestinal health.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of clinical medicine, specifically relating to a method for establishing an auxiliary diagnostic model for esophageal cancer based on the N-glycosyl fragment of immunoglobulin G Fc in serum and its application. Background Technology

[0002] Esophageal cancer is one of the leading causes of cancer-related mortality worldwide, causing more than 400,000 deaths annually. Due to its continuously increasing incidence, esophageal cancer remains a global public health concern, particularly in underdeveloped regions where its harm is even more severe. my country is a high-incidence country for esophageal cancer, accounting for approximately half of the world's new esophageal cancer cases each year. Among all malignant tumors reported annually in my country, esophageal cancer ranks fifth in incidence (approximately 21.17 per 100,000) and fourth in mortality (approximately 15.58 per 100,000). Esophageal cancer is a slow-growing cancer, and early identification and timely targeted treatment are crucial for its prevention and control. Systematic studies have shown that the one-year survival rate for T1 stage esophageal cancer can reach 95.6%, significantly higher than the one-year survival rate for T4 stage (52.8%), and the five-year survival rate for T1 stage esophageal cancer (75.2%) is even higher than the one-year survival rate for T4 stage (25.4%). Therefore, early screening is of great significance for the prognosis of esophageal cancer patients.

[0003] Currently, clinically recognized and widely used methods for early esophageal cancer screening include gastroscopy, tumor markers (such as CEA and CA19-9), and upper gastrointestinal contrast radiography. While gastroscopy has been increasingly used in the diagnosis of esophageal cancer, it is difficult to perform, requires highly skilled personnel and equipment, is inconvenient to implement, expensive, and has low acceptance among the general population, making it difficult to apply to widespread population screening. Although tumor markers and upper gastrointestinal contrast radiography are low-barrier, non-invasive examinations, their low sensitivity and specificity make them unsuitable for population screening. Therefore, there is an urgent need to develop a high-throughput, low-cost auxiliary diagnostic model for esophageal cancer that can be used for population screening.

[0004] Glycosylation is a ubiquitous post-translational modification of lipids and proteins. More than half of eukaryotes undergo protein glycosylation, and glycosylation structures are involved in the physiological functions of almost all membrane and secretory proteins. Due to the non-template-driven nature of glycosylation, the numerous monomers involved, and the presence of branched chains, research on protein glycosylation is highly complex and diverse. High-throughput glycosylation detection and analysis methods were not developed until 2008, lagging far behind genomics, transcriptomics, and proteomics research. Recent studies have shown that N-glycosylation of immunoglobulin G (IgG) is involved in the regulation of various diseases, including inflammatory diseases, autoimmune diseases, tumors, and metabolic diseases. IgG is the most important immunoglobulin in the blood, accounting for approximately 75%, and is mainly synthesized and secreted by plasma cells in the spleen and lymph nodes, playing a wide role in innate and adaptive immunity. IgG consists of two parts: an antigen-binding fragment (Fab fragment) and a crystalline fragment (Fc fragment). Each IgG molecule contains an average of 2.8 N-linked glycosyl groups, primarily linked to the 297th asparagine residue (Asn297) in the Fc region. The N-glycosyl groups of IgG include galactose, fucose, sialic acid, and diacetylglucamine, with a core fucosylated complex double-antenna glycan chain containing 0–2 galactose residues at its ends. Studies have shown that these glycosylation modifications play an important role in tumorigenesis. Currently, there are no records of constructing esophageal cancer auxiliary diagnostic models based on IgG Fc region N-glycosyl groups. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a method for establishing an esophageal cancer auxiliary diagnostic model based on the N-glycosyl fragment of immunoglobulin G Fc in serum.

[0006] The present invention also provides an application of the above-mentioned auxiliary diagnostic model for esophageal cancer.

[0007] The technical solution adopted by the present invention to achieve the above objectives is as follows:

[0008] This invention provides a biomarker for the diagnosis of esophageal cancer, wherein the biomarker is immunoglobulin G N-glycosyl GP1, GP2, GP3, GP4, GP5, GP6, GP7, GP10, GP8, GP9, GP10, GP11, GP12, GP13, GP14, GP15, GP16, GP17, GP18, GP19, GP20, GP21, GP22, GP23 and GP24.

[0009] Preferably, the biomarkers are immunoglobulin G N-glycosyl GP10, GP11, GP14, GP17 and GP23.

[0010] This invention also provides a method for establishing an esophageal cancer auxiliary diagnostic model based on the above-mentioned biomarker—the N-glycosylation of the Fc fragment of immunoglobulin G in serum, comprising the following steps:

[0011] (1) The content of N-glycosyl fragment of immunoglobulin G Fc in serum was detected by UPLC method; (2) An esophageal cancer diagnostic model was established based on “GP10”, “GP11”, “GP14”, “GP17”, and “GP23” immunoglobulin G N-glycosyl fragments by logistic regression modeling, and the model cutoff value was determined.

[0012] Further, in step (1), the specific process is as follows: first, the serum IgG is separated, bound, and eluted; then, the IgG glycans are digested and denatured, and deglycosylated; 2-aminobenzamide (2-AB) is used for fluorescent labeling; ultra-high performance liquid chromatography-tandem mass spectrometry (UPLC / MS) is used for quantitative and qualitative analysis of glycosyl groups; the position of the glycosyl peak and the glycosyl peak height are compared with the standard glycosyl structure in the glycosyl database (GlycoBase) to determine the measured glycosyl structure and sugar content, and the detection results are presented in the form of peak area percentage.

[0013] Furthermore, the model cutoff value is 0.594.

[0014] This invention also provides an application of the esophageal cancer auxiliary diagnostic model established using the above-described method, which can be used as a tool to assess the health status of the gastrointestinal tract, the risk of esophageal cancer, guide health interventions, or observe the effects of health interventions.

[0015] This invention establishes and improves a UPLC-based method for detecting IgG N-glycosyl groups. Secondly, esophageal cancer patients and healthy individuals were divided into a training set and a test set at a ratio of approximately 7:3. The method was then applied to the determination of immunoglobulin G N-glycosyl groups in human serum. Specifically, the percentage of the area occupied by 24 initial immunoglobulin G N-glycosyl peaks (GP1, GP2, GP3, GP4, GP5, GP6, GP7, GP10, GP8, GP9, GP10, GP11, GP12, GP13, GP14, GP15, GP16, GP17, GP18, GP19, GP20, GP21, GP22, GP23, and GP24) in the serum of the subjects was measured. Finally, Lasso regression of the training set samples was used to determine the immunoglobulin G N-glycosyl group. N-glycosyl groups “GP3”, “GP6”, “GP10”, “GP11”, “GP14”, “GP17”, “GP20”, and “GP23” can serve as potential biomarkers for esophageal cancer diagnosis. Finally, an esophageal cancer diagnostic model was established based on the immunoglobulin G N-glycosyl groups “GP10”, “GP11”, “GP14”, “GP17”, and “GP23” using logistic regression modeling. The cutoff value was determined, and the diagnostic model was validated using test set samples to confirm its effectiveness in various diagnostic evaluation indicators for esophageal cancer.

[0016] The 24 naïve immunoglobulin G N-glycosyl groups detected in human serum include “GP1-24”.

[0017] The UPLC-based method for detecting IgG N-glycosyl groups includes the following steps: first, separation, binding, and elution of serum IgG; then, enzymatic digestion and denaturation of IgG glycans, followed by deglycosylation; fluorescent labeling with 2-aminobenzamide (2-AB); and quantitative and qualitative analysis of glycosyl groups using ultra-high performance liquid chromatography-tandem mass spectrometry (UPLC / MS). The position and height of the glycosyl peaks are compared with standard glycosyl structures in the GlycoBase database to determine the measured glycosyl structure and sugar content. The detection results are presented as a percentage of peak area.

[0018] This invention simplifies the sample processing steps. The standard procedure for IgG isolation is as follows:

[0019] (1) Protein G extraction plate pretreatment: Discard the storage buffer, wash the plate with 2 ml of ultrapure water - vacuum pump filtration to remove waste liquid; wash the plate with 2 ml of 1×PBS - vacuum pump filtration to remove waste liquid; wash the plate with 1 ml of 0.1M formic acid - vacuum pump filtration to remove waste liquid; neutralize the plate with 2 ml of 10×PBS - vacuum pump filtration to remove waste liquid; equilibrate the plate with 2 ml of 1×PBS - vacuum pump filtration to remove waste liquid; equilibrate the plate with 2 ml of 1×PBS - vacuum pump filtration to remove waste liquid.

[0020] Before adding the sample, add 100 μL of standard serum to a 2 mL collection plate. Six standard sera were designed for each 96-well plate. Add 100 μL of ultrapure water to one well as a blank control. After thawing the serum sample stored at -80°C, add 100 μL of sample to a 2 mL collection plate; dilute with 1×PBS at a ratio of 1:7; use a vacuum pump (vacuum pressure ≤ 5 inHg) to collect the filtrate into a new 2 mL collection plate; after collection, vortex to mix the filtrate.

[0021] (2) IgG binding and washing: Transfer the filtered serum to a Protein G extraction plate, let it stand for 10 min and then filter it; wash the plate with 2 ml of 1×PBS-vacuum pump to remove waste liquid.

[0022] (3) IgG elution: Place the Protein G extraction plate on a new collection plate; elute IgG with 1 mL of 0.1 M formic acid - filter with a vacuum pump and collect the eluent; add 170 μL of neutralization buffer and 1 M ammonium bicarbonate to the collection plate, mix well, and add the mixed solution to a 1.5 mL tube.

[0023] (4) IgG deglycosylation process: denaturation: add 30 μL of 1.33% SDS to each sample and vortex to mix; incubation: wrap the plate with aluminum foil and put it in the oven at 65°C for 10 minutes; remove the plate from the oven and let it cool to room temperature for 15 minutes; add 10 μL of 4% Igepal to the sample and vortex to mix.

[0024] (5) Deglycosylation: Adjust the blank control-water for control, first add 20 μL 5×PBS, then adjust the pH to 8 with 0.1mol / L NaOH, and shake to mix; add 5 μL PNGase F enzyme solution to each sample, shake to mix; wrap the plate with aluminum foil, put it in a water bath at 37℃, and incubate for 18-20 hours. The incubated samples need to be dried.

[0025] (6) Standard procedure for glycosyl labeling and purification: Preparation of 2-aminobenzamide (2-AB) labeling reagent (prepare fresh for use): Dimethyl sulfoxide (DMSO) solution should be stored at room temperature. First, prepare acetic acid-DMSO mixed solution (add acetic acid to DMSO solution and mix well); dissolve 2-AB first with the mixed solution, and then dissolve sodium borocyanide; shake the labeling reagent and place it in a 65-degree oven until dissolved (≤3min). The whole process needs to be protected from light.

[0026] (7) Sample 2-AB labeling: Add 35 μL of labeling reagent to each sample and shake to mix; place the plate in an oven at 65°C for 3 hours; remove the plate from the oven and let it cool to room temperature for 30 minutes until the next step of adding the 0.2 μm GHP filter plate.

[0027] (8) GHP filter plate pretreatment: Add 200 μL of 70% ethanol to each well of the GHP plate and filter with a vacuum pump to remove waste liquid; add 200 μL of ultrapure water to each well of the GHP plate and filter with a vacuum pump to remove waste liquid; add 200 μL of ultrapure water to each well of the GHP plate and filter with a vacuum pump to remove waste liquid; add 200 μL of 96% acetonitrile to each well of the GHP plate and filter with a vacuum pump to remove waste liquid.

[0028] (9) Purification of 2-AB labeled glycosyl groups: Add 700 μL of 100% acetonitrile to each sample, mix by vortexing, centrifuge (5000 r / min, 5 min, 4℃) and transfer all supernatant to GHP plate; incubate in GHP plate for two minutes; filter by vacuum pump (vacuum pressure 2 inHg) to remove waste liquid.

[0029] (10) Elution of 2-AB-labeled N-glycans: Add 100 μL of ultrapure water to each well and shake at 100 rpm / min for 15 minutes (to prevent splashing); collect the first elution buffer into the PCR plate using vacuum filtration; transfer the liquid in the PCR plate to a 600 μL centrifuge tube. Add 100 μL of ultrapure water to each well and shake at 100 rpm / min for 10 minutes; centrifuge the GHP plate for 5 minutes at 1000 rpm and collect the second elution buffer into the PCR plate; transfer the liquid in the PCR plate to a 600 μL centrifuge tube. Add 100 μL of ultrapure water to each well and shake at 100 rpm / min for 10 minutes; centrifuge the GHP plate for 5 minutes at 1000 rpm and collect the third elution buffer into the PCR plate; transfer the liquid in the PCR plate to a 600 μL centrifuge tube; evaporate the sugar liquid using an evaporator.

[0030] (11) Sample preparation before UPLC detection: Sample dissolution solution: 100% acetonitrile and ultrapure water are mixed in a 2:1 ratio. This mixed solution (referred to as solution A) is used to dissolve the glycan samples (prepare fresh before use). Specific operation method: Add 25 μL of solution A to each glycan sample, shake to mix, and centrifuge in a low-temperature high-speed centrifuge for 5 min. The centrifugation conditions are 5000 r / min, 4℃, and 5 min. Take 20 μL of the supernatant, transfer it to a 150 μL glass tube and write the sample number on it. Place this tube into a 2 mL sample bottle and place it into the instrument's sample tray (record the sample placement number). Place fresh ultrapure water (referred to as sample B) in the sample tray as a blank control sample for the instrument (record the sample placement number) to check the instrument's operating status (whether the system has excessive peaks, and whether the pressure and baseline are normal).

[0031] (12) UPLC detection and quality control: Injection volume: 10 μL. Injection sequence: Sample B (injected twice) ---- blank control sample (H2O) in actual sample ---- actual sample ---- quality control sample (MIX) ---- actual sample ---- Sample B (injected once) --- Sample B (flushing tubing and shutdown method). Quality control points: Sample B test results show normal pressure, normal baseline, and no excessive peaks. Quality control samples are interspersed in the actual sample for detection. The CV of the 18 peaks in the quality control sample does not exceed 15%, and the CV of the 7 obvious main peaks does not exceed 5%.

[0032] The immunoglobulin G N-glycosylase established in this invention is used to construct an esophageal cancer diagnostic model. Logistic regression is used to build the model, and the Lasso regularization penalty function is used. Its L1 penalty can reduce the regression coefficient of each variable and eliminate variables with a coefficient of 0, thereby determining a set of optimal and representative variables. This can effectively control the complexity of the model and avoid the risk of overfitting.

[0033] An esophageal cancer diagnostic model was established based on five initial N-glycosyl indicators: "GP10", "GP11", "GP14", "GP17", and "GP23". The model cutoff value was 0.594, meaning that a patient was diagnosed with esophageal cancer when the test result was greater than the cutoff value.

[0034] Compared with the prior art, the present invention has the following beneficial technical effects:

[0035] (1) The diagnostic model uses a hydrophilic interaction ultra-high performance liquid chromatography-tandem mass spectrometry method to simultaneously detect the content of 24 immunoglobulin G N-glycosyl groups in human serum. The detection of immunoglobulin G N-glycosyl groups is highly specific, and the serum pretreatment process is simple and the analysis time is short, making it suitable for high-throughput analysis and testing of clinical samples.

[0036] (2) The human serum samples used in this diagnostic model are easy to obtain and are more psychologically acceptable to the target population than urine and fecal samples. This avoids radiation damage during imaging examinations and invasive damage during gastroscopy, resulting in good patient compliance. It is particularly suitable for large-scale population screening of gastrointestinal health.

[0037] (3) The diagnostic model is constructed based on the N-glycosylation levels of five initial immunoglobulins G in serum. The model is simple, easy to calculate and easy to judge. Moreover, the model has high accuracy, high sensitivity and strong specificity in the diagnosis of esophageal cancer, which is far superior to the diagnosis of conventional cancer markers CEA and CA199.

[0038] (4) The esophageal cancer diagnostic model constructed in this invention provides an effective, reliable and convenient method for the clinical diagnosis of gastrointestinal health, and has good auxiliary diagnostic value for gastrointestinal health. Attached Figure Description

[0039] Figure 1 ROC plots for training and testing sets of serum samples from esophageal cancer patients and healthy individuals based on an immunoglobulin G N-glycosyl index model.

[0040] Figure 2 ROC plots of serum samples from esophageal cancer patients and healthy individuals, based on glycosylation model, CEA, and CA199. Detailed Implementation

[0041] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0043] Example 1

[0044] 1. Experimental Instruments and Materials

[0045] 1.1 Instruments

[0046] Hydrophilic interaction ultra-high performance liquid chromatography (HILIC-UPLC) instrument, Walters. Pipettes, single-channel adjustable range 10μL, 20μL, 100μL, 200μL, 1000μL, Eppendorf (Germany). pH meter, PHS-3C, INESA. Shaker, THERMOMIXER C, Eppendorf (Germany). Vacuum pump, DOA-P504-BN, GAST. High-speed centrifuge, Centrifuge 5430, Eppendorf (Germany). Electronic balance, FA1104 (maximum load 101g, graduation 0.1mg), Sunny Optical Hengping Instrument Co., Ltd.

[0047] 1.2 Reagents and Consumables

[0048] PHGaseF enzyme powder, manufactured by Roche. Ammonium formate, manufactured by Mreda. Ammonium bicarbonate, manufactured by BBI Lifescience (Ammonium bicarbonate). Igepal-CA630, concentrated ammonia, dimethyl sulfoxide (DMSO), sodium cyanoborohydride, acetic acid solubility, anthranilamide, methanol (HPLC grade), acetonitrile (HPLC grade), manufactured by Sigma-Aldrich (USA). Ultrapure water, certified by Thermo Fisher Scientific (Barnstead). TM The water is processed by the EASYpure II ultrapure water system.

[0049] AcrPrep GHP 0.45 μm filter plate and AcrPrep GHP 0.2 μm filter plate, manufactured by Pall Corporation. 2 mL 96-well extraction plate, 0.2 mL clear 96-well flat-top (PCR plate), 1.5 mL disposable centrifuge tubes, and disposable pipette tips (10 μL, 200 μL, 1000 μL), manufactured by Axygen Biotechnology Co., Ltd. 150 μL sample vial sleeve, manufactured by Waters.

[0050] 2. Detection of Immunoglobulin G N-glycosyl group using hydrophilic interaction ultra-high performance liquid chromatography (HILIC-UPLC)

[0051] First, serum immunoglobulin G was isolated, bound, and eluted using G protein monolayers. Then, immunoglobulin G glycans were digested, denatured, and deglycosylated using PNGase. The samples were then labeled with 2-aminobenzamide (2-AB) fluorescently. The labeled samples were subjected to quantitative and qualitative analysis of glycosyl groups using hydrophilic interaction ultra-high performance liquid chromatography-tandem mass spectrometry. The position and height of the glycosyl peaks were compared with standard glycosyl structures in the GlycoBase database to determine the measured glycosyl structure and sugar content.

[0052] 2.1 Preparation of relevant solutions

[0053] 2.1.1 Preparation of PHGaseF enzyme solution: Dissolve 250 units of PHGaseF enzyme powder in 250 μL of ultrapure water and shake to mix.

[0054] 2.2 Chromatographic conditions

[0055] The N-glycosylation of 2-AB fluorescently labeled N-glycosyl groups was separated using a Waters ACQUITY ultra-high performance liquid chromatography (UHPLC) system based on the principle of hydrophilic interaction chromatography. The separation method was linear gradient separation with 75-62% (v / v) acetonitrile solution at a flow rate of 0.4 mL / min for 25 minutes. The chromatographic column was a Phenomenex Gemini NX-C18 (3 μm, 50*2 mm), and the column temperature was 60℃. Mobile phase A consisted of 100 mM ammonium formate solution (pH 4.4), mobile phase B consisted of 100% acetonitrile, mobile phase C consisted of 90% ultrapure water (10% methanol) solution, and mobile phase D consisted of 50% methanol solution. Phases B and C were equilibrated at a ratio of 50% B and 50% C at a flow rate of 0.2 mL / min for half an hour. Then, phases A and B were equilibrated again, initially at a low flow rate of 0.2 mL / min. After equilibrating 25%A and 75%B for 20 minutes, the flow rate was adjusted to 0.4 mL / min for further equilibration. The equilibration was considered complete if the pressure fluctuation was within 30. The autosampler's injection plate temperature was 4℃, the injection volume was 10 μL, and the injection needle aspiration rate was 5 μL / s.

[0056] 2.3 Mass Spectrometry Conditions

[0057] Glycosyl structures were identified using ultra-high performance liquid chromatography-tandem matrix-assisted time-of-flight mass spectrometry (UHPLC-MLMS). The mass spectrometry scan range was 300-2000 m / z, with a frequency of 1 Hz. Internal calibration was performed using glycopeptide standards, and the detection process was controlled by MicroOTOF 2.3 software. The mass spectrometry results were automatically assigned to glycosyl structures based on mass-to-charge ratio and retention time using Xtractor2D software (www.ms-utils.org / Xtractor2D), determining the N-glycosyl structures represented by each of the 24 chromatographic peaks.

[0058] The serum samples from the subjects consisted of serum samples from 100 esophageal cancer patients and 112 healthy individuals, totaling 212 samples.

[0059] 2.4 Data Processing and Statistical Methods

[0060] The detection data were automatically processed according to the traditional ensemble algorithm, and then each chromatogram needed to be manually corrected to ensure that all samples maintained the same interval. All the chromatograms obtained were divided into 24 chromatographic peaks (GP1-GP24) in the same way. The value of each chromatographic peak was the percentage of its area under the peak to the sum of the areas under the peaks of all chromatographic peaks. This value is the direct indicator of glycosyl groups. The glycosyl structures under different chromatographic peaks were obtained using mass spectrometry analysis software Xtractor2D.

[0061] Nonparametric tests were used to compare the levels of immunoglobulin G N-glycosylome between the esophageal cancer group and the healthy group; Lasso regression was performed using R language to screen characteristic variables; the screened variables were incorporated into the logistic regression equation to construct the model; the receiver operating characteristic curve (ROC curve) was used to evaluate the diagnostic ability of the model.

[0062] 3. Establishment and validation of esophageal cancer diagnostic models

[0063] 3.1 Establishment of a diagnostic model for esophageal cancer

[0064] Lasso is a regularized linear model whose L1 penalty reduces the regression coefficient of each variable, eliminating variables with a coefficient of 0, thereby determining a set of optimal and representative variables. When building predictive or prognostic models, Lasso can effectively control model complexity and avoid the risk of overfitting. The ROC curve is a curve plotted with the false positive rate (represented by 1-specificity) on the x-axis and the true positive rate (represented by sensitivity) on the y-axis. It is mainly used to evaluate the diagnostic efficacy of clinical indicators for diseases, to identify the optimal diagnostic cutoff value, and to compare the diagnostic efficacy of various different clinical diagnostic indicators.

[0065] The model was built by training and predicting using 145 samples in the training set (including 66 esophageal cancer patients and 79 healthy individuals).

[0066] 3.2 Efficacy evaluation of the immunoglobulin G N-glycosyl diagnostic model

[0067] ROC curves were plotted for the immunoglobulin G N-glycosylation diagnostic model, such as... Figure 1 As shown, the AUC of the diagnostic model on the training and test sets is 0.987 and 0.979, respectively. The cutoff value for the training set is 0.594, with a sensitivity of 97.5% and a specificity of 90.9%.

[0068] We also plotted ROC curves for the cancer markers CEA and CA199, with diagnostic AUC values ​​of 0.641 and 0.563, respectively. Figure 2 As shown, the diagnostic efficacy of glycosyl-based models far exceeds that of CEA and CA199.

[0069] 4. Summary

[0070] An esophageal cancer diagnostic model was established using Lasso and Logistic regression analysis based on five initial N-glycosyl indicators: “GP10”, “GP11”, “GP14”, “GP17”, and “GP23”. The cutoff value for the model was 0.409 (when the predicted probability diagnostic value determined by the esophageal cancer diagnostic model is greater than the cutoff value, the patient is diagnosed with esophageal cancer).

[0071] The esophageal cancer diagnostic model showed an AUC value of 0.979, a sensitivity of 90.2%, and a specificity of 97.0% for the serum samples tested. All diagnostic evaluation indicators were superior to those of CEA and CA199, demonstrating good auxiliary diagnostic value for esophageal cancer.

[0072] Finally, it should be noted that the diagnostic process using the esophageal cancer diagnostic model provided by this invention is as follows:

[0073] 1. Draw 1 mL of peripheral venous blood from the subject on an empty stomach, centrifuge to obtain serum;

[0074] 2. The percentage of the area occupied by five initial N-glycosyl peaks, namely “GP10”, “GP11”, “GP14”, “GP17”, and “GP23”, in serum was detected by hydrophilic interaction ultra-high performance liquid chromatography-tandem mass spectrometry.

[0075] 3. Calculate the predicted probability value based on the esophageal cancer diagnostic model, and compare the calculated predicted probability value with the cutoff value (0.409). If the predicted probability value is greater than the cutoff value (0.409), the patient is diagnosed with esophageal cancer.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A biomarker for the diagnosis of esophageal cancer, characterized in that, The biomarkers are N-glycosyl GP10, GP11, GP14, GP17 and GP23 of the immunoglobulin G Fc fragment.

2. A method for establishing an esophageal cancer auxiliary diagnostic model based on the N-glycosyl fragment of immunoglobulin G Fc in serum, as described in claim 1, characterized in that, Includes the following steps: (1) Detection of the content of N-glycosyl fragment of immunoglobulin G Fc in serum based on UPLC method; (2) An esophageal cancer diagnostic model was established based on the immunoglobulin GN-glycosyl of "GP10", "GP11", "GP14", "GP17" and "GP23" by using Logistic regression modeling, and the model cutoff value was determined.

3. The method for establishing according to claim 2, characterized in that, In step (1), the specific process is as follows: First, serum IgG is separated, bound, and eluted; then, IgG glycans are digested and denatured, and deglycosylated; 2-aminobenzamide 2-AB is used for fluorescent labeling; ultra-high performance liquid chromatography-tandem mass spectrometry is used for quantitative and qualitative analysis of glycosyl groups; the position of the glycosyl peak and the glycosyl peak height are compared with the standard glycosyl structure in the glycosyl database to determine the measured glycosyl structure and sugar content, and the detection results are presented in the form of peak area percentage.

4. The method for establishing according to claim 2, characterized in that, The cutoff value for the model is 0.594.