Model for prognosis of esophageal squamous cell carcinoma precancerous lesion and application thereof

By constructing a prognostic progression model of gene mutations and copy number variations in precancerous lesions of esophageal squamous cell carcinoma, the problems of insufficient early warning capability and resource waste in existing screening strategies have been solved, enabling a more targeted and timely endoscopic follow-up strategy and improving the efficiency of early diagnosis and treatment.

CN121191578BActive Publication Date: 2026-04-07BEIJING CANCER HOSPITAL PEKING UNIV CANCER HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing esophageal cancer screening strategies rely on histopathological results and lack early warning of changes at the molecular level, resulting in insufficient early warning capabilities. Furthermore, endoscopic monitoring programs may be too intensive, increasing the risk of adverse events and wasting resources.

Method used

By identifying gene mutations and copy number variations closely related to esophageal squamous cell carcinoma (ESCC) precancerous lesions, a prognostic progression model is constructed. Combined with clinical characteristics, an algorithm is used to conduct risk assessment, providing objective and quantitative auxiliary judgment criteria.

Benefits of technology

It optimized the risk stratification management of the population after screening, improved the efficiency of early diagnosis and treatment, reduced unnecessary endoscopic examinations, lowered the risk of invasive procedures, and enhanced molecular early warning capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121191578B_ABST
    Figure CN121191578B_ABST
Patent Text Reader

Abstract

This invention provides a model for the prognostic progression of esophageal squamous cell carcinoma precancerous lesions and its application. By introducing objective and quantifiable genomic indicators, it supplements and enhances the ability to identify the progression trend of precancerous lesions. By detecting molecular-level information such as somatic gene mutations and copy number variations (CNVs), this invention constructs a prognostic progression model for esophageal squamous cell carcinoma (ESCC) precancerous lesions, enabling early warning and accurate assessment of the risk of progression of esophageal precancerous lesions to cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a model for the prognosis and progression of precancerous lesions of esophageal squamous cell carcinoma and its application, belonging to the field of biomedicine. Background Technology

[0002] Esophageal cancer is a prevalent and distinctive cancer in my country, with esophageal squamous cell carcinoma (ESCC) being the main subtype, accounting for over 90% of all esophageal cancer cases. Etiological research on ESCC has yet to yield breakthroughs, making early diagnosis and treatment the key focus of current ESCC prevention and control efforts. Population screening can reduce the disease burden of ESCC, with a focus on accurately identifying high-risk individuals who may develop cancer in the future, implementing dynamic management of precancerous lesions, and conducting endoscopic monitoring and stratified risk intervention. Currently, the National Esophageal Cancer Screening Guidelines (2024 Edition) recommend endoscopic monitoring based solely on pathological diagnosis, suggesting that individuals with low-grade intraepithelial neoplasia (LGIN) undergo gastroscopy every 1-3 years, with those having lesions larger than 1 cm requiring annual monitoring. For patients diagnosed with non-dysplastic LULs (ND-LUL), the guidelines currently do not provide corresponding monitoring recommendations.

[0003] While current esophageal cancer screening strategies primarily rely on histopathological results to determine endoscopic follow-up intervals, offering strong operability and clinical guidance value, they still suffer from significant deficiencies in early warning capabilities. Existing monitoring pathways typically depend solely on baseline screening biopsy pathology to determine the presence of precancerous lesions, thus deciding whether and how often to repeat the screening. However, mid- to long-term follow-up data from esophageal cancer screening cohort studies show that relying solely on pathological results ignores early changes at the molecular level, and the lack of supplementary molecular diagnostic dimensions diminishes the ability to predict and stratify disease progression; significant deficiencies remain in risk identification for precancerous lesions and even earlier stages. Furthermore, the current guideline-recommended endoscopic monitoring protocols may be too intensive for most subjects with detected precancerous esophageal lesions, leading to excessive testing, wasted resources, and increased risks of adverse events such as bleeding, perforation, and anesthetic complications due to the invasive nature of endoscopic procedures.

[0004] Existing studies on biomarkers for ESCC precancerous lesions are mostly single-center, cross-sectional designs, lacking large-sample, multi-center prospective validation that integrates with histopathological grading, tissue type, and clinical outcomes, and a unified, standardized testing procedure has not yet been established. Therefore, a widely applicable clinical application system cannot yet be formed, limiting the widespread use of biomarkers in actual diagnosis and treatment. Summary of the Invention

[0005] To address the aforementioned problems, this invention provides a method and system for predicting the progression of precancerous lesions of esophageal squamous cell carcinoma (ESCC) based on joint analysis of gene mutations and copy number variations (CNVs) in tissue samples. This technical solution provides pathologists with objective and quantitative auxiliary diagnostic criteria by identifying key gene mutation features closely related to the occurrence of ESCC in tissues.

[0006] To achieve the above objectives, the specific technical solution provided by the present invention is as follows:

[0007] The first aspect of the present invention provides a method for constructing a prognostic progression model of esophageal squamous cell carcinoma precancerous lesions. The steps of the construction method include: obtaining biomarker mutation and / or copy number data and clinical characteristics of subjects in samples, and constructing a model based on the biomarker mutation and / or copy number data and clinical characteristics using an algorithm.

[0008] Table 1. Biomarkers for the prognostic progression of precancerous lesions in esophageal squamous cell carcinoma

[0009]

[0010] In an optional embodiment, the biomarker is a combination of FAT2, FSIP2, CDKN2A, CHEK2, CSMD3, KMT2D, NBPF10, NOTCH1, NOTCH3, and TP53 listed in Table 1.

[0011] In an optional embodiment, the biomarker is a combination of the CNV regions described in Table 1.

[0012] In an optional embodiment, the biomarker is a combination of FAT2, FSIP2, CDKN2A, CHEK2, CSMD3, KMT2D, NBPF10, NOTCH1, NOTCH3, TP53 and CNV region as described in Table 1.

[0013] In optional embodiments, the algorithm includes one or more of the following: Cox regression model, logistic regression analysis, random forest model, principal component analysis, deep neural network, generalized linear model, LASSO regression analysis, nearest neighbor analysis, and support vector machine.

[0014] Furthermore, the algorithm is a Cox hazard regression model.

[0015] In this invention, the clinical characteristics include individuals who do not have esophageal squamous cell carcinoma, individuals with abnormal iodine staining but without pathological dysplasia, individuals with low-grade intraepithelial neoplasia, individuals with high-grade intraepithelial neoplasia, and individuals with esophageal squamous cell carcinoma.

[0016] Furthermore, the clinical features also include tissue biopsy pathology and lesion size.

[0017] The second aspect of the present invention provides the application of reagents for detecting biomarker mutations and / or copy numbers in samples in the preparation of products for predicting the risk of progression of esophageal squamous cell carcinoma and / or for the prognosis of esophageal squamous cell carcinoma, wherein the biomarkers are those described in the first aspect of the present invention.

[0018] In optional embodiments, the product is selected from reagent kits, chips, systems, devices, readable media, and program products.

[0019] Furthermore, the kit also includes one or more of the following: sample pretreatment reagents, calibrators, quality control reagents, and diluents.

[0020] In an optional embodiment, the sample includes one or more of the following: biopsy tissue, peripheral blood, plasma, serum, cerebrospinal fluid, and cell culture medium.

[0021] Furthermore, the sample is biopsy tissue or peripheral blood.

[0022] In optional embodiments, the reagents include reagents for detecting mutations in biomarkers and / or reagents for detecting copy numbers of biomarkers.

[0023] In optional embodiments, the reagents for detecting biomarker mutations include reagents used in any of the following methods: Sanger sequencing, high-throughput sequencing, real-time quantitative PCR, amplification arrest mutation system PCR, gene chip technology, pyrosequencing, denaturing high-performance liquid chromatography, high-resolution melting curve analysis, and micro digital PCR.

[0024] In optional embodiments, the reagents for detecting biomarker copy numbers include reagents used in any of the following methods: PCR-based copy number variation analysis, fluorescence in situ hybridization, comparative genomic hybridization, high-throughput sequencing, multiplex ligation probe amplification, microsatellite instability analysis, and single nucleotide polymorphism arrays.

[0025] The third aspect of the present invention provides a computer-implemented system for predicting the progression risk of esophageal squamous cell carcinoma precancerous lesions. The system includes an analysis unit, which is configured to input the biomarker mutation data and / or copy number data of the samples described in the first aspect of the present invention into a prognostic progression model constructed by the construction method described in the first aspect of the present invention, and obtain classification results.

[0026] Furthermore, the system also includes an input unit and an output unit, wherein the input unit is configured to obtain biomarker mutation data and / or copy number data of the first aspect of the present invention in the sample, and the output unit is configured to output classification results.

[0027] Furthermore, the classification results include low, medium, and high risk groups.

[0028] A fourth aspect of the present invention provides a computer device for predicting the risk of progression of esophageal squamous cell carcinoma precancerous lesions, the computer device including a memory and a processor, the memory being used to store computer programs.

[0029] The processor executes a computer program, which, when executed, implements the following methods: acquiring data, for acquiring biomarker mutation data and / or copy number data of the biomarkers described in the first aspect of the present invention in the sample; processing data, for inputting the biomarker mutation data and / or copy number data of the biomarkers described in the first aspect of the present invention in the sample into the risk score prediction model constructed by the construction method described in the first aspect of the present invention to obtain classification results; and outputting data, for outputting analysis results.

[0030] The fifth aspect of the present invention provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the following methods: acquiring data for acquiring biomarker mutation data and / or copy number data of the biomarkers described in the first aspect of the present invention in a sample; processing data for inputting the biomarker mutation data and / or copy number data of the biomarkers described in the first aspect of the present invention in the sample into a risk score prediction model constructed by the construction method described in the first aspect of the present invention to obtain classification results; and outputting data for outputting analysis results.

[0031] The sixth aspect of the present invention provides a computer program product, including a computer program for early diagnosis of esophageal squamous cell carcinoma. When executed by a processor, the computer program performs the following methods: acquiring data, for acquiring biomarker mutation data and / or copy number data of the present invention described in the first aspect of the present invention in a sample; processing data, for inputting the biomarker mutation data and / or copy number data of the present invention described in the first aspect of the present invention in a sample into a risk score prediction model constructed by the construction method described in the first aspect of the present invention to obtain classification results; and outputting data, for outputting analysis results.

[0032] Advantages and benefits of the present invention: By introducing objective and quantifiable genomic indicators, the present invention supplements and enhances the ability to identify the progression trend of precancerous lesions, thereby optimizing the risk stratification management of the population after screening, providing a scientific basis for developing more targeted and timely endoscopic follow-up strategies, and is expected to improve the efficiency of early diagnosis and treatment, and fill the gap in the molecular early warning dimension of existing screening pathways. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the research roadmap for the present invention.

[0034] Figure 2 To identify the distribution of risk scores for different pathological grades, we present the following graphs: a) represents the risk score distribution for somatic mutation combinations, b) represents the risk score distribution for CNV gene combinations, and c) represents the risk score distribution for somatic mutations and CNV gene combinations.

[0035] Figure 3 To validate the Kaplan-Meier survival curves of different combinations of risk scores for all participants in the validation set, where a is the Kaplan-Meier survival curve for the somatic mutation combination, b is the Kaplan-Meier survival curve for the CNV gene combination, and c is the Kaplan-Meier survival curve for the somatic mutation and CNV gene combination.

[0036] Figure 4 The graph shows the results of risk identification ability verification for different combinations of risk scores in the subgroup population with baseline ND-LUL. In the graph, a is the Kaplan-Meier survival curve of somatic mutation combination, b is the Kaplan-Meier survival curve of CNV gene combination, and c is the Kaplan-Meier survival curve of somatic mutation and CNV gene combination.

[0037] Figure 5 The graph shows the results of risk identification capability verification for different combinations of risk scores in the subgroup with LGIN as baseline. In the graph, a is the Kaplan-Meier survival curve of the somatic mutation combination, b is the Kaplan-Meier survival curve of the CNV gene combination, and c is the Kaplan-Meier survival curve of the somatic mutation and CNV gene combination.

[0038] Figure 6 The discovery set is shown in the ROC curves for HGIN / ESCC, where a is the ROC curve for CNV gene combinations and b is the ROC curve for somatic mutations and CNV gene combinations.

[0039] Figure 7 The ROC curves for HGIN / ESCC early warning in the validation set are shown, where a is the ROC curve for CNV gene combinations, b is the ROC curve for somatic mutations and CNV gene combinations, and c is the ROC curve for existing monitoring protocols.

[0040] Figure 8 ROC curves for a precancerous lesion progression model combining somatic mutations, CNV gene combinations, tissue biopsy pathology, and lesion size.

[0041] Figure 9 Kaplan-Meier survival curves for the validation set stratified by the current screening strategy.

[0042] Figure 10Kaplan-Meier survival curves for risk stratification according to the Cox proportional hazards model of the present invention, used to validate the set. Detailed Implementation

[0043] In the context of this invention, the terms "esophageal squamous cell carcinoma" and "esophageal squamous cell carcinoma" (ESCC) are used interchangeably to refer to malignant tumors originating from esophageal squamous epithelial cells.

[0044] The terms "and / or," "or / and," and "and / or" as used herein include any one of two or more of the related listed items, as well as any and all combinations of the related listed items. These arbitrary and all combinations include any two related listed items, any more related listed items, or a combination of all related listed items. It should be noted that when at least three items are connected by at least two conjunctions selected from "and / or," "or / and," and "and / or," it should be understood that in this application, the technical solution undoubtedly includes technical solutions connected by "logical AND," and also undoubtedly includes technical solutions connected by "logical OR." For example, "A and / or B" includes three parallel solutions: A, B, and A+B.

[0045] In the context of this invention, the term "sample" as used refers to a composition obtained from or derived from a subject (e.g., an individual of interest) that contains cells and / or other molecular entities to be characterized and / or identified based on, for example, physical, biochemical, chemical, and / or physiological characteristics. For instance, a sample refers to any sample derived from a subject of interest that is expected or known to contain cells and / or molecular entities to be characterized.

[0046] In the context of this invention, a subject refers to any individual of interest, preferably a living organism suffering from or suspected of having esophageal squamous cell carcinoma, including humans, other mammals, preferably primates, and particularly preferably humans.

[0047] In some embodiments, the method for constructing the prognostic progression model is known to those skilled in the art and can be implemented and realized in different ways, linking biomarker mutations and / or copy numbers to a certain probability or risk. Preferably, the measured values ​​of the biomarker and one or more other biomarkers are mathematically combined, and the combined values ​​are associated with the inherent prognostic progression and risk stratification. The measured values ​​of biomarker mutations and / or copy numbers described in the first aspect of the invention can be combined using any suitable existing technical mathematical method, and a prognostic progression model can be constructed algorithmically.

[0048] In some embodiments, the computer device provided by the present invention may include: a display device for displaying information to a user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including voice input, speech input, or tactile input).

[0049] This invention provides a system or apparatus programmed to implement the methods described herein. The system or apparatus can control various aspects of biomarker analysis according to the invention, such as matching data according to a risk scoring system. The system or apparatus may be a user's electronic device or a computer system remotely located relative to that electronic device. The electronic device may be a mobile electronic device.

[0050] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0051] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0052] It should be understood that the systems, devices, and methods described in this invention can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] Example

[0055] I. Samples, Materials and Methods

[0056] 1. The discovery set was selected from esophageal cancer screening in Anyang area and surgical samples from esophageal cancer patients at Anyang Cancer Hospital, along with paired blood samples. Whole-exome sequencing was performed on biopsy tissues from multi-stage precancerous lesions and early cancers, as well as surgical resection specimens. The included samples were as follows: normal tissue (both iodine staining and pathology normal): 43 cases; ND-LUL (abnormal iodine staining lesions but pathological diagnosis non-dysplasia): 42 cases; low-grade intraepithelial neoplasia (LGIN): 36 cases; high-grade intraepithelial neoplasia (HGIN): 38 cases; esophageal squamous cell carcinoma (ESCC): 59 cases.

[0057] The inclusion criteria for screening were: 1) age 45-69 years; 2) no history of malignant tumors; 3) no contraindications for endoscopy, such as mental illness, cardiovascular and cerebrovascular diseases, blood-borne diseases (hepatitis B virus / hepatitis C virus / HIV positive); 4) no gastrointestinal endoscopy screening within 5 years prior to enrollment; and 5) biopsy tissue pathological diagnosis confirmed by the Department of Pathology of Anyang Cancer Hospital.

[0058] The inclusion criteria for surgical sample collection were: 1) patients treated at Anyang Cancer Hospital who had not undergone radiotherapy or chemotherapy before surgery; 2) the surgical resection tissue was confirmed by the pathology department of Anyang Cancer Hospital to be esophageal squamous cell carcinoma; 3) there were surgical resection tissue samples and corresponding blood tissue samples.

[0059] The validation set was derived from baseline samples of participants who underwent iodine staining endoscopic screening in the ESECC study (Endoscopic Screening for Esophageal Cancer in China). Inclusion criteria were: 1) Individuals with poorly stained areas on baseline endoscopic screening and available tissue samples; 2) Individuals lacking endoscopic lesion features, those with baseline endoscopic pathological diagnoses of HGIN or higher grade lesions, and those unable to provide adequate quality biopsy tissue DNA samples were excluded. A total of 1058 participants were included in the study, comprising biopsy tissue and paired blood samples. The study endpoint was the reporting of 113 cases of high-grade intraepithelial neoplasia, carcinoma in situ, and esophageal squamous cell carcinoma during the follow-up period (January 1, 2012 – June 30, 2024).

[0060] The ESECC study (Number: NCT01688908) was initiated in 2012 to evaluate the effectiveness and health economic value of endoscopic screening in the prevention and control of esophageal squamous cell carcinoma (ESCC). The ESECC study was approved by the Ethics Review Committee of Peking University Cancer Hospital, and all participants signed informed consent forms. The ESECC study selected a rural area in Huaxian County, Anyang City, Henan Province as the research site. After excluding 112 administrative villages with populations of either 3000 or 500, the ESECC study used simple random sampling to select 668 administrative villages with populations between 500 and 3000 from the remaining 846 administrative villages as target villages. Then, a cluster block randomization method was used to randomly assign these 668 administrative villages to either the screening group or the control group. Participants in each administrative village who met the following inclusion criteria were assigned to either the screening group or the control group accordingly.

[0061] Inclusion criteria for study participants: 1) Permanent residents of the target administrative village, aged 45-69 years; 2) No history of malignant tumors; 3) No contraindications for endoscopy, such as mental illness, cardiovascular and cerebrovascular diseases, blood-borne diseases (hepatitis B virus / hepatitis C virus / HIV positive); 4) No gastrointestinal endoscopy screening within 5 years prior to enrollment; 5) Voluntary participation in the ESECC study and agreement to randomization and endoscopic follow-up.

[0062] The ESECC study collected baseline data between 2012 and 2016. Participants in the screening group underwent upper gastrointestinal endoscopy with white light and iodine staining, and targeted biopsies were performed on areas with suspected abnormal iodine staining. Tissue samples were fixed in 10% neutral formalin, embedded in paraffin, stained with H&E, and then interpreted by pathologists at Anyang Cancer Hospital. Simultaneously, corresponding fresh tissue samples were collected and stored at -80°C for subsequent molecular testing.

[0063] 2. Prospective follow-up and determination of progress and outcome

[0064] The ESECC study employed a combination of active and passive follow-up to collect tumor incidence information from study participants. Active follow-up involved on-site team members conducting home visits to record the time of tumor onset, tumor type, treatment, time of death, and cause of death. Passive follow-up involved accessing the databases of the "New Rural Cooperative Medical Reimbursement System" and the "Death Registration System" in Huaxian County, Anyang City, Henan Province, to obtain relevant tumor incidence and mortality data. Furthermore, ESECC conducted its first and second rounds of endoscopic follow-ups in 2017-2018 and 2023-2024, respectively. During these follow-ups, endoscopists used white light endoscopy combined with iodine staining technology and performed repeated tissue biopsies at the same esophageal anatomical location based on the lesion sites recorded at baseline or during the initial follow-up to obtain pathological diagnostic information.

[0065] Progression outcome is defined as high-grade intraepithelial neoplasia (HGIN) or higher lesions detected during follow-up or re-examination.

[0066] 3. DNA extraction and sequencing methods

[0067] Genomic DNA was extracted from fresh frozen tissue and matched blood samples using the DNeasy Blood & Tissue Kit (QIAGEN, USA) according to the manufacturer's instructions. The extracted DNA was then sent to Beijing Maikino Gene Technology Co., Ltd. for sequencing analysis. Genomic DNA sequencing was performed using the Illumina Novaseq platform (Illumina Inc., San Diego, CA, USA), generating 150-bp paired-end reads.

[0068] 4. Bioinformatics Analysis Workflow

[0069] The discovery set samples were sequenced using whole exome sequencing (WES). The data processing workflow is as follows:

[0070] 1) The original fastq file was aligned to the hg19 reference genome using BWA; 2) Samtools was used for sorting, and Picard was used to remove duplicates and build an index; 3) GATK was used for base quality recalibration, local realignment, and variant recall; 4) MuTect2 was used to detect somatic SNVs, Platypus was used to identify indels, and peripheral blood samples were used as controls; 5) ANNOVAR was used for functional annotation, including mutation regions, types, and amino acid changes; 6) CNV analysis was performed using CNVkit, LOH analysis was performed using VarScan, and LOH results were functionally annotated using ANNOVAR.

[0071] The validation set samples were obtained using whole-genome sequencing (WGS, ×10) and deep targeted panel sequencing (68 genes). The analysis workflow was as follows: 1) The alignment and sorting process was the same as that of the test set (BWA, Samtools, Picard, GATK); 2) MuTect2 and Platypus were used to identify somatic mutations in the targeted sequencing data, with paired blood samples used as a reference; 3) ANNOVAR was used to annotate all variants; 4) Control-FreeC was used to perform copy number variation (CNV) detection and LOH analysis on the tissue whole-genome sequencing data.

[0072] 5. Marker screening

[0073] 1) Somatic Gene Mutation

[0074] For candidate mutant genes with a concentrated variation frequency ≥5%, a two-step screening strategy is adopted:

[0075] First, biomarkers showing differences between the normal and cancer groups were selected, and Fisher's exact test was performed on the Normal and ESCC groups. Then, biomarkers exhibiting trends across five pathological stages were selected, and the Cochran-Armitage Trend Test was performed on samples from each stage to assess the trend; the Spearman rank correlation coefficient was then calculated. The p-values ​​from the two steps were integrated using the sum-log method, and the free descent rate (FDR) was corrected using the Benjamini-Hochberg method. Finally, significant somatic mutation genes were identified.

[0076] 2) CNV Peak

[0077] Significant peak regions of chromosomes were extracted using GISTIC 2.0, and a two-step screening strategy was applied to these peaks, consistent with that used for somatic gene mutations. The top 20 significant peak regions were then selected, and correlation analysis was performed. Highly correlated regions (r > 0.8) were removed, and the remaining regions were included.

[0078] Significant somatic gene mutations and CNV markers were finally screened (Table 2), which were used to construct a targeting panel to identify molecular features at different stages from precancerous lesions to ESCC.

[0079] Table 2. Differential markers obtained from the discovery set screening

[0080]

[0081]

[0082]

[0083] Note: a. The Q value was calculated using the BH (Benjamini-Hochberg, a multiple test correction method).

[0084] b. TP53 is presented in the form of mutant / double mutant / total (%).

[0085] c. The reference genome is hg19.

[0086] 6. Construction of the scoring system

[0087] Based on the Somatic Gene Mutation and CNV Peak obtained from the final screening, a pathological progression scoring system was constructed. First, Spearman rank correlation analysis was performed between each biomarker and the pathological stage level (five stages from normal to ESCC), and the correlation coefficient ρ value was extracted (see Tables 3 and 4), which represents the direction and intensity of its monotonic trend in disease progression (Spearman coefficient).

[0088] A risk score is constructed using a weighted linear combination approach: each individual's state value on the biomarker (whether there is a mutation, whether there is gain / loss) is multiplied by the corresponding Spearman correlation coefficient ρ and then summed to obtain the comprehensive molecular characteristic score of the individual.

[0089]

[0090]

[0091]

[0092] Table 3. CNV Peak Weighting Coefficient ρ

[0093]

[0094] Table 4. Gene weighting coefficients ρ

[0095]

[0096] First, in the discovery set, we calculated the distribution of different risk scores (Mutation score, CNV score, Genomicscore) according to five pathological grades. Each risk score showed significant distributional differences across different pathological types, demonstrating good discriminatory power for esophageal cancer progression. Results are as follows: Figure 2 As shown, the esophageal squamous cell carcinoma (ESCC) group showed a significant high-risk distribution across all risk scores, particularly concentrated in the high-risk region (0.75–1.0).

[0097] In the validation set, we also described the distribution of different risk scores and divided them into five risk levels (each group accounting for 20%) based on quantiles. Subsequently, we calculated the cumulative incidence in each risk group and plotted Kaplan-Meier survival curves to compare the outcome differences between the highest risk level (0-20%) and the lowest risk level (80-100%). The results are shown in Table 5. Figure 3 As shown.

[0098] Table 5. Cumulative incidence rates for different risk groups as defined by risk scores

[0099]

[0100]

[0101] The hazard ratios (HRs) for the two groups corresponding to the Mutation score were 2.77 (95% CI: 1.56–4.91, P<0.0001), the HRs for the two groups corresponding to the CNV score were 7.38 (95% CI: 4.55–11.98, P<0.0001), and the HRs for the two groups corresponding to the Genomicscore were 7.43 (95% CI: 3.77–14.65, P<0.0001). The differences in the incidence of progression events between the high-risk and low-risk groups in terms of Mutation score, CNV score, and Genomic score were all statistically significant, suggesting that these scores have good discriminative ability in stratifying the risk of progression of precancerous esophageal lesions.

[0102] Furthermore, we evaluated the risk identification and enrichment capabilities of the three risk scores in subgroups at different precancerous stages at baseline; in the population with ND-LUL at baseline, all three risk scores showed good discriminative power (HR). Mutation score = 2.92 95% CI: 1.426-5.98; HR CNV score = 7.01 95% CI: 3.923-12.52; HR Genomic score = 6.55 95% CI: 3.063-13.99) ( Figure 4 In the LGIN subgroup, the CNV score (HR = 4.427, 95% CI: 1.544–12.69) and Genomic score (HR = 6.709, 95% CI: 0.91–49.54) still showed good discriminative power. Figure 5 ).

[0103] 7. Construction of a prognostic progression model for precancerous esophageal lesions

[0104] In the findings set, we used HGIN / ESCC as the study outcome and fitted logistic regression models using both CNV score and Genomic score. The results are as follows: Figure 6As shown, the CNV score model alone achieved an AUC of 0.895 (95% CI: 0.846–0.935), while the Genomic score model showed an AUC of 0.900 (95% CI: 0.852–0.937).

[0105] 8. Model performance verification

[0106] In the validation set, the CNV score and Genomic score were used to validate the model, and the results were compared with the existing surveillance program (China's Early Diagnosis and Treatment Technology Program for Esophageal Cancer, 2024). Using the number of HGIN-level lesions reported during the follow-up period and the corresponding onset time as the dependent variable, and the CNV score and Genomic score as independent variables, a Cox proportional hazards model was constructed to predict the progression of precancerous lesions in ESCC.

[0107] Figure 7 The ROC curves of the Cox regression model based on CNV score and Genomic score at different time points are shown. Under CNV score prediction alone, the AUCs for 5, 8, and 12 years are 0.797, 0.811, and 0.760, respectively. At the optimal cutoff, the sensitivity is 80.9% and the specificity is 71.3% for 5 years; 81.8% and 70.1% for 8 years; and 61.6% and 83% for 12 years. Figure 7 (a) Using the Genomic score for progress risk prediction, the AUCs for 5, 8, and 12 years were 0.803, 0.808, and 0.765, respectively. At the optimal cutoff, the sensitivity was 80.9% and the specificity was 69.9% at 5 years; 81.8% and 67.9% at 8 years; and 67.1% and 80.1% at 12 years. Figure 7 (b) The results suggest that the model has relatively stable discrimination ability at different follow-up time points, and its medium-term prediction effect is relatively better. Meanwhile, compared with the risk stratification strategy recommended in the current guidelines (based on AUC, sensitivity, and specificity at different time points), the two models mentioned above ( Figure 7 In the c-values, both showed better predictive ability for the progression of precancerous lesions (p<0.001, nonparametric U-statistic).

[0108] Building upon this, we further integrated the two warning dimensions recommended by the current monitoring protocol (tissue biopsy pathology and lesion size) with the Genomic score into a single model to validate the risk identification capability of the precancerous lesion progression model that combines macroscopic and microscopic information. The model's AUCs at 5, 8, and 12 years were 0.847, 0.849, and 0.775, respectively. Under the optimal cutoff, the sensitivity was 85.7% and specificity was 77.3% at 5 years; 78.7% and 82.4% at 8 years; and 72.5% and 77.3% at 12 years. Figure 8 ).

[0109] 8. Verification of the risk stratification effect of the scoring system

[0110] The validation set was divided into three groups based on baseline pathological characteristics: the ND-LUL group (887 cases, with 70 cases progressing during the follow-up period, a cumulative progression rate of 7.89%), the LGIN group (141 cases, with 34 cases progressing during the follow-up period, a cumulative progression rate of 24.11%), and the LGIN group with lesions larger than 1 cm in diameter (30 cases, with 9 cases progressing, a cumulative progression rate of 30.00%). Results are as follows... Figure 9 As shown, after a follow-up of approximately 12.5 years, there was a significant difference in the cumulative progression rate between ND-LUL and LUL. However, there was no significant difference in the progression rate between the LGIN group and the LGIN group with lesions larger than 1 cm under microscopy. This indicates that the current post-screening monitoring strategy is not effective in stratifying subjects at high risk of progression, and that there are still high-risk subgroups among a large number of subjects who are classified as low-risk.

[0111] To improve the accuracy of risk prediction, we used the Cox proportional hazards model, which integrates molecular risk scores, pathological grades, and LUL lesion size, to divide subjects into low, intermediate, and high risk groups. In the validation cohort, this model showed significant discriminative ability in predicting the cumulative risk of future ESCC (log-rank P < 0.0001). Its advantage over current guideline recommendations lies in its ability to effectively distinguish both the subgroup with the highest risk of progression and the truly low-risk subgroup. Figure 10 Based on our risk restratification, we have developed a new monitoring strategy; we recommend that high-risk groups undergo endoscopy every two years, medium-risk groups every three years, and low-risk groups every eight years.

[0112] 9. Endoscopic monitoring strategy

[0113] Based on the results of the natural history study of esophageal cancer, we first make the following hypothesis:

[0114] 1) It takes an average of 3 years for high-grade esophageal intraepithelial neoplasia to progress to esophageal squamous cell carcinoma. That is, ESCC can be detected by endoscopy within three years before the onset of the disease.

[0115] 2) High-grade esophageal intraepithelial neoplasia is stable within 3 years, meaning that HGIN (or carcinoma in situ) can be detected by endoscopy within ±1.5 years from the current detection time.

[0116] Based on the above assumptions, we used follow-up data from a validation cohort to simulate two endoscopic monitoring (follow-up) strategies for patients with precancerous lesions: one following current guideline recommendations, and the other guided by our strategy. Our strategy, with a median follow-up of 10.8 years (95% CI: 10.7–10.8) in 1058 patients with precancerous lesions, detected 93 cases of esophageal cancer progression (82.30% of all progression cases) through 1866 endoscopy examinations. In contrast, the existing guideline-based strategy, with the same population and follow-up period, detected only 83 cases (73.45% of all progression cases) through 2187 endoscopy examinations (Table 6). Therefore, our strategy is more effective in improving the identification rate of high-risk populations and is suitable for post-endoscopic screening and monitoring.

[0117] Table 6. Performance comparison of the two endoscopic monitoring strategies

[0118]

[0119] The above description of the embodiments is only for understanding the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from the principles of the invention, and these improvements and modifications will also fall within the protection scope of the claims of the present invention.

Claims

1. A method for constructing a prognostic progression model of esophageal squamous cell carcinoma precancerous lesions, characterized in that, The construction method includes the following steps: obtaining biomarker mutation and / or copy number data and clinical characteristics of subjects in the sample, and constructing a model based on the biomarker mutation and / or copy number data and clinical characteristics using an algorithm; Table 1. Biomarkers for the prognostic progression of precancerous lesions in esophageal squamous cell carcinoma The biomarkers are combinations of gene mutations in FAT2, FSIP2, CDKN2A, CHEK2, CSMD3, KMT2D, NBPF10, NOTCH1, NOTCH3, and TP53 listed in Table 1. Alternatively, the biomarkers may be a combination of CNVs listed in Table 1; Alternatively, the biomarkers may be a combination of FAT2, FSIP2, CDKN2A, CHEK2, CSMD3, KMT2D, NBPF10, NOTCH1, NOTCH3, TP53 and CNV as listed in Table 1. The algorithm includes one or more of the following: Cox regression model, logistic regression analysis, random forest model, principal component analysis, deep neural network, generalized linear model, LASSO regression analysis, nearest neighbor analysis, and support vector machine.

2. The construction method according to claim 1, characterized in that, The clinical characteristics include individuals who do not have esophageal squamous cell carcinoma, individuals with abnormal iodine staining but no pathological dysplasia, individuals with low-grade intraepithelial neoplasia, individuals with high-grade intraepithelial neoplasia, and individuals with esophageal squamous cell carcinoma. The clinical features also include tissue biopsy pathology and lesion size.

3. The application of reagents for detecting biomarker mutations and / or copy numbers in samples in the preparation of products for predicting the risk of progression of esophageal squamous cell carcinoma precancerous lesions and / or for the prognosis of esophageal squamous cell carcinoma, characterized in that, The biomarker is the biomarker described in claim 1.

4. The application according to claim 3, characterized in that, The product is selected from any one of the following: reagent kit, chip, system, device, readable medium, and program product.

5. The application according to claim 4, characterized in that, The kit also includes one or more of the following: sample pretreatment reagents, calibrators, quality control reagents, and diluents.

6. The application according to claim 3, characterized in that, The samples include one or more of the following: biopsy tissue, peripheral blood, plasma, serum, cerebrospinal fluid, and cell culture medium.

7. The application according to claim 3, characterized in that, The sample is either biopsy tissue or peripheral blood.

8. The application according to claim 3, characterized in that, The reagents include reagents for detecting mutations in biomarkers and / or reagents for detecting copy numbers of biomarkers.

9. The application according to claim 8, characterized in that, The reagents used to detect biomarker mutations include those used in any of the following methods: Sanger sequencing, high-throughput sequencing, real-time quantitative PCR, amplification arrest mutation system PCR, gene chip technology, pyrosequencing, denaturing high-performance liquid chromatography, high-resolution melting curve analysis, and micro digital PCR.

10. The application according to claim 8, characterized in that, The reagents for detecting the copy number of biomarkers include those used in any of the following methods: PCR-based copy number variation analysis, fluorescence in situ hybridization, comparative genomic hybridization, high-throughput sequencing, multiplex ligation probe amplification, microsatellite instability analysis, and single nucleotide polymorphism arrays.

11. A computer-based system for predicting the progression of esophageal squamous cell carcinoma precancerous lesions, characterized in that, The system includes an analysis unit, which is configured to input the biomarker mutation data and / or copy number data of the sample as described in claim 1 into the prognostic progression model constructed by the construction method of any one of claims 1-2, and obtain classification results.

12. The system according to claim 11, characterized in that, The classification results include low, medium, and high risk groups.

13. The system according to claim 11, characterized in that, The system further includes an input unit and an output unit. The input unit is configured to obtain the biomarker mutation data and / or copy number data of the biomarker as described in claim 1 in the sample, and the output unit is configured to output the classification result.

14. A computer device for predicting the risk of progression of esophageal squamous cell carcinoma precancerous lesions, characterized in that, The computer device includes a memory and a processor, the memory being used to store computer programs; The processor executes a computer program, which, when executed, implements the following method: Acquire data to obtain mutation data and / or copy number data of the biomarkers described in claim 1 in the sample; The data is processed to input the biomarker mutation data and / or copy number data of the sample as described in claim 1 into the risk score prediction model constructed by the construction method of any one of claims 1-2, and to obtain the classification result; Output data, used to output analysis results.

15. A computer-readable medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the following method: Acquire data to obtain mutation data and / or copy number data of the biomarkers described in claim 1 in the sample; The data is processed to input the biomarker mutation data and / or copy number data of the sample as described in claim 1 into the risk score prediction model constructed by the construction method of any one of claims 1-2, and to obtain the classification result; Output data, used to output analysis results.

16. A computer program product comprising a computer program for the early diagnosis of esophageal squamous cell carcinoma, characterized in that, When this computer program is executed by the processor, it implements the following method: Acquire data to obtain mutation data and / or copy number data of the biomarkers described in claim 1 in the sample; The data is processed to input the biomarker mutation data and / or copy number data of the sample as described in claim 1 into the risk score prediction model constructed by the construction method of any one of claims 1-2, and to obtain the classification result; Output data, used to output analysis results.

Citation Information

Patent Citations

  • Biomarker for esophageal cancer typing and application thereof

    CN114214409A

  • Molecular typing diagnosis marker for esophageal squamous carcinoma and application

    CN115232877A