A method, apparatus, medium for predicting the absolute risk of esophageal squamous cell carcinoma
By constructing an adaptive absolute risk prediction model for esophageal squamous cell carcinoma, and combining specific age and long-term follow-up data, the problem of insufficient regional applicability of existing models is solved, realizing individualized and dynamic risk assessment, and supporting precise screening and long-term health management.
Patent Information
- Application Number
- CN202510124471.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-01-26
AI Technical Summary
Existing absolute risk prediction models for esophageal squamous cell carcinoma have failed to effectively adapt to the differences in disease burden in different regions, resulting in poor risk stratification prediction performance in high-incidence or low-incidence areas, and lack of individualized and dynamic risk assessment methods.
An adaptive absolute risk prediction model was developed, which combines age-specific morbidity data and up to 10 years of follow-up data, and incorporates factors such as preference for hard foods, preference for hot foods and other clinical characteristics to construct an unhealthy eating habit score for accurately estimating individualized absolute risk. The model is also adapted to the disease burden in different regions by adjusting the baseline hazard rate.
It enables efficient and individualized absolute risk assessment in different regions, improves the model's generalization and predictive accuracy, can dynamically monitor future disease risk, and supports regular screening and long-term health management.
Smart Images

Figure CN120072294B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medicine, and more specifically, to a method, device, medium and program product for predicting the absolute risk of esophageal squamous cell carcinoma. Background Art
[0002] Esophageal cancer (EC) currently ranks seventh in incidence and sixth in mortality worldwide. EC is divided into two histological subtypes: squamous cell carcinoma (ESCC) and adenocarcinoma (EAC). EAC is the predominant subtype of EC in Western countries. In contrast, ESCC is the predominant subtype of EC in East to Central Asia, along the East African Rift Valley, and in South Africa, accounting for over 85% of confirmed cases. The prevalence of ESCC varies significantly across geographic regions, with rates nearly 20-fold higher in very high-risk areas compared to low-risk areas. Endoscopic screening reduces mortality from ESCC, highlighting the importance of early diagnosis and treatment as an effective approach for its prevention and control. Although a number of organized ESCC screening programs have been implemented in high-risk areas, the practical and economic efficiency of universal endoscopic screening is low, and more refined approaches based on individualized risk assessment using risk prediction models are needed.
[0003] We developed a series of models to identify common cases of ESCC, focusing on the current diagnosis. These models can aid in immediate decisions regarding endoscopic screening. However, for ESCC, as with most cancers with a long, multistage natural history, periodic screening may be more effective than a single screening test. Long-term risk assessment is crucial for a regular screening strategy, particularly for individuals at increased risk for future disease. Absolute risk models can be used for comprehensive, continuous, and dynamic risk assessment, informing individuals of their long-term or even lifelong risk and encouraging engagement in repeated self-assessment and regular screening.
[0004] To date, several absolute risk prediction models for esophageal squamous cell carcinoma have been reported, most of which were developed based on European populations. Only one model was developed based on a Chinese community cohort - "Electronic Health Record-Based Absolute Risk Prediction Model for Esophageal Cancer in the Chinese Population: Model Development and External Validation" (DOI:10.2196 / 43725). However, this model has major limitations. First, it ignores the huge differences in disease burden in different regions, which will lead to huge errors in the estimation of individualized EC risk. Second, the main predictor of this model (the factor with the highest weight / contribution) is "whether to live in a high-risk area", which inevitably leads to poor predictive performance when the model is used for risk stratification in high-incidence or low-incidence areas, greatly weakening its applicability in the real world. Summary of the Invention
[0005] In response to these challenges, the present invention provides an adaptive absolute risk prediction model that incorporates age-specific incidence rates and is based on 10 years of follow-up data from a large-scale community screening trial conducted in a high-risk area for ESCC. This model accurately estimates individualized absolute risk while also adapting to regional disease burden. We recommend it as a tool for dynamic ESCC risk assessment and precise, regular screening.
[0006] The present application (first aspect) discloses a method for predicting the absolute risk of esophageal squamous cell carcinoma, the method comprising:
[0007] Obtaining the subject's preference score for hard food, high-temperature food preference score, and other clinical characteristics, including age;
[0008] The unhealthy eating habit score was calculated based on the preference score for hard food and the preference score for high-temperature food;
[0009] The unhealthy eating habit score and the other clinical characteristics are input into the absolute risk prediction model for esophageal squamous cell carcinoma to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma. The absolute risk prediction model for esophageal squamous cell carcinoma is an absolute risk prediction model constructed based on the unhealthy eating habit score and other clinical characteristics of the training set.
[0010] Furthermore, the unhealthy diet score is obtained based on the sum of the hard food preference score and the high-temperature food preference score. If the sum of the hard food preference score and the high-temperature food preference score is 0 or 1, the unhealthy diet score is 0, otherwise it is 1.
[0011] Optionally, the score for preference for hard food is expressed as follows: if the subject's preference for hard food is "occasionally", the score is 0; if the subject's preference for hard food is "often", the score is 1; the score for preference for hot food is expressed as follows: if the subject's preference for hot food is "occasionally", the score is 0; if the subject's preference for hot food is "often", the score is 1;
[0012] Optionally, if the preference for hard food and high-temperature food is less than or equal to 1 time per week, the situation is "occasionally".
[0013] Furthermore, the other clinical characteristics also include: gender, whether the diet is regular; the unhealthy eating habit score and the other clinical characteristics are input into the absolute risk prediction model of esophageal squamous cell carcinoma to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma, wherein the absolute risk prediction model of esophageal squamous cell carcinoma is an absolute risk prediction model constructed based on the unhealthy eating habit score of the training set and the other clinical characteristics.
[0014] Furthermore, the other clinical characteristics also include: BMI, family history of esophageal cancer; the unhealthy eating habit score and the other clinical characteristics are input into the esophageal squamous cell carcinoma absolute risk prediction model to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma, wherein the esophageal squamous cell carcinoma absolute risk prediction model is an absolute risk prediction model constructed based on the unhealthy eating habit score of the training set and the other clinical characteristics;
[0015] Optionally, whether the diet is regular is classified as "yes" or "no";
[0016] Optionally, the BMI is divided into two categories based on a division threshold;
[0017] Optionally, the family history of esophageal cancer is divided into three categories based on the number of cases of esophageal squamous cell carcinoma in direct blood relatives within 3 generations, namely, the number of cases is equal to 0, the number of cases is equal to 1, and the number of cases is greater than 1.
[0018] Furthermore, the absolute risk prediction model outputs the absolute risk of the subject suffering from esophageal squamous cell carcinoma at different times, wherein the different times include one or more of the following: six months later, one year later, three years later, five years later, and time later;
[0019] Optionally, the unhealthy eating habit score, the other clinical characteristics, and the predicted years are input into an absolute risk prediction model for esophageal squamous cell carcinoma to obtain the absolute risk of the subject developing esophageal squamous cell carcinoma in the predicted years.
[0020] Furthermore, the method of constructing an absolute risk prediction model based on the unhealthy eating habit score and other clinical characteristics of the training set includes:
[0021] Step 1: Obtain the unhealthy eating habit score and other clinical characteristics, overall incidence, age-specific incidence, population-attributable risk, and age-specific all-cause mortality excluding esophageal squamous cell carcinoma for the training set;
[0022] Step 2: Obtain relative risk based on the unhealthy eating habit score, other clinical characteristics, and overall morbidity.
[0023] Step 3: Calculate the baseline hazard rate based on age-specific incidence and population attributable risk;
[0024] Step 3: Obtain the absolute risk prediction model based on the baseline hazard rate, relative risk, and competing risk of death;
[0025] Optionally, the baseline hazard rate is expressed as:
[0026] h1(t)=h0(t){1-PAR(t)%}
[0027] Where h1(t) represents the baseline hazard rate; h0(t) represents the age-specific incidence of esophageal cancer; PAR(t) represents the population attributable risk;
[0028] Optionally, the absolute risk prediction model is expressed as:
[0029]
[0030] Among them, AR i (a, a+τ) represents the absolute risk of esophageal squamous cell carcinoma in the age range [a, a+τ] in the i-th risk group; h1(t) represents the baseline hazard rate; h2(t) represents the age-specific all-cause mortality rate excluding esophageal squamous cell carcinoma; r i represents the relative risk of the i-th risk group; S i (t-1) represents the survival risk of age t-1 in the i-th risk group; S i (0)=1.
[0031] Furthermore, the step of inputting the unhealthy eating habit score and the other clinical characteristics into the esophageal squamous cell carcinoma absolute risk prediction model to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma is as follows:
[0032] Step 1: determining an applicable absolute risk prediction model for esophageal squamous cell carcinoma based on the unhealthy eating habit score and the other clinical characteristics;
[0033] Step 2: Determine the baseline risk and all-cause mortality rate corresponding to the age of the subject based on the subject's age; determine the risk group to which the subject belongs, the relative risk of the risk group, and the survival risk of the subject at age t-1 in the risk group based on the subject's unhealthy eating habit score and the other clinical characteristics;
[0034] Step 3: Calculate the absolute risk based on the baseline risk corresponding to the age of the subject, all-cause mortality rate, risk group, relative risk of the risk group, and survival risk of the risk group at age t-1, where τ is six months, one year, three years, five years, and ten years later, respectively.
[0035] The second aspect of the present application discloses a prediction system for the absolute risk of esophageal squamous cell carcinoma, comprising:
[0036] Acquisition module 201: used to obtain the subject's preference score for hard food, high-temperature food preference score and other clinical characteristics, wherein the other clinical characteristics include age;
[0037] Feature extraction module 202: used to calculate the unhealthy eating habit score based on the hard food preference score and the high-temperature food preference score;
[0038] Prediction module 203: used to input the unhealthy eating habit score and the other clinical characteristics into the esophageal squamous cell carcinoma absolute risk prediction model to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma, wherein the esophageal squamous cell carcinoma absolute risk prediction model is an absolute risk prediction model constructed based on the unhealthy eating habit score and other clinical characteristics of the training set.
[0039] The third aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, it is used to perform the steps of the above method.
[0040] In a fourth aspect, the present application discloses a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned method when the computer program is executed by a processor.
[0041] In a fifth aspect, the present application discloses a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.
[0042] This application has the following beneficial effects:
[0043] (1) This application calculated an unhealthy eating habit score based on the preference score for hard food and the preference score for hot food, and determined that the unhealthy eating habit score is an effective predictor of the absolute risk of esophageal squamous cell carcinoma;
[0044] (2) We modified h0(t) to obtain various baseline risks h1(t), and calculated the absolute risk of a specific area based on the local disease burden. This can adapt to the risk burden of different risk areas. The absolute prediction model constructed in this application has strong generalization. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0046] Figure 1 This is a schematic diagram of the method flow provided by the first aspect of the embodiment of the present invention;
[0047] Figure 2 is a schematic diagram of a program product provided by the second aspect of an embodiment of the present invention;
[0048] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the present invention;
[0049] Figure 4 is a schematic diagram of the architecture of an exemplary computing device provided by an embodiment of the present invention;
[0050] Figure 5 is a schematic diagram of a storage medium provided by an embodiment of the present invention;
[0051] Figure 6 1 is an ROC diagram of an absolute risk prediction model in different groups provided by an embodiment of the present invention;
[0052] Figure 7 This is a comparison chart of absolute risks in different regions (a) and a comparison chart of populations covered by different absolute risks (b) provided by an embodiment of the present invention;
[0053] Figure 8 is an AUC value of a different risk factor provided by an embodiment of the present invention;
[0054] Figure 9 5-year (a) and 10-year risk prediction model calibration graphs (b) of a control arm model calibration assessment provided by an embodiment of the present invention;
[0055] Figure 10 This is a schematic diagram of online prediction of an absolute risk prediction method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0057] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0059] Figure 1 1 is a flow chart of a method for predicting the absolute risk of esophageal squamous cell carcinoma provided by an embodiment of the present invention. Specifically, the method comprises the following steps:
[0060] S101: Obtaining the subject's preference score for hard food, high-temperature food preference score, and other clinical characteristics, wherein the other clinical characteristics include age;
[0061] S102: Calculate the unhealthy eating habit score based on the preference score for hard food and the preference score for high-temperature food;
[0062] S103: The unhealthy eating habit score and the other clinical characteristics are input into an absolute risk prediction model for esophageal squamous cell carcinoma to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma. The absolute risk prediction model for esophageal squamous cell carcinoma is an absolute risk prediction model constructed based on the unhealthy eating habit score and other clinical characteristics of the training set.
[0063] This application is based on the following basic research.
[0064] The basic research is summarized as follows:
[0065] 1. Materials and Methods
[0066] 1.1 Subjects
[0067] The subjects we studied were participants in the ESECC trial, which was initiated in January 2012 in rural Huaxian County, Henan Province, China. The design of the trial has been reported in detail previously (see the article "Efficacy of endoscopic screening foresophageal cancer in China (ESECC)"). Briefly, 668 villages were randomly assigned in a 1:1 ratio to either a screening group or a control group using block randomization according to their population size. The ESECC trial enrolled 17,151 participants in the screening group and 16,797 in the control group. In the screening group, all participants were invited to undergo endoscopic screening; 1544 withdrew before enrollment, and 303 failed to complete the test. An additional 5 participants in the screening group and 33 in the control group did not complete the questionnaires. Thus, the number of participants who completed all examinations and surveys in the screening and control groups was 15,299 and 16,764, respectively.
[0068] Cases of esophageal malignancy were diagnosed by endoscopy combined with Lugol's iodine staining, biopsy, and pathological examination. Lugol's iodine staining has been shown to have a sensitivity of 96% and a specificity of 63% for population-based screening. In this study, we further excluded subjects with missing values for body mass index (BMI), resulting in 15,191 subjects in the screening group and 16,692 subjects in the control group.
[0069] 1.2 Data Collection and Follow-up
[0070] Trained interviewers administered a computer-assisted, personalized questionnaire to each participant at baseline, collecting information on demographic characteristics and potential risk factors for ESCC, such as lifestyle, dietary habits, and family history of ESCC. A comprehensive physical examination, including height, weight, and blood pressure measurements, was also performed.
[0071] Outcome events in this study were defined as severe esophageal dysplasia, carcinoma in situ, and squamous cell carcinoma (collectively referred to as SDA) detected during endoscopic screening and follow-up from January 2012 to August 2022. Events occurring during follow-up were ascertained from two data sources: annual door-to-door interviews (i.e., active follow-up) and a comparison with New Rural Cooperative Medical Insurance (NCMS) reimbursement data (i.e., passive follow-up), which have nearly 100% population coverage in the region. The performance of the follow-up framework has been fully evaluated previously.
[0072] 2. Statistical Analysis
[0073] We estimated individualized absolute risks of ESCC in three steps: 1) constructing a relative risk prediction model; 2) calculating region-specific age-based baseline hazard ratios; and 3) calculating individualized absolute risks and adjusting for competing mortality risks.
[0074] Step 1: Build a relative risk prediction model
[0075] Candidate predictors were selected based on the results of relevant studies and were determined by considering the applicability of each predictor variable in different populations.
[0076] Candidate variables included demographic characteristics, dietary habits, lifestyle variables, and family history of ESCC. A two-stage selection approach was used to determine the final set of variables to be included in the relative risk prediction model.
[0077] First, univariate logistic regression was used to evaluate candidate variables. Variables with an odds ratio (OR) >1.3 and a P value <0.5 or a P value <0.05 were subsequently included in a multivariate logistic regression model. Age and sex were deterministically added to the model as a priori confounders. The final relative risk prediction model was determined using backward selection using the Akaike information criterion (AIC). The area under the receiver operating characteristic curve (AUC) calculated using the bootstrap method was used to evaluate the performance of the relative risk model. AUC values were generated for all included ESECC trial participants, the screening group, and the control group.
[0078] Step 2: Calculate age-specific baseline hazard rates
[0079] The baseline hazard 1(t) represents the potential risk of an event, without considering any risk factors, and is calculated based on the age-specific incidence rate and the population attributable risk (PAR). The variable t refers to age. The age-specific incidence rate of esophageal cancer h0(t) data are derived from the reimbursement data of the new rural cooperative medical system in Huaxian County. PAR is the proportion of ESCC incidence in the population due to exposure to a series of known risk factors presented in the selected final relative risk model. Therefore, based on the relative risk model constructed in the previous step, there are a total of l risk groups. Here, l represents the total number of risk factor combinations, which is equal to the product of the number of categories for each risk factor. For example, if there are k risk factors and each risk factor has n i categories (where i = 1, 2…, k), then:
[0080] Define r i As the relative risk of the i-th risk group compared with the baseline group (the baseline group represents the group without risk factors), where r1=1 represents the baseline group, x i is the number of cases in the ith group, and x is the total number of cases in the entire cohort.
[0081] Assume that the proportion of subjects of age t belonging to the i-th risk group is P i (t), and the incidence rate h0(t) can be expressed as:
[0082]
[0083] Then, assuming s i (t) represents the proportion of cases with age t in the i-th risk group, and the calculation formula is:
[0084] s i (t) = P i (t)h1(t)r i / h0(t)
[0085] According to this formula, we get the baseline hazard 1(t) as follows:
[0086]
[0087] Step 3: Calculate individualized absolute risk and adjust for competing risk of death
[0088] The calculation of absolute risk is based on the following information:
[0089] 1) Relative risk (r i ): calculated using the relative risk prediction model;
[0090] 2) Baseline hazard h1(t)
[0091] 3) Age-specific all-cause mortality h2(t) excluding esophageal squamous cell carcinoma.
[0092] All-cause mortality was estimated based on the Huaxian death registration data, and we assumed that h2(t) was the same for all subjects.
[0093] For subjects in the age range [a, a+τ] in the i-th risk group, the absolute risk (AR) of ESCC is defined as:
[0094]
[0095] The above equation predicts the probability that a subject will experience an event of interest within a given time interval [a, a + τ], provided that the subject does not die before age a. In our study, we consider age in discrete one-year intervals. Therefore, we replace the integral with an appropriate summation to reflect the discrete nature of the age variable.
[0096] The absolute risk (AR) of ESCC development in the subjects in the age range [a, a+τ] in the adjusted risk group i is defined as:
[0097] Cumhaz(a)i =h1(a)r i +h2(a)
[0098] S i (a) = S i (a-1)exp{-Cumhaz(a) i}
[0099]
[0100] Calibration analysis was performed by stratifying the predicted AR into five quantiles, calculating the mean predicted risk and comparing it with the observed risk in each group. Calibration curves were visually inspected using calibration plots, and statistical testing was performed using the Hosmer-Lemeshow test.
[0101] Uncertainty in our model arises primarily from β (the coefficient of the relative risk model), h1(t), and h2(t). Because h1(t) and h2(t) are derived from a large population and are assumed to be relatively stable, we focus on the uncertainty in the estimate of β. We use parametric bootstrapping to assess model uncertainty (details are described in the supplementary material).
[0102] Performance of the ESCC Absolute Risk Model in Very High, High, and Low-Risk Areas. Estimation of absolute risk involves not only relative risk but also regional incidence levels. Given the significant geographic heterogeneity in ESCC incidence, adapting to regional age-specific incidence rates is crucial for the application of our model. In our study, we modified h0(t) to derive various baseline hazards h1(t) to calculate region-specific absolute risks based on the local disease burden (local esophageal cancer incidence / prevalence).
[0103] Here, we provide examples of applying the absolute risk model in ESCC very high risk (Linzhou, Henan), high risk (Huaxian, Henan), and low risk (Beijing) areas. The age-standardized incidence rates (ASRs) in Linzhou, Huaxian, and Beijing were 72.46 / 100,000 (2008-2013)
[27] , 30.03 / 100,000 (2014-2018)
[23] , and 3.78 / 100,000 (2017)
[28] , respectively. We calculated individualized absolute risks assuming that the age distributions in the three areas were comparable and that the ratios of age-specific incidence rates were the same as the ratios of ASR, i.e.
[0104]
[0105] We then used their average 5-year cumulative risk as an indicator to calculate the proportion of the population in different regions that would require screening at a range of cumulative incidence cutoffs. All analyses were performed using R software (version 4.4.3, RRID: SCR_001905), and all significance tests were two-sided tests with a P value of 0.05.
[0106] Ethical approval and consent to participate: This study was conducted in accordance with the Declaration of Helsinki and was approved by the Institutional Review Board of Peking University Oncology School [2011KT27]. Written informed consent was obtained from all participants. Data availability. The data generated for this study are available in the article and its supplementary data files. The datasets used and / or analyzed during the current study are available from the corresponding author upon reasonable request.
[0107] 3. Results
[0108] 3.1 Characteristics of study participants
[0109] A total of 31,883 participants from the ESECC trial were included in this study, of whom 151 were diagnosed as SDA cases at baseline endoscopic screening and follow-up, and 144 were diagnosed as EC cases during follow-up until August 12, 2022. The baseline distribution of candidate predictors is shown in Table 1. Compared with the non-case group (n = 31,588), individuals in the case group (n = 295) were older, predominantly male, and more likely to have a lower BMI (≤ 22 kg / m2), a family history of esophageal cancer, an irregular diet, and a preference for high-temperature and hard foods.
[0110] Table 1. Demographic and behavioral characteristics of the 31,883 participants in the Huaxian Rural Esophageal Endoscopic Screening for Cancer (ESECC) trial in China, 2012–2022.
[0111]
[0112] aOccasionally: ≤ once a week; Frequency: > once a week;
[0113] b Number of esophageal squamous cell carcinoma cases among first-degree relatives and relatives within 3 generations;
[0114] cP values were derived from the Pearson chi-square test.
[0115] 3.2 Structure and performance of relative risk models
[0116] As shown in Table 2, the final relative risk model consisted of six risk factors:
[0117] Advanced age (OR adjusted :2.00,95%CI:1.81-2.24)
[0118] Male (OR adjusted :1.56,95%CI:1.09-2.22)
[0119] Irregular eating patterns (OR adjusted :1.40,95%CI:1.09-1.79)
[0120] Prefer hot or hard food (OR adjusted :1.48,95%CI:1.13-1.94)
[0121] BMI < 22 kg / m 2 (OR adjusted :1.48,95%CI:1.13-1.94),
[0122] Family history of esophageal cancer (one case of esophageal cancer in three generations) adjusted :1.77,95%CI:1.25-2.51;
[0123] >1 ESCC case among three generations of blood relatives (OR adjusted :3.60,95%CI:2.07-6.27).
[0124] Table 2: Crude and adjusted odds ratios for ESCC risk predictors in the Huaxian ESECC trial in China from 2012 to 2022
[0125]
[0126] Abbreviations: number, number; CI, confidence interval.
[0127] aThe unhealthy eating habits score is the sum of the two variables “hard food preference” and “high-temperature food preference”.
[0128] b Number of cases of esophageal squamous cell carcinoma among direct relatives within three generations.
[0129] In the ESECC trial, the AUC of this model was 0.753 (95% CI: 0.749-0.757) ( Figure 6 -a). Specifically, the AUC of the model in the screening group was 0.759 (95% CI: 0.732-0.786) ( Figure 6 -b), and the control group was 0.749 (95% CI: 0.705-0.794) ( Figure 6 -c). Subgroup analysis showed that the AUC within the predictor variables was relatively robust ( Figure 9 ).
[0130] 3.3 Individualized absolute risk assessment
[0131] We calculated the individualized 1- to 10-year absolute risk of esophageal cancer based on the relative risk of the predictors, age-specific incidence rates, baseline risk, and all-cause mortality excluding esophageal cancer. Among all participants in the ESECC trial, the mean absolute risk of developing ESCC at 3, 5, and 10 years was 0.28%, 0.53%, and 1.30%, respectively. A total of 48 risk factor combinations were considered within each age group. Figure 7 As shown, area under the curve (AUC) analysis showed the predictive performance of risk factors including sex, dietary pattern, eating habits, family history and BMI. The model was well calibrated for the predicted 5-year absolute risk in the control arm (p = 0.460), while the 10-year prediction tended to overestimate the actual risk in the control arm of the ESECC trial (p < 0.001).
[0132] Taking into account the large geographic variation in disease burden, we calculated an absolute risk-adjusted range for ASR of 1 to 80 per 100,000.
[0133] Based on the above estimates, we developed a user-friendly online tool (https: / / pkugenetics.shinyapps.io / escc_risk_prediction / ) that can estimate an individual's absolute risk of ESCC in different regions.
[0134] The user is first asked to manually fill in the regional ASR or select from a list of regions of residence with known ASRs in the database (including Asian and African countries). The tool will then prompt the user to enter information about the selected predictors, including age, sex, height, weight, dietary habits, and family history of EC. After submitting this data, the tool will calculate and present a personalized risk profile, showing the user's predicted probability of developing ESCC within 3, 5, and 10 years ( Figure 10 ).
[0135] 3.4 Application of the absolute risk model in different ESCC burden regions
[0136] We selected three typical regions in China with different burdens of esophageal squamous cell carcinoma, namely Linzhou (formerly Lin County, an extremely high-risk area), Hua County (a high-risk area), and Beijing (a low-risk area) to validate the performance of our model. We calculated the adjusted absolute risk of subjects from these three regions over 1 to 10 years and found that the average cumulative risk of subjects in Linzhou was significantly higher than that in Hua County and Beijing ( Figure 7 Taking the 10-year cumulative risk as an example, the average absolute risk in Linzhou City is 3.21%, 2.4, and 18.9 times that of Huaxian County (1.30%) and Beijing (0.17%), respectively.
[0137] We then calculated the population proportions at different critical values using the average 5-year cumulative risk as an indicator.
[0138] Table 3: Validity of risk stratification for the 5-year cumulative risk in Beijing, Huaxian, and Linzhou calculated based on the absolute risk prediction model
[0139]
[0140] a Population coverage is the proportion of individuals whose predicted 5-year cumulative risk exceeds the cutoff value.
[0141] It is clear that the percentage of the population covered above a given absolute risk level varies widely across regions. For example, when considering
[0142] When the 5-year absolute risk threshold is 1%, in areas with extremely high incidence such as Linzhou, more than 50% of the population exceeds
[0143] The specified risk level ( Figure 7 -b and Table 3), so screening is needed. In contrast, people living in Beijing
[0144] (low-incidence areas) do not exceed this threshold. In high-incidence areas such as Huaxian, the number of people who exceed this risk threshold
[0145] The proportion of test takers is about 17%.
[0146] discuss
[0147] The prognosis of advanced esophageal squamous cell carcinoma is poor, and early diagnosis and treatment are needed to improve the prognosis.
[0148] One method is screening, but for a country with a large population like China, popularizing endoscopic screening is not practical and economically feasible.
[0149] Therefore, precise strategies based on individual risk prediction and stratification provide a practical solution. In addition, for chronic diseases such as esophageal squamous cell carcinoma, which have a long latency period from the appearance of tumor cells to the onset of symptoms, the long-term absolute risk of developing the disease in the future may become a more robust basis for regular screening and monitoring.
[0150] indicators. There is a risk of prevalent lesions. In addition, accurate prediction of absolute risk must take into account the significant geographical heterogeneity of ESCC burden. Based on a large-scale community screening trial with a 10-year follow-up, this study constructed an absolute risk prediction model for the long-term risk of ESCC that is personalized and adapted to the regional disease burden, using regional incidence as a parameter. We first developed a relative risk prediction model that incorporated six individual-level risk factors: older age, male sex, irregular dietary pattern, preference for hot or hard food, BMI < 22 kg / m2, and family history of esophageal cancer. These factors were selected based on their contribution to improving the model's predictive performance and are generally consistent with factors identified in previous studies [24,29-31], supporting their relevance in the risk prediction model.
[0151] It is worth mentioning that smoking has been shown to be a well-established risk factor in Western countries and the United States. However, in high-risk areas, particularly in the screening setting of the Taihang Mountains, smoking has not been conclusively proven to be a risk factor for ESCC, not only in the ESECC trial but also in a community survey conducted in Lin County. First, the increased risk of ESCC appears to be more closely related to smoking duration, and current smoking status may not adequately reflect the cumulative effects of long-term smoking. Second, differences in tobacco processing between cigarettes and pipes may also affect cancer risk. Finally, the lower smoking prevalence among rural women in China (who account for nearly half of ESCC cases) further reduces the overall impact of smoking as a significant risk factor in these areas.
[0152] Our model showed good calibration, with 5-year predictions closely matching the observed risk, whereas 10-year predictions tended to overestimate the actual risk. However, we believe that this discrepancy may be partly due to the selection of participants during the recruitment phase of the ESECC trial, which enrolled approximately 20% of eligible subjects in the target villages [8] who were generally healthier, more health-conscious, and engaged in fewer risky behaviors than the general population, resulting in a lower long-term incidence of ESCC in the control group of the ESECC trial.
[0153] The risk factors evaluated in our study were based on a systematic literature review and the results of a previous large-scale population-based study conducted in a high-risk area for esophageal squamous cell carcinoma. This supports the potential robustness and cross-population applicability of the model's risk factor framework. However, additional risk factors may exist in populations with different ethnic distributions and genetic backgrounds, and if these factors can be added to specific populations, they will help improve predictive power. Furthermore, despite differences in risk factor distribution and outcome ascertainment between the screening and control groups, the model demonstrated equally satisfactory predictive accuracy in both groups, indicating that the model is a robust case for risk prediction for both screening detection and subsequent identification.
[0154] To date, five absolute risk models for ESCC have been published, three of which were based on cohorts from Northern Europe and the United Kingdom. For the remaining models developed in China, only one model was built based on a community cohort, but it has significant limitations. First, the model was not developed in a cohort specifically designed for ESCC, and therefore failed to systematically collect information on potential predictive factors associated with the development of ESCC, such as preference for hard foods, eating speed, etc. Second, the model estimated the absolute risk based on the same baseline hazard. This "one-size-fits-all" approach largely ignores the huge regional differences in the burden of esophageal cancer disease. In this study, we illustrate the significant differences in absolute risk among three regions with different incidence rates of esophageal squamous cell carcinoma (ESCC). Figure 7 -a), highlighting the inaccuracy of previously reported absolute risk prediction models, which do not account for regional heterogeneity in disease burden. Third, the model combined individuals from both high- and low-incidence areas during model development and generated a "high-risk area or not" variable, resulting in a higher area under the curve (AUC). However, the inclusion of this binary variable diluted the substantial geographic variation observed in real-world practice when the model was applied to high- or low-incidence areas.
[0155] Our model further incorporates flexible incidence parameters to reflect the disease burden in different regions. Our model further incorporates regional incidence rates, allowing participants to select their region and calculate their absolute risk based on their region. This makes it a more precise and flexible tool for assessing the long-term absolute risk of ESCC. This enhances its applicability across geographic regions worldwide, enabling the development of more targeted and effective prevention and control strategies. For example, if we define high-risk individuals as those with a 5-year risk greater than 1%, our model shows that half of the population in Linzhou is considered high-risk, compared to approximately 17% in Huaxian County and 0% in Beijing.
[0156] Our online tool helps proactively estimate an individual's absolute risk of ESCC across diverse geographic regions, not only in China but also in other Asian and African countries where ESCC is the predominant subtype of esophageal cancer. Furthermore, for public health policymakers, this model can identify high-risk individuals based on local disease burden and healthcare resources, guiding targeted screening policies in both high- and low-incidence areas.
[0157] Conclusions: In conclusion, we developed a robust model that predicts the individualized absolute risk of ESCC, performs well, and can be adjusted to reflect the global regional burden of the disease. This model can help identify individuals at high risk and empower them to proactively undergo regular endoscopic screening. Furthermore, our model can help healthcare professionals plan long-term care and surveillance for screened individuals, thereby ensuring continued lifelong health management and monitoring.
[0158] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the present invention, such as Figure 3 As shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein the memories store computer-readable codes, and when the computer-readable codes are run by the one or more processors, they may execute the method described above.
[0159] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. It can implement or execute the various methods, operations, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or any conventional processor, etc., and can be an X86 architecture or an ARM architecture.
[0160] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0161] For example, the method or apparatus according to the embodiment of the present disclosure may also be implemented by Figure 4 The architecture of the computing device 3000 shown in FIG. Figure 4As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method provided in the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 4 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 4 One or more components of a computing device are shown.
[0162] The embodiment of the present invention further provides a computer-readable storage medium, such as Figure 5 As shown, it is a schematic diagram of a storage medium 4000 provided in an embodiment of the present invention, and computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are executed by the processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0163] The present disclosure also provides a computer program product or a computer program, which implements the steps of the above method when executed by a processor, such as Figure 2 As shown, the computer program product or computer program includes:
[0164] Acquisition module 201: used to obtain the subject's preference score for hard food, high-temperature food preference score and other clinical characteristics, wherein the other clinical characteristics include age;
[0165] Feature extraction module 202: used to calculate the unhealthy eating habit score based on the hard food preference score and the high-temperature food preference score;
[0166] Prediction module 203: used to input the unhealthy eating habit score and the other clinical characteristics into the esophageal squamous cell carcinoma absolute risk prediction model to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma, wherein the esophageal squamous cell carcinoma absolute risk prediction model is an absolute risk prediction model constructed based on the unhealthy eating habit score and other clinical characteristics of the training set.
[0167] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.
[0168] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0169] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0170] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0171] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0172] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0173] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will appreciate that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for predicting the absolute risk of esophageal squamous cell carcinoma, characterized in that: The method comprises: Obtaining the subject's preference score for hard food, high-temperature food preference score, and other clinical characteristics, including age; The unhealthy eating habit score was calculated based on the preference score for hard food and the preference score for high-temperature food; The unhealthy eating habit score and the other clinical characteristics are input into an absolute risk prediction model for esophageal squamous cell carcinoma to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma. The absolute risk prediction model for esophageal squamous cell carcinoma is an absolute risk prediction model constructed based on the unhealthy eating habit score and other clinical characteristics of the training set. The method for constructing an absolute risk prediction model based on the unhealthy eating habit score and other clinical characteristics of the training set includes: Step 1: Obtain the unhealthy eating habit score and other clinical characteristics, overall incidence, age-specific incidence, population-attributable risk, and age-specific all-cause mortality excluding esophageal squamous cell carcinoma for the training set; Step 2: Obtain relative risk based on the unhealthy eating habit score, other clinical characteristics, and overall morbidity. Step 3: Calculate the baseline hazard ratio based on the age-specific incidence rate and the population attributable risk; the baseline hazard ratio is expressed as: ;in, represents the baseline hazard rate; represents the age-specific incidence of esophageal cancer; represents the population-attributable risk; Step 4: Obtain an absolute risk prediction model based on the baseline hazard rate, relative risk, and competing risk of death; the absolute risk prediction model is expressed as: ;in, represents the absolute risk of esophageal squamous cell carcinoma for subjects in the age range [𝑎,𝑎+𝜏] in the i-th risk group; the training set will not die before the age of a, and age is considered in discrete one-year intervals to adjust for the competing risk of death; represents the baseline hazard rate, represents age-specific all-cause mortality excluding esophageal squamous cell carcinoma; represents the relative risk of the i-th risk group; represents the survival risk of age t-1 in the i-th risk group; .
2. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 1, characterized in that: The unhealthy eating habit score is obtained based on the sum of the preference score for hard food and the preference score for hot food. If the sum of the preference score for hard food and the preference score for hot food is 0 or 1, the unhealthy eating habit score is 0, otherwise it is 1.
3. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 2, characterized in that: The score of preference for hard food is expressed as follows: if the subject prefers hard food "occasionally", the score is 0; if the subject prefers hard food "often", the score is 1; The high-temperature food preference score is expressed as follows: if the subject's high-temperature food preference is "occasionally", the score is 0; if it is "often", the score is 1.
4. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 3, characterized in that: If the preference for hard food and high-temperature food is less than or equal to 1 time per week, the situation is "occasionally".
5. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 1, characterized in that: The other clinical characteristics also include: gender, whether the diet is regular; the unhealthy eating habit score and the other clinical characteristics are input into the absolute risk prediction model of esophageal squamous cell carcinoma to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma, wherein the absolute risk prediction model of esophageal squamous cell carcinoma is an absolute risk prediction model constructed based on the unhealthy eating habit score of the training set and the other clinical characteristics.
6. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 1, characterized in that: The other clinical characteristics also include: BMI, family history of esophageal cancer; the unhealthy eating habit score and the other clinical characteristics are input into the absolute risk prediction model for esophageal squamous cell carcinoma to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma, wherein the absolute risk prediction model for esophageal squamous cell carcinoma is an absolute risk prediction model constructed based on the unhealthy eating habit score of the training set and the other clinical characteristics.
7. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 5, characterized in that: Whether the diet is regular is classified as "yes" or "no".
8. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 6, characterized in that: The BMI is divided into two categories based on a division threshold.
9. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 6, characterized in that: The family history of esophageal cancer is divided into three categories based on the number of cases of esophageal squamous cell carcinoma in direct blood relatives within 3 generations, namely, the number of cases is equal to 0, the number of cases is equal to 1, and the number of cases is greater than 1.
10. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 1, characterized in that: The absolute risk prediction model outputs the absolute risk of the subject suffering from esophageal squamous cell carcinoma at different times, where the different times include one or more of the following: six months later, one year later, three years later, five years later, and ten years later.
11. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 1, characterized in that: The unhealthy eating habit score, the other clinical characteristics, and the predicted years are input into the esophageal squamous cell carcinoma absolute risk prediction model to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma in the predicted years.
12. The method for predicting the absolute risk of esophageal squamous cell carcinoma according to claim 1, characterized in that: The steps of inputting the unhealthy eating habit score and the other clinical characteristics into the esophageal squamous cell carcinoma absolute risk prediction model to obtain the absolute risk of the subject suffering from esophageal squamous cell carcinoma are as follows: Step 1: determining an applicable absolute risk prediction model for esophageal squamous cell carcinoma based on the unhealthy eating habit score and the other clinical characteristics; Step 2: Determine the baseline risk and all-cause mortality rate corresponding to the age of the subject based on the subject's age; determine the risk group to which the subject belongs, the relative risk of the risk group, and the survival risk of the subject at age t-1 in the risk group based on the subject's unhealthy eating habit score and the other clinical characteristics; Step 3: Calculate the absolute risk based on the baseline risk corresponding to the age of the subject, all-cause mortality, risk group, relative risk of the risk group, and survival risk of the risk group at age t-1, where: They are half a year later, one year later, three years later, five years later, and ten years later respectively.
13. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Esophageal squamous cell carcinoma incidence risk predicting method, constructing method and application
CN111383765A