Prediction system and method for early screening of esophageal squamous carcinoma, cardia cancer and precancerous lesions
By constructing an AI prediction system based on multidimensional feature variables of exfoliated cells, the problems of efficiency and safety in the early screening of esophageal squamous cell carcinoma and gastric cardia cancer have been solved. It has achieved screening with high sensitivity and specificity, and provides risk stratification and early lesion detection for high-risk groups.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies are insufficient for efficient, safe, convenient, and economical early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions, especially in large-scale populations where screening needs are difficult to meet, and AI technology has not yet been widely applied in this field.
An AI-based prediction system was used to collect exfoliated cells from the esophagus and cardia. Multidimensional candidate feature variables were used to train risk assessment and lesion differentiation models, including cytological characteristics, CpG site methylation levels, and epidemiological risk factors. LASSO Logistic regression and elastic network penalized logistic regression models were constructed to conduct risk assessment and lesion type differentiation.
It significantly improves the sensitivity and specificity of screening, can accurately distinguish the origin of lesion tissue, provide risk stratification and minimally invasive triage for high-risk groups, reduce the endoscopic examination rate for low-risk groups, and improve the early lesion detection rate.
Smart Images

Figure CN121938601A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and in particular to a predictive system and method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions. Background Technology
[0002] Upper gastrointestinal cancers (including esophageal squamous cell carcinoma and gastric cardia cancer) are highly prevalent malignant tumors, and early screening is crucial for reducing mortality and improving patient prognosis. Currently, endoscopy combined with iodine staining biopsy is the main screening method in high-incidence areas of upper gastrointestinal cancer in my country. However, this method is limited by its high invasiveness, high technical cost, and low detection rate of target lesions, making it difficult to meet the screening needs of large-scale populations. Currently, there is a lack of a specimen collection protocol that is simple, economical, and non-invasive. To address these issues, there is an urgent need to develop an efficient and safe primary screening method to effectively triage high-risk individuals before endoscopy, allowing only high-risk individuals to undergo endoscopy. This would reduce overall screening costs and improve the utilization efficiency of endoscopic resources.
[0003] With the deep application of artificial intelligence (AI) technology in the medical field, its potential in automated pathological diagnosis is becoming increasingly prominent, especially in cervical cancer screening, where it has achieved remarkable results. Practice has shown that AI technology can effectively alleviate the workload of pathologists, reduce subjective bias in the diagnostic process, and significantly improve the homogeneity and efficiency of diagnostic results. However, currently, AI technology has not been applied to the early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions, nor has a mature AI-assisted diagnostic solution been developed.
[0004] The above description of the background technology is only for the purpose of facilitating a deeper understanding of the technical solution of the present invention (the technical means used, the technical problems solved, and the technical effects produced, etc.), and should not be regarded as an admission or in any form an implication that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a predictive system and method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer and precancerous lesions, which can efficiently and safely triage and screen high-risk groups before endoscopy.
[0006] According to an embodiment of the present invention, a predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions is provided, comprising: a collection module configured to collect multiple samples including exfoliated cells from the esophagus and gastric cardia of a subject; a preprocessing module configured to preprocess the collected multiple samples to obtain multiple training samples; a candidate feature variable screening module configured to use the multiple training samples to screen candidate feature variables for cytological feature categories, CpG site methylation level categories, and epidemiological risk factor categories of the subject; and a risk assessment model construction module configured to construct a risk assessment model, wherein the risk assessment model is LASSO. A logistic regression model; a risk assessment model training module, configured to train the risk assessment model using candidate feature variables selected from multiple training samples, to obtain the input feature variables and risk score results of the risk assessment model. The input feature variables include: the number of abnormal cells in the cytological feature category; the mean and variance of the DNA index of the top ten cells based on DNA index; the mean of the nuclear width of the top ten cells based on DNA index; the variance of the DNA index of all cells; the variance of the average gray level of the nuclear nucleus of all cells; the variance of the nuclear area of all cells; and the variance of the nuclear height of all cells; the methylation levels of cg27284428, cg11798358, cg00011482, and cg14633892 in the CpG site methylation level category; and age, sex, and missing teeth in the epidemiological risk factor category.
[0007] Preferably, the prediction system further includes a risk assessment model validation module, which is configured to use the Bootstrap resampling method to perform repeated sampling a predetermined number of times to evaluate the stability of the risk assessment model; or the risk assessment model validation module is configured to use independent validation samples to evaluate the generalization ability of the risk assessment model.
[0008] Preferably, the prediction system further includes: a lesion differentiation model construction module, configured to construct a lesion differentiation model, wherein the lesion differentiation model is an elastic network penalized logistic regression model; a lesion differentiation model training module, configured to train the lesion differentiation model based on risk score results and using candidate feature variables of CpG site methylation level to obtain lesion types; and a lesion differentiation model validation module, configured to use 10-fold cross-validation to evaluate the discrimination performance and generalization ability of the lesion differentiation model.
[0009] Preferably, the acquisition module is further configured to: simultaneously acquire exfoliated cells from the esophagus and cardia of the subject using a novel esophageal cell acquisition device.
[0010] Preferably, the preprocessing module is further configured to: quantitatively analyze the total number of esophageal and cardiac detached cells in the collected samples based on the SegNet algorithm to determine whether the total number of esophageal and cardiac detached cells reaches the predetermined minimum detection standard; specifically identify gastric glandular epithelial cells or cell clusters in the samples based on the CenterNet algorithm to determine whether the sample enrichment material reached the stomach during the sampling process; when it is determined that the total number of esophageal and cardiac detached cells does not reach the predetermined minimum detection standard, or when it is determined that the sample enrichment material did not reach the stomach during the sampling process, the subject... Repeated sampling is performed on the subjects; when the total number of exfoliated cells in the esophagus and cardia of the repeated samples still does not reach the predetermined minimum detection standard, or when the sample enrichment material does not reach the stomach during the sampling process, the samples are excluded, thereby obtaining multiple training samples. Each of the multiple training samples includes the value of a characteristic variable and the corresponding pathological diagnosis result. The pathological diagnosis result includes normal, low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma. The characteristic variables include cytological feature category, CpG site methylation level category, and the subject's epidemiological risk factor category.
[0011] Preferably, the candidate feature variable screening module is further configured as follows: taking the diagnostic results of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma as the target outcome, the Boruta algorithm is used to screen 105 cytological features of each sample in the training samples; the top ten cytological features with the median importance score are selected as candidate feature variables for the cytological feature category, wherein the candidate feature variables for the cytological feature category include: the number of abnormal cells, the mean and variance of the DNA index of the top ten cells based on the DNA index, the mean of the nuclear width of the top ten cells based on the DNA index, the mean of the DNA index of all cells, the variance of the DNA index of all cells, the variance of the average gray level of the nuclear nucleus of all cells, the variance of the nuclear area of all cells, the variance of the nuclear width of all cells, and the variance of the nuclear height of all cells.
[0012] Preferably, the candidate feature variable screening module is further configured to: divide multiple training samples into a case group and a control group according to the pathological diagnosis results, and match the ratio of the number of samples in the case group to the number of samples in the control group to be 1:3. The case group includes multiple training samples with pathological diagnoses of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma, while the control group includes multiple training samples with normal pathological diagnoses. Pyrosequencing technology is used to quantitatively detect the methylation levels of specific CpG sites in all training samples in the case group and the control group, respectively. Specific CpG sites may include cg27284428, cg11798358, etc. The methylation levels of training samples were analyzed between the case group and the control group. Based on the fact that the CpG site methylation level of the training samples in the case group was significantly higher than that in the corresponding CpG site methylation level of the training samples in the control group, candidate feature variables for CpG site methylation level categories were identified. Among them, the candidate feature variables for CpG site methylation level categories include: methylation levels of cg27284428, cg11798358, cg07880787, cg00011482, and cg14633892.
[0013] Preferably, the candidate feature variable screening module is further configured as follows: Multiple training samples are divided into a case group and a control group based on pathological diagnosis results, with the ratio of the number of samples in the case group to the number of samples in the control group being 1:3. The case group includes multiple training samples with pathological diagnoses of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma, while the control group includes multiple training samples with normal pathological diagnoses. For all training samples in both the case and control groups, using the diagnosis of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma as the target outcome, a logistic regression model is used to perform univariate analysis of epidemiological risk factors. These epidemiological risk factors include age, sex, smoking status, alcohol consumption, consumption of hot foods, pickled or dried foods, fried foods, missing teeth, history of upper gastrointestinal diseases, body mass index, and family history of cancer. The selected samples are then screened out... P Epidemiological risk factors with values less than 0.05 were considered as candidate characteristic variables for the epidemiological risk factor category. Among these, the candidate characteristic variables for the epidemiological risk factor category included age, sex, and missing teeth.
[0014] Preferably, the risk assessment model training module is further configured to: when training the risk assessment model using candidate feature variables selected from multiple training samples, determine the weight coefficient of each candidate feature variable, wherein candidate feature variables whose weight coefficients are compressed to 0 are automatically removed from the risk assessment model, thereby obtaining the input feature variables of the risk assessment model.
[0015] According to an embodiment of the present invention, a predictive method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions is provided, comprising the following steps: collecting multiple samples including exfoliated cells from the esophagus and gastric cardia of the subject; preprocessing the collected multiple samples to obtain multiple training samples; using the multiple training samples to screen candidate feature variables for cytological characteristic categories, CpG site methylation level categories, and epidemiological risk factor categories of the subject; and constructing a risk assessment model, wherein the risk assessment model is LASSO. A logistic regression model is used; candidate feature variables selected from multiple training samples are used to train a risk assessment model to obtain the input feature variables and risk score results of the risk assessment model. The input feature variables include: the number of abnormal cells in the cytological feature category; the mean and variance of the DNA index of the top ten cells based on the DNA index; the mean of the nuclear width of the top ten cells based on the DNA index; the variance of the DNA index of all cells; the variance of the average gray level of the nuclear nucleus of all cells; the variance of the nuclear area of all cells; and the variance of the nuclear height of all cells; the methylation levels of cg27284428, cg11798358, cg00011482, and cg14633892 in the CpG site methylation level category; and age, sex, and missing teeth in the epidemiological risk factor category.
[0016] Preferably, the prediction method further includes: using the Bootstrap resampling method to perform repeated sampling a predetermined number of times to evaluate the stability of the risk assessment model; or using independent validation samples to evaluate the generalization ability of the risk assessment model.
[0017] Preferably, the prediction method further includes: constructing a lesion differentiation model, wherein the lesion differentiation model is an elastic network penalized logistic regression model; training the lesion differentiation model based on the risk score results using candidate feature variables of CpG site methylation level to obtain lesion types; and using 10-fold cross-validation to evaluate the discrimination performance and generalization ability of the lesion differentiation model.
[0018] The present invention adopts the above technical solution, which has the following beneficial effects: This invention uses multiple training samples of exfoliated cells from the esophagus and cardia of subjects to screen out multi-dimensional candidate feature variables covering cytological characteristics, CpG site methylation levels, and subjects' epidemiological risk factors. These multi-dimensional candidate feature variables are then used to train a risk assessment model, significantly improving the model's sensitivity and specificity in screening for esophageal squamous cell carcinoma, cardia cancer, and precancerous lesions. Furthermore, by training a lesion differentiation model based on candidate feature variables of CpG site methylation levels, this invention can accurately distinguish the origin of lesion tissue, demonstrating stable differentiation efficacy and generalization.
[0019] In summary, the predictive system and method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to the embodiments of the present invention can provide an important basis for risk stratification and minimally invasive triage of high-risk groups for esophageal squamous cell carcinoma and gastric cardia cancer before endoscopy. It can not only significantly reduce the endoscopy rate and medical costs of low-risk groups, but also improve the early lesion detection rate of high-risk groups, and has important clinical application value for achieving early diagnosis and treatment of esophageal squamous cell carcinoma and gastric cardia cancer. Attached Figure Description
[0020] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. For clarity, the same components in different drawings are shown with the same reference numerals. It should be noted that the drawings are for illustrative purposes only and are not necessarily drawn to scale. In these drawings: Figure 1 This is a schematic diagram of the architecture of a predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to an embodiment of the present invention. Figure 2A The ROC curve of the risk assessment model with low-grade intraepithelial neoplasia and above as the target outcome is shown. Figure 2B The ROC curve of the risk assessment model with high-grade intraepithelial neoplasia and above as the target outcome is shown. Figure 3 The ROC curves for the lesion differentiation model after 10-fold cross-validation are shown. Figure 4 This is a flowchart of a predictive method for early screening of esophageal squamous cell carcinoma, cardia cancer, and precancerous lesions according to an embodiment of the present invention. Detailed Implementation
[0021] The following provides a detailed description of the embodiments of the present invention. These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes. However, the scope of protection of the present invention is not limited to the following embodiments.
[0022] Figure 1This is a schematic diagram of the architecture of a predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to an embodiment of the present invention.
[0023] A predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to an embodiment of the present invention may include a data acquisition module 10, a preprocessing module 20, a candidate feature variable screening module 30, a risk assessment model construction module 40, and a risk assessment model training module 50.
[0024] The data collection module 10 can be configured to collect multiple samples, including exfoliated cells from the esophagus and cardia of the subject. The preprocessing module 20 can be configured to preprocess the collected samples to obtain multiple training samples. The candidate feature variable screening module 30 can be configured to use the multiple training samples to screen candidate feature variables for cytological feature categories, CpG site methylation level categories, and epidemiological risk factor categories of the subject. The risk assessment model construction module 40 can be configured to construct a risk assessment model, which can be a LASSO Logistic Regression model. The risk assessment model training module 50 can be configured to use the screened candidate feature variables from the multiple training samples to train the risk assessment model, thereby obtaining the input feature variables and risk score results of the risk assessment model.
[0025] The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to an embodiment of the present invention may further include a risk assessment model validation module 80. The risk assessment model validation module 80 may be configured to evaluate the stability of the risk assessment model by performing repeated sampling a predetermined number of times using the Bootstrap resampling method; or to evaluate the generalization ability of the risk assessment model using independent validation samples.
[0026] The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to embodiments of the present invention is based on multiple training samples of exfoliated cells from the esophagus and gastric cardia of subjects. It screens out multi-dimensional candidate feature variables covering cytological characteristic categories, CpG site methylation level categories, and subject epidemiological risk factor categories. The risk assessment model is trained using the multi-dimensional candidate feature variables, which can significantly improve the model's sensitivity and specificity for screening esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions.
[0027] To further differentiate the origin of lesions in positive individuals identified by the risk assessment model, the prediction system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to an embodiment of the present invention may further include: a lesion differentiation model construction module 60, a lesion differentiation model training module 70, and a lesion differentiation model validation module 90. The lesion differentiation model construction module 60 can be configured to construct a lesion differentiation model, which can be an elastic network penalized logistic regression model. The lesion differentiation model training module 70 can be configured to train the lesion differentiation model based on risk score results, using candidate feature variables of CpG site methylation levels to obtain lesion types. The lesion differentiation model validation module 90 can be configured to use ten-fold cross-validation to evaluate the discrimination performance and generalization ability of the lesion differentiation model.
[0028] The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to embodiments of the present invention further constructs a lesion differentiation model and uses candidate feature variables of CpG site methylation level categories to train the lesion differentiation model, thereby accurately distinguishing the origin of lesion tissue and exhibiting stable differentiation efficacy and generalization.
[0029] The processing of each module of the predictive system for early screening of esophageal squamous cell carcinoma, cardia cancer and precancerous lesions according to embodiments of the present invention is described in detail below.
[0030] This invention employs a novel esophageal cell collection device to simultaneously obtain exfoliated cells from the esophagus and cardia of subjects. Subjects are drawn from natural populations in areas with a high incidence of upper gastrointestinal cancer (e.g., Linzhou in Henan, Cixian in Hebei, Feicheng in Shandong, Yangzhou in Jiangsu, and Yanting in Sichuan).
[0031] The novel esophageal cell collection device may include the Esoheal 1.0 esophageal cell collector, which consists of a hollow capsule shell, cell-enriching material compressed within the capsule, and a traction thread connected to the cell-enriching material. Subjects must fast for at least 2 hours before sampling. During the procedure, the medical personnel fold the traction thread in an S-shape and place it along with the capsule at the base of the subject's tongue, with the blue mark on the traction thread pointing outside the subject's mouth. The subject is then instructed to swallow the capsule with 55°C water. After the capsule enters the stomach with the water, the capsule shell completely dissolves, and the cell-enriching material expands into a dome-shaped trefoil. Subsequently, the medical personnel slowly pull back the traction thread, allowing the cell-enriching material to pass sequentially through the cardia, esophagus, and oral cavity, collecting esophageal and cardia mucosal cells during this process. Next, the medical personnel use sterile surgical scissors to cut the traction thread 1 cm from the cell-enriching material, placing the material in a fixative solution and storing it at room temperature.
[0032] In one implementation scheme, after sample collection, medical personnel assess the sample quality based on the degree of unfolding of the cell-enriched material. If it fully unfolds into a standard dome-shaped trefoil structure, it is considered a qualified sample; if it partially unfolds but does not reach the standard shape or fails to unfold, it is considered an unqualified sample.
[0033] In another implementation, a fully automated scanning analyzer scans the slides of the sample and uses intelligent stitching technology to generate a digital full-field cytology image. This image is then transmitted to a digital pathology diagnostic system, which assesses sample quality using an automated quality control algorithm. Specifically, the SegNet algorithm is used to quantitatively analyze the total number of esophageal and cardia cells in the sample to determine whether the total number of esophageal and cardia cells meets a predetermined minimum detection standard (e.g., 3 million). The CenterNet algorithm is used to specifically identify gastric glandular epithelial cells or cell clusters in the sample to determine whether the enriched material reached the stomach during sampling. When it is determined that the total number of esophageal and cardia cells does not meet the predetermined minimum detection standard, or when it is determined that the enriched material did not reach the stomach during sampling, the subject is resampled. When it is determined that for the resampled sample, the total number of esophageal and cardia cells still does not meet the predetermined minimum detection standard, or the enriched material still does not reach the stomach during sampling, the sample is excluded, thus obtaining multiple training samples that meet the quality control standards.
[0034] All subjects underwent endoscopic examination and pathological diagnosis on the same day after completing cytological sampling, thereby obtaining the corresponding pathological diagnosis results.
[0035] Accordingly, each of the multiple training samples may include the values of feature variables and the corresponding pathological diagnosis results. Feature variables include cytological feature categories, CpG site methylation level categories, and epidemiological risk factor categories of the subjects. Pathological diagnosis results are histopathological diagnosis results, including normal, low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma.
[0036] After obtaining multiple training samples, candidate feature variables were determined for cytological feature categories, CpG site methylation level categories, and subjects' epidemiological risk factor categories, respectively.
[0037] When determining candidate feature variables for cytological feature categories, the pathological diagnoses of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma were used as the target outcomes. The Boruta algorithm was employed to filter 105 cytological features for each sample from multiple training samples. The top ten cytological features with the highest median importance scores were selected as candidate feature variables for cytological feature categories. Table 1 presents the median importance scores of cytological features as an example.
[0038] [Table 1] Accordingly, candidate feature variables for cytological feature categories may include: the number of abnormal cells, the mean and variance of the DNA index of the top ten cells based on the DNA index, the mean of the nuclear width of the top ten cells based on the DNA index, the mean of the DNA index of all cells, the variance of the DNA index of all cells, the variance of the average gray level of the nuclear nuclei of all cells, the variance of the nuclear area of all cells, the variance of the nuclear width of all cells, and the variance of the nuclear height of all cells.
[0039] When determining candidate feature variables for CpG site methylation level categories, multiple training samples can be divided into case and control groups based on pathological diagnosis results. The case group can include multiple training samples with pathological diagnoses of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma, while the control group can include multiple training samples with normal pathological diagnoses. The ratio of the number of training samples in the case group to the number of training samples in the control group should be approximately 1:3. For example, the case group might have 117 samples, and the control group might have 408 samples.
[0040] Pyrosequencing was used to quantitatively detect the methylation levels of specific CpG sites in all training samples from both the case and control groups. These specific CpG sites included cg27284428, cg11798358, cg07880787, cg00011482, cg14633892, and cg04415798. The differences in methylation levels between the case and control groups were analyzed. Based on the significantly higher methylation levels of cg27284428, cg11798358, cg07880787, cg00011482, and cg14633892 in the case group compared to the corresponding CpG sites in the control group, candidate characteristic variables for CpG site methylation level categories were identified.
[0041] Table 2 exemplarily presents the intergroup comparison results of methylation levels at specific CpG sites in the training samples from the case group and the control group. As shown in Table 2, compared with the control group, the methylation levels at all five CpG sites (cg27284428, cg11798358, cg07880787, cg00011482, and cg14633892) were significantly increased in the case group. P < 0.05, suggesting its potential as a biomarker for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions. However, the methylation level of cg04415798 was low in both groups, with no significant difference ( P= 0.444), therefore, the methylation level at site cg04415798 is not considered a candidate feature variable for the methylation level category of CpG sites.
[0042] [Table 2] Accordingly, candidate feature variables for the CpG site methylation level category may include: methylation levels of cg27284428, cg11798358, cg07880787, cg00011482, and cg14633892.
[0043] When identifying candidate characteristic variables for epidemiological risk factor categories, all training samples from the case and control groups can be used. The diagnostic outcomes of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and cancer can be used as target outcomes. A logistic regression model can be employed to perform univariate analysis of epidemiological risk factors. Epidemiological risk factors may include age, sex, smoking status, alcohol consumption, consumption of hot foods, pickled or dried foods, fried foods, missing teeth, history of upper gastrointestinal diseases, body mass index, and family history of cancer, etc. Screening is then performed. P Epidemiological risk factors with values less than 0.05 were used as candidate characteristic variables for epidemiological risk factor categories.
[0044] Table 3 presents epidemiological risk factor information for the training samples as an example. As shown in Table 3, the case group had a higher mean age compared to the normal control group ( P < 0.001), and the proportion of males was significantly higher than that of females ( P = 0.045). Furthermore, the proportion of missing teeth in the case group was significantly higher than that in the control group ( P < 0.05). However, no significant differences were observed between the two groups in terms of smoking and drinking.
[0045] [Table 3] Accordingly, candidate characteristic variables for epidemiological risk factor categories may include age, sex, and missing teeth.
[0046] This invention screens candidate feature variables in three categories: cellular characteristics, CpG site methylation level, and epidemiological risk factors. It constructs a multi-dimensional joint feature system from three dimensions: cell morphology, molecular biology, and population exposure risk. Thus, it uses candidate feature variables from the categories of cellular characteristics, CpG site methylation level, and the subject's epidemiological risk factors to construct a multimodal risk assessment model.
[0047] The risk assessment model according to the embodiments of the present invention can be a LASSO Logistic regression model, which can be expressed as: Logit is a risk score for esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions. These are the values of the characteristic variables. is the weight coefficient of the feature variable, and k is the number of feature variables.
[0048] According to an embodiment of the present invention, a risk assessment model is trained using candidate feature variables selected from multiple training samples to determine the weight coefficient of each candidate feature variable. The LASSO Logistic Regression model is a machine learning method that uses L1 regularization to select features. It compresses regression coefficients to zero out the coefficients of unimportant variables, while preventing overfitting. Table 4 presents the weight coefficients of each candidate feature variable as an example. .
[0049] [Table 4] As can be seen from Table 4, the weighting coefficients of the mean DNA index of all cells, the variance of the cell nuclear width of all cells, and the methylation level of cg07880787 were compressed to 0. Candidate feature variables whose weighting coefficients were compressed to 0 were automatically removed from the risk assessment model.
[0050] Accordingly, the input feature variables of the risk assessment model—the LASSO Logistic Regression Model—may include: the number of abnormal cells in the cytological feature category; the mean and variance of the DNA index of the top ten cells based on the DNA index; the mean of the nuclear width of the top ten cells based on the DNA index; the variance of the DNA index of all cells; the variance of the average gray level of the nuclear nucleus of all cells; the variance of the nuclear area of all cells; the variance of the nuclear height of all cells; the methylation levels of CpG sites cg27284428, cg11798358, cg00011482, and cg14633892 in the category; and age, sex, and missing teeth in the category of epidemiological risk factors.
[0051] The risk assessment model can output risk scores, which can be either positive (i.e., the risk score is higher than a preset threshold) or negative (i.e., the risk score is lower than a preset threshold).
[0052] After obtaining the trained risk assessment model, the risk assessment model validation module can evaluate the model's stability and generalization ability. Specifically, the stability of the risk assessment model can be evaluated by repeatedly sampling a predetermined number of times (e.g., 1000 times) using the Bootstrap resampling method.
[0053] Furthermore, independent validation samples can be used to evaluate the generalization ability of the risk assessment model. According to an embodiment of the present invention, samples from subjects in Cixian County, Hebei Province, Feicheng City, Shandong Province, Yanting County, Sichuan Province, and Yangzhou City, Jiangsu Province can be preprocessed and used as training samples, while samples from subjects in Linzhou City, Henan Province can be preprocessed and used as validation samples.
[0054] The model evaluation results show that the risk assessment model according to the embodiment of the present invention exhibits excellent predictive efficacy for low-grade intraepithelial neoplasia and higher lesions. Figure 2A The ROC (Receiver Operating Characteristic) curve of the risk assessment model according to an embodiment of the present invention, with low-grade intraepithelial neoplasia and above as the target outcome, is shown. Figure 2A As shown, the results indicated an AUC (area under the curve) of 0.922 (95% CI: 0.894–0.947). In terms of diagnostic accuracy, the sensitivity was 0.855 (95% CI: 0.780–0.907), and the specificity was 0.873 (95% CI: 0.837–0.901).
[0055] Furthermore, the risk assessment model according to the embodiments of the present invention also demonstrates excellent predictive efficacy for HGD and above lesions. Figure 2B The ROC curve of the risk assessment model according to an embodiment of the present invention, with high-grade intraepithelial neoplasia and above as the target outcome, is shown. Figure 2B As shown, the AUC value was 0.876 (95% CI: 0.820–0.933). In terms of diagnostic accuracy, the sensitivity was 0.929 (95% CI: 0.774–0.980), and the specificity was 0.746 (95% CI: 0.706–0.783).
[0056] This invention relies on a screening population in high-incidence areas of upper gastrointestinal cancer and employs a case-control study design. The evaluation results show that the LASSO Logistic regression model, constructed by combining candidate characteristic variables from three categories—cytological characteristics, CpG site methylation levels, and epidemiological risk factors—has significant predictive power.
[0057] While the risk assessment model according to the embodiments of the present invention can effectively identify positive individuals with lesion risk, it cannot further distinguish whether the lesion originates from the esophagus or the cardia. To further differentiate the origin of lesion tissue in positive individuals identified by the risk assessment model, the present invention constructs a lesion differentiation model, which can be a resilient network penalized logistic regression model.
[0058] The Elastic Network Penalized Logistic Regression (ELR) model, an improved regression model that combines the advantages of L1 regularization (LASSO) and L2 regularization (Ridge), can accurately screen key feature variables closely related to lesion type in high-dimensional feature space and effectively eliminate redundant information by organically combining loss function and double penalty term. It can also specifically eliminate multicollinearity problem between features and avoid mutual interference between dependent variables that leads to model prediction bias. At the same time, the introduction of penalty term effectively suppresses the risk of model overfitting and significantly improves the model's generalization ability on unknown samples, ensuring that it has stable and reliable discrimination power in actual clinical applications.
[0059] This invention utilizes candidate feature variables of CpG site methylation levels from multiple training samples identified as positive by a risk assessment model to train an elastic network penalized logistic regression model, thereby further obtaining the lesion type (esophageal lesion or cardia lesion).
[0060] Specifically, the risk assessment model identified 88 cases of esophageal lesions and 12 cases of cardia lesions among the positive individuals. In the selection of feature variables, this invention uses candidate feature variables based on the methylation levels of five CpG sites (i.e., cg27284428, cg11798358, cg07880787, cg00011482, and cg14633892) to train an elastic network penalized logistic regression model. By setting reasonable model hyperparameters, the elastic network penalized logistic regression model can fully learn the mapping relationship between the methylation levels of the five CpG sites and the lesion type.
[0061] After training, the Elastic Network Penalized Logistic Regression Model can obtain the input feature variables, including the methylation level of cg00011482, and can also obtain the lesion type (esophageal lesion or cardia lesion).
[0062] To verify the robustness and reliability of the trained Elastic Network Penalized Logistic Regression Model, this invention employs the 10-fold cross-validation method to evaluate the performance of the Elastic Network Penalized Logistic Regression Model.
[0063] Model evaluation results show that the elastic network penalized logistic regression model exhibits stable and excellent discriminative power, with an average area under the receiver operating characteristic (AUC) of 0.745 after 10-fold cross-validation.
[0064] Furthermore, to more intuitively verify the model performance, this invention plots a combined receiver operating characteristic (ROC) curve based on the predicted probabilities of each fold validation set. Figure 3 The ROC curve for ten-fold cross-validation of the lesion differentiation model is shown. Figure 3 As shown, the overall ROC curve exhibits a good upward trend, with a corresponding cross-validation AUC value of 0.739 (95% CI: 0.567-0.911). This fully demonstrates that the lesion differentiation model constructed in this invention has good robustness and can stably and accurately differentiate between esophageal and cardia lesions in different batches of positive samples, providing reliable technical support for subsequent precise clinical diagnosis and treatment.
[0065] The present invention also provides a predictive method for early screening of esophageal squamous cell carcinoma, cardia cancer and precancerous lesions. Figure 4 This is a flowchart illustrating a predictive method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to an embodiment of the present invention. See also... Figure 4 The prediction method may include the following steps: S100: Collect multiple samples, including exfoliated cells from the subject's esophagus and cardia; S200: Preprocess the collected samples to obtain multiple training samples; S300: Candidate feature variables are selected from multiple training samples, including cytological feature categories, CpG site methylation level categories, and subjects' epidemiological risk factor categories; S400: Construct a risk assessment model, which is a LASSO Logistic regression model; S500: Use candidate feature variables selected from multiple training samples to train the risk assessment model in order to obtain the input feature variables and risk score results of the risk assessment model; When the risk score is positive, the prediction method further includes: S600: Construct a lesion differentiation model, which is an elastic network penalized logistic regression model; S700: Use candidate feature variables of CpG site methylation level to train a lesion differentiation model to obtain lesion type (esophageal lesion or cardia lesion).
[0066] The prediction method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to embodiments of the present invention may further include the following steps: using the Bootstrap resampling method to perform repeated sampling a predetermined number of times to evaluate the stability of the risk assessment model; or using independent validation samples to evaluate the generalization ability of the risk assessment model; and using ten-fold cross-validation to evaluate the discrimination performance and generalization ability of the lesion differentiation model.
[0067] The specific processing procedures for each step are similar to those for each module of the predictive system used for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions, and will not be repeated here.
[0068] In the predictive system and method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to the present invention, integrating DNA methylation markers, epidemiological risk factors, and cytological characteristics into the risk assessment model can significantly improve the predictive efficacy of the model. Further analysis shows that DNA methylation markers have the potential to differentiate the tissue origin of esophageal and gastric cardia lesions, thus a resilient network penalized logistic regression model based on CpG sites is constructed to differentiate the tissue origin of lesions. Therefore, this can provide important scientific basis for risk stratification and minimally invasive triage of high-risk populations for esophageal squamous cell carcinoma and gastric cardia cancer before endoscopy.
[0069] The various embodiments of the present invention are not an exhaustive list of all possible combinations, but are intended to describe representative aspects of the invention, and the contents described in the various embodiments can be applied independently or in two or more combinations.
[0070] The description of the exemplary embodiments presented above is merely illustrative of the technical solutions of the present invention and is not intended to be exhaustive, nor is it intended to limit the invention to the precise forms described. Obviously, those skilled in the art can make many changes and variations based on the above teachings. The exemplary embodiments were chosen and described to explain the specific principles of the invention and its practical applications, thereby enabling others skilled in the art to understand, implement, and utilize the various exemplary embodiments of the invention and their various alternatives and modifications. The scope of protection of the present invention is intended to be defined by the appended claims and their equivalents.
Claims
1. A predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions, comprising: The collection module is configured to collect multiple samples, including exfoliated cells from the subject's esophagus and cardia; The preprocessing module is configured to preprocess multiple collected samples to obtain multiple training samples. The candidate feature variable screening module is configured to use multiple training samples to screen candidate feature variables for cytological feature categories, CpG site methylation level categories, and subject epidemiological risk factor categories; The risk assessment model building module is configured to build a risk assessment model, which is a LASSO Logistic regression model. The risk assessment model training module is configured to train the risk assessment model using candidate feature variables selected from multiple training samples, in order to obtain the input feature variables and risk score results of the risk assessment model. The input feature variables include: the number of abnormal cells in the cytological feature category; the mean and variance of the DNA index of the top ten cells based on the DNA index; the mean of the nuclear width of the top ten cells based on the DNA index; the variance of the DNA index of all cells; the variance of the average gray level of the nuclear nucleus of all cells; the variance of the nuclear area of all cells; and the variance of the nuclear height of all cells; the methylation levels of cg27284428, cg11798358, cg00011482, and cg14633892 in the CpG site methylation level category; and age, sex, and missing teeth in the epidemiological risk factor category.
2. The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 1, wherein, The prediction system further includes a risk assessment model validation module. The risk assessment model verification module is configured to use the Bootstrap resampling method to perform repeated sampling a predetermined number of times, thereby evaluating the stability of the risk assessment model. or The risk assessment model validation module is configured to use independent validation samples to evaluate the generalization ability of the risk assessment model.
3. The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 1, wherein, The prediction system further includes: The lesion differentiation model construction module is configured to construct a lesion differentiation model, wherein the lesion differentiation model is an elastic network penalized logistic regression model. The lesion differentiation model training module is configured to train the lesion differentiation model based on the risk score results and using candidate feature variables of CpG site methylation level to obtain lesion types. The lesion differentiation model validation module is configured to use ten-fold cross-validation to evaluate the discrimination performance and generalization ability of the lesion differentiation model.
4. The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 1, wherein, The acquisition module is further configured to simultaneously acquire exfoliated cells from the esophagus and cardia of the subject using a novel esophageal cell acquisition device.
5. The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 1, wherein, The preprocessing module is further configured as follows: The SegNet algorithm is used to quantitatively analyze the total number of esophageal and cardiac detached cells in the collected samples to determine whether the total number of esophageal and cardiac detached cells has reached the minimum detection standard of the predetermined number. The CenterNet algorithm is used to specifically identify gastric glandular epithelial cells or cell clusters in the sample to determine whether the sample enrichment material reaches the stomach during the sampling process. If the total number of exfoliated cells in the esophagus and cardia does not meet the predetermined minimum detection standard, or if the sample enrichment material does not reach the stomach during the sampling process, the subject should be resampled. If, in repeated sampling, the total number of exfoliated cells from the esophagus and cardia still does not reach the predetermined minimum detection standard, or if the sample enrichment material does not reach the stomach during sampling, the sample is excluded, thus obtaining multiple training samples. Each of the multiple training samples includes the value of a feature variable and the corresponding pathological diagnosis result. The pathological diagnosis result includes normal, low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma. The feature variables include cytological feature categories, CpG site methylation level categories, and the subject's epidemiological risk factor categories.
6. The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 5, wherein, The candidate feature variable screening module is further configured as follows: Using the diagnostic results of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia, and carcinoma as the target outcomes, the Boruta algorithm was used to screen 105 cytological features for each sample in the training samples. The top ten cytological features with the highest median importance scores were selected as candidate feature variables for the cytological feature category. The candidate feature variables for the cytological feature category include: the number of abnormal cells, the mean and variance of the DNA index of the top ten cells based on the DNA index, the mean of the nuclear width of the top ten cells based on the DNA index, the mean of the DNA index of all cells, the variance of the DNA index of all cells, the variance of the average gray level of the nuclear nucleus of all cells, the variance of the nuclear area of all cells, the variance of the nuclear width of all cells, and the variance of the nuclear height of all cells.
7. The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 5, wherein, The candidate feature variable screening module is further configured as follows: Based on the pathological diagnosis results, multiple training samples were divided into a case group and a control group, and the ratio of the number of samples in the case group to the number of samples in the control group was matched to 1:
3. The case group included multiple training samples with pathological diagnoses of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia and carcinoma, while the control group included multiple training samples with pathological diagnoses of normal. Pyrosequencing was used to quantitatively detect the methylation levels of specific CpG sites in all training samples from the case group and the control group. The specific CpG sites may include cg27284428, cg11798358, cg07880787, cg00011482, cg14633892 and cg04415798. The differences in methylation levels in training samples between the case group and the control group were analyzed; Based on the fact that the CpG site methylation level in the training samples of the case group was significantly higher than the corresponding CpG site methylation level in the training samples of the control group, candidate feature variables for CpG site methylation level categories were identified. Among them, the candidate feature variables for CpG site methylation level categories include: methylation levels of cg27284428, cg11798358, cg07880787, cg00011482, and cg14633892.
8. The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 5, wherein, The candidate feature variable screening module is further configured as follows: Based on the pathological diagnosis results, multiple training samples were divided into a case group and a control group, and the ratio of the number of samples in the case group to the number of samples in the control group was matched to 1:
3. The case group included multiple training samples with pathological diagnoses of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia and carcinoma, while the control group included multiple training samples with pathological diagnoses of normal. For all training samples in the case group and control group, the diagnosis results of low-grade intraepithelial neoplasia, high-grade intraepithelial neoplasia and carcinoma were used as the target outcomes. A logistic regression model was used to perform univariate analysis on epidemiological risk factors, which included age, sex, smoking status, alcohol consumption, hot food, pickled and dried food, fried food, missing teeth, history of upper gastrointestinal disease, body mass index and family history of cancer. Filter out P Epidemiological risk factors with values less than 0.05 were considered as candidate characteristic variables for the epidemiological risk factor category. Among these, the candidate characteristic variables for the epidemiological risk factor category included age, sex, and missing teeth.
9. The predictive system for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 1, wherein, The risk assessment model training module is further configured as follows: when training the risk assessment model using candidate feature variables selected from multiple training samples, the weight coefficient of each candidate feature variable is determined. Candidate feature variables whose weight coefficients are compressed to 0 are automatically removed from the risk assessment model, thereby obtaining the input feature variables of the risk assessment model.
10. A predictive method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions, comprising the following steps: Multiple samples were collected, including exfoliated cells from the esophagus and cardia of the subjects; Multiple collected samples are preprocessed to obtain multiple training samples; Candidate feature variables were selected from multiple training samples, including cytological feature categories, CpG site methylation level categories, and subjects' epidemiological risk factor categories. A risk assessment model is constructed, which is a LASSO Logistic regression model. The risk assessment model is trained using candidate feature variables selected from multiple training samples to obtain the input feature variables and risk score results for the risk assessment model. The input feature variables include: the number of abnormal cells in the cytological feature category; the mean and variance of the DNA index of the top ten cells based on the DNA index; the mean of the nuclear width of the top ten cells based on the DNA index; the variance of the DNA index of all cells; the variance of the average gray level of the nuclear nucleus of all cells; the variance of the nuclear area of all cells; and the variance of the nuclear height of all cells; the methylation levels of cg27284428, cg11798358, cg00011482, and cg14633892 in the CpG site methylation level category; and age, sex, and tooth loss in the epidemiological risk factor category.
11. The predictive method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 10, wherein, The prediction method further includes: The stability of the risk assessment model can be evaluated by using the Bootstrap resampling method, which involves repeated sampling a predetermined number of times; or The generalization ability of the risk assessment model is evaluated using independent validation samples.
12. The predictive method for early screening of esophageal squamous cell carcinoma, gastric cardia cancer, and precancerous lesions according to claim 10, wherein, The prediction method further includes: A lesion differentiation model is constructed, which is an elastic network penalized logistic regression model. Based on the risk score results, a lesion differentiation model was trained using candidate feature variables of CpG site methylation level to obtain lesion types; Ten-fold cross-validation was used to evaluate the discrimination performance and generalization ability of the lesion differentiation model.