A stepwise gastric function and gastric organic lesion combined screening system
Patent Information
- Application Number
- CN202610836444.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-25
AI Technical Summary
目前国内外已建立的胃部疾病筛查模型多聚焦于胃癌单一病种,主要采用Logistic回归、Cox比例风险模型等方法,仅整合年龄、血清标志物等少量指标,缺乏对饮食、生活习惯、基础疾病等临床多维度信息的综合考量,且未形成分层、阶梯式的筛查流程,存在筛查针对性不足、医疗资源浪费、漏诊风险较高等问题
本发明构建了问答式胃功能预测模型(GFPSBQA)和整合问卷与血清标志物的胃器质性病变预测模型(GDPSIQSM),GFPSBQA模型基于问卷特征向量构建,用于胃功能风险筛查;GDPSIQSM模型整合问卷(人口学、生活习惯、症状、既往病史)与血清胃功能四项数据,通过机器学习方法构建,用于胃器质性病变风险筛查,相比现有单一指标模型,预测更全面、准确。GFPSBQA模型灵敏度达91.16%,GDPSIQSM模型灵敏度达89.55%,漏诊机会少,能有效识别胃部疾病高风险人群,对早期胃部器质性病变的诊断和预防具有重要意义。
Smart Images

Figure CN122822320A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digestive system disease screening technology, specifically to a step-by-step combined screening system for gastric function and organic gastric lesions. Background Technology
[0002] Endoscopic tissue biopsy is the "gold standard" for diagnosing organic gastric lesions, but it is difficult to popularize as a routine screening method due to factors such as uneven distribution of medical resources and low public acceptance of endoscopy.
[0003] Serum gastric function tests (pepsinogen I, pepsinogen II, gastrin G17, and Helicobacter pylori antibody) are important indicators for detecting gastric diseases, and their levels are closely related to the severity of gastric mucosal lesions. Currently, most established gastric disease screening models, both domestically and internationally, focus on gastric cancer alone, primarily employing methods such as logistic regression and Cox proportional hazards models. These models integrate only a limited number of indicators, including age and serum biomarkers, lacking comprehensive consideration of multi-dimensional clinical information such as diet, lifestyle habits, and underlying diseases. Furthermore, they lack a tiered, step-by-step screening process, resulting in insufficient screening targeting, wasted medical resources, and a high risk of missed diagnoses.
[0004] Therefore, it is urgent to construct a combined predictive model for gastric function and organic gastric lesions that can integrate multi-dimensional clinical information, has high sensitivity, and is easy to operate, and to establish a tiered screening method to achieve hierarchical management of gastric diseases, improve the early diagnosis and treatment rate, and make rational use of medical resources such as gastroscopy and serum testing. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a step-by-step screening system for gastric function and organic gastric lesions. This model integrates questionnaire information and serum biomarker data, has high sensitivity and low false negative rate, and the step-by-step screening process can realize risk stratification management of gastric diseases, improve the rational use of gastroscopy and serum testing, reduce the waste of medical resources, and improve the early diagnosis and treatment rate of gastric diseases.
[0006] The main idea of the technical solution adopted in this invention is as follows: This invention constructs a question-and-answer gastric function prediction model (GFPSBQA) and a gastric organic lesion prediction model integrating questionnaires and serum biomarkers (GDPSIQSM). The GFPSBQA model is constructed based on questionnaire feature vectors and is used for gastric function risk screening; the GDPSIQSM model integrates four data points from the questionnaire and serum gastric function, constructed through machine learning methods, and is used for gastric organic lesion risk screening. Simultaneously, a tiered screening process is established based on the dual models: first, the GFPSBQA model performs initial screening using questionnaires, and high-risk individuals undergo serum testing; then, the GDPSIQSM model performs comprehensive screening, and high-risk individuals undergo gastroscopy and pathological examination, achieving stratified management of gastric disease risk. The models of this invention have high sensitivity and low false negative rate; the screening process is simple to operate and low in cost, which can improve the early diagnosis and treatment rate of gastric diseases, rationally utilize medical resources, and can be promoted and applied in medical institutions at all levels, especially at the grassroots level.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A step-by-step combined screening system for gastric function and organic gastric lesions, including The first-stage screening module is used to conduct an initial assessment of the gastric function of the target individuals and output the initial assessment results, including: The first data acquisition module is used to acquire questionnaire data of the target subjects. The questionnaire data includes demographic information, epidemiological history, dietary and lifestyle habits, previous examination and treatment history, and gastrointestinal symptoms in the past 3 months. The first assessment module is used to input the questionnaire data of the target subjects into the trained first machine learning model to output the first assessment result representing the gastric function of the target subjects. The first assessment result includes low risk and high risk. The judgment module is used to judge the results of the first assessment. If the first assessment result is determined to be low risk, the first assessment result is output directly; otherwise, the second-stage screening module performs a second assessment on the target and outputs the second assessment result, which includes: The second data acquisition module is used to acquire the serum gastric function four test data of the target object, including the detection data of PGⅠ, PGⅡ, G17 and HP antibody; The second assessment module is used to input the questionnaire data and serum gastric function test data of the target subjects into the trained second machine learning model to output the second assessment results that characterize the risk level of organic gastric lesions of the target subjects.
[0008] Based on the above technical solutions, the method for obtaining the well-trained first machine learning model is as follows: Obtain the questionnaire data of the sample subjects and their corresponding serum gastric function test result labels; Feature processing was performed on the questionnaire data of the sample subjects to identify key questions; The first feature vector is constructed based on the selected key issues; Using the first feature vector as input and the corresponding serum gastric function test result label as output, a preset first classification model is trained to obtain the first machine learning model.
[0009] Based on the above technical solutions, the method for obtaining the well-trained second machine learning model is as follows: Obtain questionnaire data, corresponding serum gastric function test data, and corresponding gastroscopy and pathological examination result labels from the sample subjects; The questionnaire data was combined with the corresponding serum gastric function test data to form a comprehensive feature set; Using the comprehensive feature set as input and the corresponding gastroscopy and pathology examination result labels as output, the model is trained using at least one machine learning algorithm selected from support vector machine, random forest or decision tree to obtain the second machine learning model.
[0010] Furthermore, based on the above technical solution, the second machine learning model is a decision tree model, and its training parameters are set as follows: global pruning mode, pruning severity of 75, and minimum number of records for each sub-branch of 5.
[0011] Furthermore, based on the above technical solution, the second machine learning model is a random forest model, and its training parameters are set as follows: the number of models constructed is 20, the maximum number of nodes in the tree growth is 50, the maximum tree depth is 10, and the minimum number of nodes is 5.
[0012] Furthermore, based on the above technical solution, the second machine learning model is a support vector machine model, and its training parameters are set as follows: the kernel type is radial basis function, the regularization parameter is 5, and the radial basis function gamma value is 0.1.
[0013] Furthermore, the judgment module is also used to: when the first assessment result is low risk, output a prompt message suggesting that the target object be screened again after a preset time interval.
[0014] Furthermore, the judgment module is also used to: when the target object obtains a low-risk assessment result three times in a row in the first assessment result, output a prompt message to conduct a second assessment of the target through the second-stage screening module.
[0015] The present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the functions of the system as described above.
[0016] The beneficial effects of this invention are: This invention constructs a question-and-answer gastric function prediction model (GFPSBQA) and an integrated questionnaire and serum biomarker gastric organic lesion prediction model (GDPSIQSM). The GFPSBQA model is built based on questionnaire feature vectors and is used for gastric function risk screening. The GDPSIQSM model integrates four data points (demographics, lifestyle habits, symptoms, and past medical history) and serum gastric function data, and is constructed using machine learning methods for gastric organic lesion risk screening. Compared with existing single-indicator models, the predictions are more comprehensive and accurate. The GFPSBQA model has a sensitivity of 91.16%, and the GDPSIQSM model has a sensitivity of 89.55%, with a low chance of missed diagnoses. It can effectively identify high-risk groups for gastric diseases and is of great significance for the diagnosis and prevention of early gastric organic lesions.
[0017] Meanwhile, a tiered screening process was established based on a dual-model approach. First, a questionnaire using the GFPSBQA model was used for initial screening, followed by serum testing for high-risk individuals. Then, a comprehensive screening using the GDPSIQSM model was conducted, with gastroscopy and pathological examination performed on high-risk individuals. This approach enables risk-stratified management of gastric diseases, avoids excessive testing for low-risk individuals, reduces the economic expenditure on blood tests and gastroscopy for patients, and significantly reduces the waste of medical resources.
[0018] Among them, the questionnaire screening method is simple to operate, low in cost, and highly operable. Serum test samples are easy to obtain and have high public acceptance. It can be promoted in areas with relatively weak medical resources, such as primary hospitals, communities, and rural areas, to improve the popularity of gastric disease screening.
[0019] The model is implemented using computer software, enabling rapid data entry and result output. It can assist primary care physicians in completing professional screening and assessment of organic gastric lesions, thereby improving the gastric disease screening capabilities of primary healthcare institutions. The model of this invention has high sensitivity and low false negative rate. The screening process is simple to operate and low in cost. It can improve the early diagnosis and treatment rate of gastric diseases, make rational use of medical resources, and can be promoted and applied in medical institutions at all levels, especially at the grassroots level. Attached Figure Description
[0020] Figure 1 Flowchart of a step-by-step screening model for gastric function and organic gastric lesions; Figure 2 This is a comparison between the prediction results of the model in this invention and clinical outcomes; Figure 3 A schematic diagram of the ROC curve for predicting serum gastric function using the GFPSBQA model; Figure 4 A schematic diagram of the ROC curve for predicting gastroscopy (including pathology) examination results using the GDPSIQSM model. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0022] Example 1 See Figures 1-4 This application discloses a step-by-step predictive model for gastric function and organic gastric lesions, including a question-and-answer gastric function predictive model (GFPSBQA) and a gastric organic lesion predictive model integrating questionnaires and serum biomarkers (GDPSIQSM). Based on the two models, a step-by-step screening system for gastric function and organic gastric lesions is constructed. The specific technical solution is as follows: 1. Construction method of question-and-answer gastric function prediction model (GFPSBQA) Questionnaire data was collected from patients seeking medical treatment. The questionnaire covered general demographic information (including: name, gender, age, ethnicity, weight, height, occupation, education level, daily language, and long-term residence), epidemiological history (including: family history of esophageal cancer, gastric cancer, and colorectal cancer (e.g., parents, brothers, sisters, etc., have gastric cancer), high-salt diet (daily salt intake greater than 10 grams / day, equivalent to one Coke bottle cap), intake of pickled foods, fried foods, smoked foods, fresh fruits, fresh vegetables, tea, smoking, alcohol consumption history, drinking water source, eating speed, history of diabetes, hypertension, hypertriglyceridemia, gastric mucosal erosion, gastric polyps, gastric mucosal atrophy, intestinal metaplasia, gastric ulcer, gastric intraepithelial neoplasia, gastric cancer, stromal tumor, duodenal polyps, duodenal bulb ulcer, reflux esophagitis, Barrett's esophagus, esophageal varices, and hiatal hernia). Principal component analysis, entropy method, or analytic hierarchy process (AHP) were used to perform weight analysis on each question in the questionnaire to identify key questions, low-relevance questions, and negative impact questions. Among these, key questions (epidemiological history (including family history of esophageal cancer, gastric cancer, colorectal cancer (e.g., parents, siblings, etc., with first-degree relatives having gastric cancer), diabetes, hypertension, hypertriglyceridemia, gastric mucosal erosion, gastric polyps, gastric mucosal atrophy, intestinal metaplasia, gastric ulcer, gastric intraepithelial neoplasia, gastric cancer, stromal tumor, duodenal polyps, duodenal bulb ulcer, reflux esophagitis, Barrett's esophagus, esophageal varices, hiatal hernia) were used as the core input variables to form the basic questionnaire model for GFPSBQA. Questions with low relevance (gender, age, ethnicity, weight, height, occupation, education level, daily language, long-term residence), fresh fruit intake, fresh vegetable intake, tea intake, drinking water source, and eating speed) were reviewed by clinical and statistical experts to determine whether to simplify and merge or directly remove them, in order to reduce the redundancy and complexity of questionnaire completion. Questions with negative impact (high salt diet (usual salt intake greater than 10 grams / day, equivalent to one Coke bottle cap), pickled food intake, fried food intake, smoked food intake, smoking, and drinking history) were re-optimized in terms of wording or adjusted in terms of option settings to avoid data collection bias due to flawed question design, which could affect the accuracy of model predictions.
[0023] Subsequently, based on the selected set of valid questions, an initial prediction model was constructed using the random forest algorithm, and weight analysis was performed. The model parameters were iteratively tuned using training set data, and the model's performance metrics, such as AUC, sensitivity, and specificity, were evaluated using validation set data to ensure that the model can effectively identify risk factors for gastric dysfunction and organic lesions.
[0024] High-quality questionnaire data features were constructed to form feature vectors based on key questions. Parameters such as gender, family history of gastric cancer, and smoking status were numerically converted using one-hot encoding or label encoding. Continuous variables (such as age, BMI, and daily salt intake) were standardized or normalized to eliminate the impact of dimensional differences on model training. Multiple-choice questions (such as frequency of pickled food intake and combinations of digestive symptoms) were transformed into multi-valued Boolean feature vectors. The processed feature vectors were input into a random forest model, and the top 30 core features were selected based on feature importance, further simplifying the model input dimensions.
[0025] Low-quality questionnaire data features are constructed by combining key questions with low-relevance questions. These low-quality features are basic questions that have a potential association with abnormal gastric function or organic gastric lesions, but they also contain a large number of redundant questions with very low relevance to the screening target, such as daily entertainment preferences and minor discomforts not related to the digestive system. If these low-quality features are directly included in the model, they are likely to introduce data noise and interfere with the accuracy of feature importance ranking. Therefore, in the early stages of model construction, it is necessary to use Pearson correlation coefficient analysis or variance filtering to remove questions with a relevance below the threshold, retaining only the sub-items in the key questions that are significantly related to the risk of gastric lesions, in order to reduce the negative impact of invalid features on model training and form a feature vector. Based on the aforementioned feature vectors, a question-and-answer gastric function prediction model was constructed and trained using the support vector algorithm. The sample data was divided into training and validation sets in a 7:3 ratio, with diagnostic and pathological results as labels. Model parameters (maximum tree depth, minimum number of sample splits, etc.) were optimized and tuned using a grid search method. After training, the model performance was comprehensively evaluated using metrics such as the area under the receiver operating characteristic (ROC) curve (AUC), sensitivity, specificity, and accuracy. This ensured that the model's AUC value on the validation set was ≥0.85 and its sensitivity was ≥90%, meeting the needs of clinical screening. The final GFPSBQA model can output a gastric function risk level (low or high risk) based on the questionnaire data of the target population, providing an initial screening basis for subsequent tiered screening procedures.
[0026] 2. Methodology for constructing a gastric organic lesion prediction model integrating questionnaires and serum biomarkers (GDPSIQSM) The questionnaire data and serum gastric function test data (PGⅠ, PGⅡ, G17, HP antibody) were integrated to form a comprehensive dataset; The comprehensive dataset was divided into training, testing, and validation groups using a random method, and a 10-fold cross-validation experiment was conducted. We employed machine learning methods such as support vector machines, random forests, and decision trees, and used a validation set and a test set to verify the learning effect and application performance of the model on the training set. The decision tree model utilizes C5.0 decision tree nodes to build the model. To ensure the model is built using only training partition data, partition data is selected. Cross-validation is chosen, and 10 folds are set to estimate the accuracy of the built model. The final output model type is selected as decision tree. Expert mode is chosen for training to better fit the actual target. Global pruning mode is used to achieve a concise yet accurate small tree pruning severity, with a pruning severity of 75. To prevent overtraining with noisy data, the minimum number of records per sub-branch is increased to 5, limiting the number of splits in any branch of the tree. Misclassification cost analysis is not used because the specific cost of misclassification of the suitable population for gastroscopy cannot be accurately estimated in practice. The model for calculating the importance of predictor variables is selected.
[0027] Use the Random Forest model nodes to build a Random Forest model. Ensure the training model data matches real-world conditions; set the model sample size to 1, making the sample size equal to the original training data partitions. Considering training time and model fit, set the number of models to 20, the maximum number of tree nodes to 50, the maximum tree depth to 10, and the minimum node size to 5. Select "Stop building trees when accuracy no longer improves." Advanced data settings: Set the maximum percentage of missing values to 70%; if any variable has more than 70% missing values, it will be excluded from the model. Set the exclusion threshold for fields where most values in a single category exceed 95% to 95%; if any category has more than 95% of its values, it will be excluded from the model. Set the maximum number of categories in a field to 50; if there are more than 50 categories, they will be excluded from the model. Set the minimum field variance to 0.05; if the variance coefficient of a continuous variable is 0.05, it will be excluded from the model. Set the number of tiers to 10, arranging the observations of the continuous variable in ascending order and dividing them into 10 equal parts based on the number of observations, with each part considered a tier.
[0028] Use the Support Vector Machine (SVM) model node to build the SVM model. Ensure that the model is built using only the data from the training partition; select the option to build the model using partitioned data. Based on practical experience, select expert mode for model building, with the following parameter settings: Set "Append All Probabilities" to "No," displaying the probability of predicted values only for the nominal or flag target field; set the termination condition value to 1.0E–3 to determine the criterion for stopping the optimization algorithm; reduce the regularization parameter from the preset value of 10 to 5 to reduce the risk of overfitting without significantly reducing the classification accuracy of the training data; set the kernel type to Radial Basis Function (RBF), and increase the RBF gamma from the preset value of 0.01 to 0.1, approaching the maximum value, to increase model accuracy. Select the model for calculating the importance of predictor variables.
[0029] After prediction, the model information was evaluated in detail by using the coincidence matrix and evaluation metrics to calculate the model's sensitivity, specificity, positive predictive value, negative predictive value, accuracy, and area under the receiver operating characteristic curve (AUC).
[0030] A feature combination analysis method was established, and a support vector algorithm was selected to construct a prediction model for gastric organic lesions, which was used to predict the risk level (low risk / high risk) of patients with gastric organic lesions.
[0031] 3. A stepwise screening method for gastric function and organic gastric lesions based on a dual-model approach, such as... Figure 1 This invention demonstrates two stages of a tiered screening process. The first stage is screening using the GFPSBQA model questionnaire, with follow-up or serum testing conducted based on low / high risk results. The second stage is comprehensive screening using the GDPSIQSM model, with follow-up or gastroscopy + pathological examination conducted based on low / high risk results, clarifying the connection and processing methods of each step.
[0032] Phase 1: GFPSBQA model questionnaire screening For patients presenting with gastrointestinal symptoms, a questionnaire was completed via face-to-face interview. The questionnaire data was then input into the GFPSBQA model to output a low-risk / high-risk conclusion for gastric function. Low-risk individuals: undergo GFPSBQA model questionnaire screening again after 1 year; if they are low-risk for 3 consecutive years, they will undergo serum gastric function four-item test and enter the second stage GDPSIQSM model screening. High-risk individuals: Immediately undergo serum gastric function tests (four items) and proceed to the second phase of GDPSIQSM model screening.
[0033] The definitions of normal and abnormal gastric function test results are shown in Table 1.
[0034]
[0035]
[0036] Phase Two: Comprehensive Screening Using the GDPSIQSM Model Patient questionnaire data and serum gastric function test data were input into the GDPSIQSM model to output a low-risk / high-risk conclusion for organic gastric lesions: Low-risk individuals: undergo the first-stage GFPSBQA model questionnaire screening again after 1 year; if they are low-risk for 3 consecutive years, they will undergo detailed gastroscopy and pathological examination. High-risk individuals: Immediately undergo detailed gastroscopy and pathological examination, and take appropriate measures according to clinical guidelines based on the examination results.
[0037] The definitions of normal and abnormal gastroscopy (including pathology) results are shown in Table 2.
[0038] Table 2 Definitions of Normal and Abnormal Gastroscopy (Including Pathology) Results
[0039] Example 2: Specific Application Case The present invention will be further described in detail below with reference to specific experimental data. The scope of protection of the present invention is not limited to the following embodiments.
[0040] (a) Research Subjects This study selected 1677 patients admitted to the First Affiliated Hospital of Guangdong Pharmaceutical University, Foshan Traditional Chinese Medicine Hospital, and Nanxiong Traditional Chinese Medicine Hospital from June 2015 to March 2023 with gastrointestinal symptoms. Inclusion criteria: ① Age ≥ 18 years; ② All patients underwent serum gastric function tests (PGⅠ, PGⅡ, G17, and HP antibody); ③ All patients underwent gastroscopy and histological examination; ④ All patients signed informed consent forms. Exclusion criteria: ① Patients with mental disorders or other reasons preventing cooperation with the investigation; ② Patients with a history of other malignant tumors.
[0041] (II) Model Construction and Validation Complete the model building according to the above GFPSBQA and GDPSIQSM model building methods; The patients were interviewed by designated personnel to complete questionnaires. The questionnaire data was entered into the GFPSBQA model. At the same time, the patients' serum gastric function was detected using ELISA kits and detection instruments, including PGI, PGII, G-17 and HP antibody detection. The questionnaire data and serum data were entered into the GDPSIQSM model. Blind gastroscopy with pathological examination performed by a specialist is considered the gold standard for clinical diagnosis. The endoscopist is proficient in gastroscopy and is unaware of the patient's GFPSBQA and GDPSIQSM model interpretations. If the patient consents to a mucosal biopsy and signs a consent form before the gastroscopy, a routine mucosal biopsy is performed. Otherwise, a mucosal biopsy is not performed. For patients who consent to pathological examination, the biopsy criteria are as follows: If suspicious lesions are found endoscopically, a biopsy tissue sample is taken using disposable biopsy forceps after magnified chromoendoscopic observation. If no potentially lesioned mucosa is found endoscopically, a biopsy is routinely performed 2 cm from the pylorus on the greater curvature of the gastric antrum.
[0042] Statistical analysis was performed using Excel 2010 and SPSS 23.0 software. Normally distributed quantitative data were described using mean ± standard deviation, and t-tests were used for comparisons between groups. Non-normally distributed quantitative data were described using median and quartiles, and Mann-Whitney U tests were used for comparisons between groups. Qualitative data were described using frequency and percentage, and chi-square tests were used for comparisons between groups. The predictive power of the model was evaluated using sensitivity, specificity, positive predictive value, negative predictive value, area under the ROC curve, and Youden's index. A p-value < 0.05 was considered statistically significant.
[0043] (III) Implementation Results 1. Basic information of all patients 1.1 Basic Patient Information From June 2015 to March 2023, a total of 1677 patients with complete questionnaires, serum gastric function tests, and gastroscopy examinations including pathological data were included. Among them, 801 were male (47.8%) and 876 were female (52.2%), with a male-to-female ratio of 0.9:1.0. 274 patients (16.3%) were 40 years of age or younger, and 1403 patients (82.7%) were over 40 years of age (Table 3).
[0044]
[0045]
[0046]
[0047] 1.2. Distribution of patient gastroscopy diagnoses: Of the 1677 subjects included in the study, 1081 cases were positive for gastroscopy, with a positive detection rate of 64.46%. The most common diagnoses were gastric mucosal erosion, gastric polyps, and gastric mucosal atrophy, with 297, 286, and 144 cases, respectively, accounting for 17.71%, 17.05%, and 8.59% of the total cases. Other gastric-related diagnoses included intestinal metaplasia (69 cases), gastric ulcer (65 cases), gastric intraepithelial neoplasia (52 cases), gastric cancer (21 cases), and stromal tumor (4 cases). Duodenal-related diagnoses included duodenal polyps (28 cases), duodenal bulb ulcers (136 cases), reflux esophagitis (51 cases), Barrett's esophagus (13 cases), esophageal varices (27 cases), and hiatal hernia (23 cases) (Table 4).
[0048]
[0049] 2. Based on the conclusions output by the GFPSBQA and GDPSIQSM models, groups were formed, and the differences in the basic characteristics of the two groups of patients were compared. 2.1 Based on the output results of the GFPSBQA model, patients were divided into a low-risk group and a high-risk group for gastric function. The differences in the basic characteristics of the two groups were as follows: Of the 1,677 participants included, 266 were in the low-risk group and 1,411 were in the high-risk group. A comparative analysis of the baseline characteristics of the two groups revealed significant differences (P < 0.05) between the low-risk and high-risk groups in terms of place of residence, occupation, intake of pickled foods, intake of fried and smoked foods, intake of fresh vegetables, tea intake, smoking history, diabetes, hypertension, and hypertriglyceridemia (Table 5).
[0050]
[0051]
[0052]
[0053] 2.2 Based on the output results of the GDPSIQSM model, patients were divided into a low-risk group and a high-risk group for organic gastric lesions. The differences in basic characteristics between the two groups were as follows: Of the 1,677 participants included, 311 were in the low-risk group for organic gastric lesions and 1,366 were in the high-risk group. A comparative analysis of the baseline characteristics of the two groups revealed significant differences between the low-risk and high-risk groups in terms of gender, age, place of residence, education level, family history of cancer, high-salt diet, intake of pickled foods, fried foods, smoked foods, fresh fruit, fresh vegetables, tea, smoking history, alcohol consumption, eating speed, diabetes, hypertension, hypertriglyceridemia, HP (antibody), PGⅠ, and PGⅡ (P < 0.05) (Table 6).
[0054]
[0055]
[0056]
[0057]
[0058] 3. The output of the predictive model is compared with the results of blood tests for gastric function and gastroscopy, including pathological results: From December 2015 to June 2023, a total of 1677 patients had their data entered, including those with complete data on gastric function prediction questionnaires, blood gastric function tests, and gastroscopy examinations (including pathological results). The GFPSBQA model identified 1411 patients as high-risk and 266 as low-risk; 1289 patients had abnormal blood gastric function test results, and 388 had normal results. The GDPSIQSM model identified 1366 patients as high-risk and 311 as low-risk; 1081 patients had abnormal gastroscopy examinations (including pathological results), and 596 had normal results. Figure 2 This is a comparison chart between model predictions and clinical outcomes. The chart uses case number as the horizontal axis to show the comparison trends between the prediction results (low / high risk) of the GFPSBQA model and the GDPSIQSM model and the results of serum gastric function tests and gastroscopy + pathological examinations (normal / abnormal), respectively, intuitively demonstrating the consistency between model predictions and actual clinical outcomes.
[0059] 3.1 Accuracy of GFPSBQA output results: The GFPSBQA prediction model showed a sensitivity of 91.16%, a specificity of 39.18%, a positive predictive value of 83.27%, a negative predictive value of 57.14%, an area under the ROC curve of 0.702, a 95% confidence interval of (0.665, 0.740), and a Youden index of 0.303 (Table 7). Figure 3 ). Figure 3 The ROC curve of the GFPSBQA model predicting serum gastric function results is shown in the figure. The horizontal axis represents the false positive rate and the vertical axis represents the true positive rate. The ROC curve of the support vector algorithm model has an area under the curve of 0.702 and a 95% confidence interval of (0.665, 0.740), which reflects the predictive efficacy of the model for serum gastric function results.
[0060]
[0061] 3.2 Accuracy of GDPSIQSM output results: The GDPSIQSM prediction model showed a sensitivity of 89.55%, a specificity of 33.22%, a positive predictive value of 70.86%, a negative predictive value of 63.67%, an area under the ROC curve of 0.672, a 95% confidence interval of (0.639, 0.707), and a Youden index of 0.228 (Table 8). Figure 4 ). Figure 4The ROC curve for predicting gastroscopy (including pathology) examination results using the GDPSIQSM model is shown. The horizontal axis represents the false positive rate, and the vertical axis represents the true positive rate. The ROC curve of the support vector algorithm model has an area under the curve of 0.672 and a 95% confidence interval of (0.639, 0.707), which reflects the model's predictive efficacy for organic gastric lesions.
[0062]
[0063] In summary, the efficacy of the GFPSBQA model was as follows: sensitivity of 91.16%, specificity of 39.18%, positive predictive value of 83.27%, negative predictive value of 57.14%, area under the ROC curve of 0.702 (95% CI: 0.665-0.740), and Youden index of 0.303. GDPSIQSM model performance: Sensitivity for predicting gastroscopy + pathology examination results was 89.55%, specificity was 33.22%, positive predictive value was 70.86%, negative predictive value was 63.67%, area under the ROC curve was 0.672 (95% CI: 0.639-0.707), and Youden index was 0.228; Results of the tiered screening method: Among 1677 patients, 266 were identified as low-risk and 1411 as high-risk using the GFPSBQA model, and 311 as low-risk and 1366 as high-risk using the GDPSIQSM model. After gastroscopy and pathological examination, the positive detection rate of gastroscopy in high-risk patients was 64.46%, effectively identifying patients with organic lesions such as gastric mucosal erosion, polyps, atrophy, and gastric cancer, thus achieving accurate screening of high-risk groups.
[0064] (iv) Application instructions for the model and screening methods The GFPSBQA and GDPSIQSM models of this invention can be implemented using computer software and developed into a medical terminal system. This system supports manual entry / batch import of questionnaire and serum test data, and automatically outputs risk level conclusions. The tiered screening method is suitable for hospitals at all levels, community health service centers, physical examination centers, and other medical institutions, and is especially suitable for primary healthcare institutions to conduct initial screening for gastric diseases. The questionnaire content can be fine-tuned according to regional lifestyles and disease spectrum characteristics. The model can be iteratively optimized by supplementing datasets from different regions and populations to improve predictive efficacy.
[0065] This study presents an innovative disease management model that uses a three-step step-by-step approach to predict the risk of gastric function and gastric diseases, and employs a stratified management and step-by-step gastric disease screening model for patients with different risks of gastric disease.
[0066] This study's model incorporates questionnaire surveys and serological testing of gastric function. In practical clinical applications, questionnaire screening, based on the entire population, is simple, highly operable, low-cost, and feasible. After questionnaire screening, low-risk individuals are advised to undergo regular follow-up, while high-risk individuals are recommended to undergo further serological screening. Serum biomarkers are readily available from blood samples, and other indicators can be tested simultaneously, making them well-accepted by the public. This allows general practitioners in rural areas to combine questionnaire surveys with professional screening and assessment of organic gastric lesions—a quick, convenient, and low-cost approach. Following serological and questionnaire screening, low-risk individuals are advised to undergo regular follow-up, while high-risk individuals are recommended to undergo detailed gastroscopy. This targeted approach saves medical resources and aligns with the screening goals of "cost-effectiveness, social affordability, and public acceptance." The results show that the predictive performance of the GFPSBQA and GDPSIQSM models, compared with blood sample gastric function and gastroscopy examination results including pathological diagnosis, demonstrates good predictive performance for gastric diseases. The GFPSBQA and GDPSIQSM models exhibit high sensitivity, meaning a low chance of missed diagnoses, which is significant for the diagnosis and prevention of early organic gastric lesions.
[0067] This model is based on a population seeking medical attention for gastrointestinal symptoms. This study also validates the model using this population. If it is widely applied in clinical practice, it can reduce the workload of blood tests and gastroscopy for patients, thereby reducing their financial burden and the waste of medical resources.
[0068] In summary, GFPSBQA and GDPSIQSM demonstrate excellent predictive capabilities for gastric diseases in patients presenting with gastrointestinal symptoms. They can help patients detect functional and organic lesions promptly, reducing the waste of medical resources and showing broad clinical application prospects. They are expected to make greater contributions to the screening and early diagnosis of functional and organic gastric diseases in the future. Currently, the models are mainly used in hospital settings. Further optimization of data collection is possible, such as extending their application to communities and rural areas, and promoting their use among individuals with gastrointestinal symptoms who have difficulty accessing medical care. This would increase data volume and diversity, expand the sample size while continuously improving system accuracy, and simplify system operation techniques, ultimately providing gastric disease screening services to a wider population.
[0069] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A step-by-step combined screening system for gastric function and organic gastric lesions, characterized in that, include: The first-stage screening module is used to conduct an initial assessment of the gastric function of the target individuals and output the initial assessment results, including: The first data acquisition module is used to acquire questionnaire data of the target subjects. The questionnaire data includes demographic information, epidemiological history, dietary and lifestyle habits, previous examination and treatment history, and gastrointestinal symptoms in the past 3 months. The first assessment module is used to input the questionnaire data of the target subjects into the trained first machine learning model to output the first assessment result representing the gastric function of the target subjects. The first assessment result includes low risk and high risk. The judgment module is used to judge the results of the first assessment. If the first assessment result is determined to be low risk, the first assessment result is output directly; otherwise, the second-stage screening module performs a second assessment on the target and outputs the second assessment result, which includes: The second data acquisition module is used to acquire the serum gastric function four test data of the target object, including the detection data of PGⅠ, PGⅡ, G17 and HP antibody; The second assessment module is used to input the questionnaire data and serum gastric function test data of the target subjects into the trained second machine learning model to output the second assessment results that characterize the risk level of organic gastric lesions of the target subjects.
2. The system according to claim 1, characterized in that, The method for obtaining the first trained machine learning model is as follows: Obtain the questionnaire data of the sample subjects and their corresponding serum gastric function test result labels; Feature processing was performed on the questionnaire data of the sample subjects to identify key questions; The first feature vector is constructed based on the selected key issues; Using the first feature vector as input and the corresponding serum gastric function test result label as output, a preset first classification model is trained to obtain the first machine learning model.
3. The system according to claim 2, characterized in that, The method for obtaining the trained second machine learning model is as follows: Obtain questionnaire data, corresponding serum gastric function test data, and corresponding gastroscopy and pathological examination result labels from the sample subjects; The questionnaire data was combined with the corresponding serum gastric function test data to form a comprehensive feature set; Using the comprehensive feature set as input and the corresponding gastroscopy and pathology examination result labels as output, the model is trained using at least one machine learning algorithm selected from support vector machine, random forest or decision tree to obtain the second machine learning model.
4. The system according to claim 3, characterized in that, The second machine learning model is a decision tree model, and its training parameters are set as follows: global pruning mode, pruning severity of 75, and minimum number of records in each sub-branch of 5.
5. The system according to claim 3, characterized in that, The second machine learning model is a random forest model, and its training parameters are set as follows: the number of models constructed is 20, the maximum number of nodes in the tree growth is 50, the maximum tree depth is 10, and the minimum number of nodes is 5.
6. The system according to claim 3, characterized in that, The second machine learning model is a support vector machine model, and its training parameters are set as follows: the kernel type is radial basis function, the regularization parameter is 5, and the radial basis function gamma value is 0.
1.
7. The system according to claim 3, characterized in that, The judgment module is also used to: when the first assessment result is low risk, output a prompt message suggesting that the target object be screened again after a preset time interval.
8. The system according to claim 7, characterized in that, The judgment module is also used to: when the target object receives a low-risk assessment result three times in a row in the first assessment result, output a prompt message to conduct a second assessment of the target through the second-stage screening module.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program performs the functions of the system described in any one of claims 1 to 8.