Remote metastasis prediction system and method based on multiple examinations of gastric cancer patient
By integrating multimodal data and multivariate algorithms, the imaging, endoscopic and blood test data of gastric cancer patients were screened, and an interpretable machine learning model was constructed, which solved the misdiagnosis and misdiagnosis of the prediction of distal metastasis of gastric cancer in the existing technology, and achieved real-time and accurate prediction and clinical support.
Patent Information
- Application Number
- CN202510976138.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing gastric cancer distal metastasis prediction model relies on single mode inspection data, resulting in information island effect, making it difficult to fully capture complex biological mechanisms, there is a risk of misdiagnosis and misdiagnosis, and lacks multi-algorithm cross-validation and interpretability, making it difficult to achieve real-time prediction and clinical decision support.
Multimodal data on imaging, endoscopy and blood examination are integrated, and features are screened using LassoCV, recursive feature elimination and Boruta algorithms, and model training is combined with multiple machine learning algorithms such as Logistic classification and XGBoost. Multi-dimensional evaluation curves are drawn through hierarchical nested cross-validation and automatic mesh parameter search to generate interpretable prediction results.
It improves the accuracy and reliability of remote metastasis prediction, reduces the risk of misdiagnosis and misdiagnosis, realizes real-time prediction and clinical decision-making support, simplifies the diagnosis and treatment process, and enhances the clinical credibility of the model.
Smart Images

Figure CN120496857A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical data analysis and machine learning technology, and in particular to a system and method for predicting distal metastasis based on multiple examinations of gastric cancer patients. Background Art
[0002] Gastric cancer is a highly prevalent malignant tumor worldwide, and the early and accurate prediction of its distant metastasis directly impacts the formulation of patient treatment plans and prognosis assessment. In current clinical practice, traditional prediction of distant metastasis of gastric cancer mainly relies on single-modality examination data, such as imaging examinations to present tumor anatomical characteristics, endoscopic pathology to determine the degree of tissue differentiation, or blood indicators to reflect the body's status. However, this single-dimensional assessment method has significant limitations. Although imaging examinations can intuitively display tumor size and lymph node metastasis, they are difficult to capture the molecular biological characteristics of tumor cells; although endoscopic pathology can clearly determine the degree of tumor differentiation and infiltration depth, it cannot dynamically reflect the body's immune-inflammatory microenvironment; and although blood indicators such as tumor markers and blood routine tests can quantify the body's physiological state, due to the lack of specificity of a single indicator, it is difficult to accurately associate metastasis risk. This information island effect of single-modality data makes it difficult for traditional prediction methods to fully capture the complex biological mechanisms of distant metastasis of gastric cancer, often resulting in missed or misdiagnoses, and failing to meet the needs of clinical precision diagnosis and treatment.
[0003] With the expanding application of machine learning in healthcare, existing gastric cancer metastasis prediction models, while attempting to incorporate algorithmic modeling, still face numerous technical bottlenecks. Firstly, most models rely solely on a single algorithm for feature selection and modeling, such as the Lasso or recursive feature elimination, lacking cross-validation across multiple algorithms. This results in instability in feature selection and susceptibility to data noise. Secondly, when dealing with class imbalance in gastric cancer metastasis data, traditional models often prioritize accuracy, neglecting the ability to identify minority metastatic cases, leading to missed diagnoses of high-risk patients. Existing models generally lack interpretability, making black-box prediction models difficult to gain clinical trust. They also lack generalization capabilities for multi-center, heterogeneous data and are unable to adapt to the variability of diagnostic and treatment data across different medical institutions. Furthermore, existing systems are often limited to laboratory research and lack deep integration with clinical diagnostic and treatment processes. This results in low data integration efficiency and high computational costs, making it difficult to achieve real-time prediction of metastasis risk and clinical decision support. These technical limitations pose multiple challenges to the accuracy, reliability, and practicality of existing gastric cancer distant metastasis prediction models in real-world clinical applications. Summary of the Invention
[0004] In order to address the deficiencies of the existing technology, the present invention discloses a distal metastasis prediction system and method based on multiple examinations of gastric cancer patients, which integrates multimodal data, fuses multiple algorithms and has both interpretability.
[0005] The present invention discloses a distal metastasis prediction system based on multiple examinations of gastric cancer patients, which includes a data acquisition module: for collecting preoperative data of gastric cancer patients;
[0006] Data preprocessing module: used to clean and normalize the collected gastric cancer patient data;
[0007] Feature selection module: Using three machine learning algorithms, LassoCV, recursive feature elimination, and Boruta, the preprocessed data is filtered to identify the overlapping feature indicators most relevant to gastric cancer metastasis and construct a feature vector.
[0008] Model training module: This module uses multiple machine learning algorithms, with the selected feature vectors as input and the patient's actual metastasis status as output. Parameters are automatically grid-searched and validated through stratified nested cross-validation. Standardized random initialization is used to establish multiple machine learning models for predicting distant metastasis of gastric cancer based on the selected feature vectors.
[0009] Model screening module: used to screen the best machine learning model among multiple machine learning models for predicting distant metastasis of gastric cancer, and conduct multi-dimensional evaluation and optimization of the machine learning model by drawing relevant curves;
[0010] Decision support module: Develop an online prediction website for the generated machine learning model or connect to the clinical data database on the hospital intranet to automatically generate prediction results based on feature variables.
[0011] Furthermore, gastric cancer patient data include imaging data, endoscopic data, and blood test report data;
[0012] Imaging data included: tumor location, tumor size, invasion depth, and lymph node metastasis;
[0013] Endoscopic data included: tumor location, tumor size, pathological type, Lauren classification, degree of differentiation, immunohistochemistry, depth of invasion, and lymph node metastasis;
[0014] Blood test report data includes: tumor markers, routine blood indicators and nutritional indicators.
[0015] Furthermore, the blood test report data includes five optimization indices:
[0016] Tumor marker index TMI, nutritional index PNI, platelet-lymphocyte ratio PLR, lymphocyte-monocyte ratio LMR and neutrophil-lymphocyte ratio NLR.
[0017] Furthermore, the model screening module screens the best machine learning model among multiple machine learning models for predicting distant metastasis of gastric cancer. The evaluation criteria are: area under the receiver operating characteristic curve greater than 0.9, accuracy>0.8, sensitivity>0.8, specificity>0.8, F1>0.6, and area under the precision-recall curve>0.7. Based on the evaluation criteria, the model that meets the criteria and has the largest area under the curve is selected as the best model.
[0018] Furthermore, drawing the correlation curve includes: drawing a calibration curve to evaluate the consistency between the model predicted probability and the actual probability;
[0019] Draw a net benefit decision curve to evaluate the clinical practicality of the model at different decision thresholds and help select the optimal decision threshold;
[0020] Draw a learning curve to show how model performance changes with the amount of training data or training time, which is used to diagnose whether the model is underfitting or overfitting;
[0021] Draw the KS curve to evaluate the ability of the classification model to distinguish between positive and negative samples;
[0022] A SHAP diagram was drawn to quantify the contribution of each feature to the model prediction and improve the clinical practicality of the diagnostic model.
[0023] Furthermore, a variety of machine learning algorithms include: Logistic classification, XGBoost classification, LightGBM classification, random forest classification, AdaBoost classification, decision tree classification, GBDT classification, Gaussian naive Bayes classification, neural network classification, support vector machine classification, and k-nearest neighbor classification.
[0024] The present invention discloses a method for predicting distal metastasis based on multiple examinations of gastric cancer patients, using any one of the aforementioned distal metastasis prediction systems based on multiple examinations of gastric cancer patients, comprising:
[0025] S1: Collect preoperative imaging data, endoscopic data, and blood test report data of gastric cancer patients through the data acquisition module;
[0026] S2: Use the data preprocessing module to clean and normalize the collected data and remove outliers and missing values;
[0027] S3: The feature selection module uses three machine learning algorithms, LassoCV, recursive feature elimination, and Boruta, to screen the overlapping feature indicators most relevant to gastric cancer distant metastasis from the preprocessed data and construct a feature vector.
[0028] S4: Using the model training module, the screened feature vectors are used as input and the patient's actual metastasis status is used as output. Through automatic grid parameter search and hierarchical nested cross-validation, a prediction model is established based on multiple machine learning algorithms such as logistic classification, XGBoost classification, and random forest classification.
[0029] S5: Through the model screening module, the best model is selected according to the preset evaluation criteria, and the model is evaluated and optimized in multiple dimensions by drawing calibration curves, decision curves, learning curves, KS curves and SHAP graphs;
[0030] S6: Develop the optimal model into an online prediction website or connect it to the hospital intranet using the decision support module to automatically generate metastasis risk prediction results based on the input feature variables.
[0031] Furthermore, after S5 selects the best model through the model screening module, it also conducts internal and external test set verification;
[0032] Internal test set validation includes: inputting a test set from the same center, independent of the training and validation sets, into the best model, and verifying the generalization ability under the same data distribution by calculating indicators such as AUC and accuracy;
[0033] External test set validation includes: inputting the test set of another center into the model, comparing the AUC values of data from different centers, and evaluating the actual application effect of the model in unknown data distribution.
[0034] Furthermore, the data collected in S1 includes:
[0035] Imaging data: tumor location, tumor size, invasion depth, and lymph node metastasis;
[0036] Endoscopic data: tumor location, tumor size, pathological type, Lauren classification, degree of differentiation, immunohistochemistry, invasion depth, and lymph node metastasis;
[0037] Blood test report data: tumor markers, routine blood indicators and nutritional indicators, and the blood test report data is optimized into five indices: tumor marker index, nutritional index, platelet-lymphocyte ratio, lymphocyte-monocyte ratio and neutrophil-lymphocyte ratio.
[0038] Furthermore, the specific steps for screening overlapping feature indicators in S3 are as follows: feature screening is performed using the LassoCV, RFECV, and Boruta algorithms respectively, and the intersection features of the results of the three algorithms are taken;
[0039] S4 includes various machine learning algorithms, including Logistic classification, XGBoost classification, LightGBM classification, random forest classification, AdaBoost classification, and support vector machine classification. Parameter optimization uses automatic grid parameter search, and the verification method is stratified nested cross-validation.
[0040] Beneficial effects of the present invention:
[0041] This patent integrates multimodal data from imaging, endoscopy, and blood tests of gastric cancer patients to construct a multi-algorithm collaborative feature screening strategy and machine learning model, effectively improving the accuracy and reliability of distant metastasis prediction. Multi-algorithm cross-validation screens core features to avoid bias from a single algorithm. Automatic grid search is combined with hierarchical nested cross-validation to optimize model parameters. Multi-dimensional evaluations such as calibration curves and decision curves ensure the model's performance under imbalanced data, solving the problem of insufficient prediction accuracy of traditional single-examination methods.
[0042] At the clinical application level, this patent integrates the optimized model into an online platform or hospital intranet to realize automatic analysis of patient data and real-time prediction of transfer risk, improve diagnosis and treatment efficiency and reduce the risk of missed diagnosis and misdiagnosis. The standardized prediction process reduces the diagnostic bias caused by individual experience differences. Explanatory analysis such as SHAP graphs enhances the clinical credibility of the model. Moreover, the data are all from routine examination items, without additional cost, which is convenient for clinical promotion and application, and provides a scientific data-driven decision-making basis for the formulation of personalized diagnosis and treatment plans for gastric cancer patients. The personalized diagnosis and treatment plan provided by the present invention for gastric cancer patients provides a quantitative basis for clinical decision-making, and the specific diagnosis and treatment plan is determined by the doctor. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a lasso model coefficient graph obtained by the LassoCV algorithm in the embodiment of the implementation manner of this application.
[0044] Figure 2 In the embodiment of the present application, a recursive feature elimination cross-validation curve is obtained by the RFECV algorithm.
[0045] Figure 3 This is a Boruta feature importance evaluation graph obtained by Boruta in the embodiment of the implementation manner of this application.
[0046] Figure 4 This is a Venn diagram of overlapping features related to distant metastasis of gastric cancer screened out by the three feature screening methods of LassoCV, RFECV, and Boruta in the examples of the embodiments of the present application.
[0047] Figure 5 Shown in the figure is a ROC curve diagram of a training set in an embodiment of the present application.
[0048] Figure 6 Shown in is a ROC curve diagram of a validation set in an embodiment of the present application.
[0049] Figure 7 Shown in the figure is a training set precision-recall curve, i.e., a training set PR curve, in an embodiment of the present application.
[0050] Figure 8 Shown in the figure is a verification set precision-recall curve, i.e., a verification set PR curve, in an embodiment of the present application.
[0051] Figure 9 Shown in the figure is a learning curve diagram of a logistic regression model in an embodiment of the present application, which shows the trend of model performance changing with the number of training samples.
[0052] Figure 10 Shown in FIG is a calibration curve diagram of a logistic regression model in an embodiment of the present application.
[0053] Figure 11 Shown in the figure is a test set decision curve, namely a DCA curve diagram, in an embodiment of the present application.
[0054] Figure 12 Shown in is a KS statistic graph of a test set in an embodiment of the present application.
[0055] Figure 13 Shown in FIG is a SHAP summary diagram of an embodiment of the present application.
[0056] Figure 14 Shown in FIG is a SHAP force-directed graph in an embodiment of the present application.
[0057] Figure 15 Shown in FIG is another SHAP force-directed graph in an embodiment of the present application.
[0058] Figure 16 Shown in the figure is a ROC curve diagram of an internal test set in an embodiment of the present application.
[0059] Figure 17 Shown in the figure is a graph of the receiver operating curve of an external test set in an embodiment of the present application.
[0060] Figure 18 Shown in the figure is a diagram of formulating corresponding examination and treatment plans based on prediction results in an embodiment of the present application. DETAILED DESCRIPTION
[0061] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the specific implementation manner of the present invention will be clearly and completely described below.
[0062] The present invention discloses a distal metastasis prediction system based on multiple examinations of gastric cancer patients, which comprises: a data acquisition module: used for collecting preoperative data of gastric cancer patients;
[0063] Data preprocessing module: used to clean and normalize the collected gastric cancer patient data;
[0064] Feature selection module: Using three machine learning algorithms, LassoCV, recursive feature elimination, and Boruta, the preprocessed data is filtered to identify the overlapping feature indicators most relevant to gastric cancer metastasis and construct a feature vector.
[0065] Model training module: Utilizes multiple machine learning algorithms, using the selected feature vectors as input and the patient's actual metastasis status as output. Parameters are automatically grid-searched and validated through hierarchical nested cross-validation. Standardized random initialization is used to establish multiple machine learning models for predicting gastric cancer distant metastasis based on the selected feature vectors. Standardized random initialization can be performed using Xavier standardized random initialization.
[0066] Model screening module: used to screen the best machine learning model among multiple machine learning models for predicting distant metastasis of gastric cancer, and conduct multi-dimensional evaluation and optimization of the machine learning model by drawing relevant curves;
[0067] Decision support module: Develop an online prediction website for the generated machine learning model or connect to the clinical data database on the hospital intranet, automatically generate prediction results based on characteristic variables, and provide clinical decision support.
[0068] The present invention uses a multi-algorithm collaborative feature screening strategy to accurately locate the core indicators most relevant to distal metastasis of gastric cancer, eliminate redundant information interference, and lay a solid foundation for model construction. The model training stage relies on technologies such as automatic grid parameter search and hierarchical nested cross-validation, combined with a standardized initialization process, to effectively improve the generalization ability and stability of the model; the multi-dimensional evaluation system not only verifies the prediction accuracy, but also optimizes around clinical practicality and interpretability to ensure that the model output has both accuracy and medical interpretability, providing support for the establishment of clinical trust. At the clinical translation level, the present invention integrates the optimized model into an online platform or hospital intranet system to achieve automatic analysis of patient data and real-time output of prediction results. This model breaks through the efficiency bottleneck of traditional manual analysis, helps clinicians quickly obtain quantitative decision-making basis, standardizes the diagnostic process of distal metastasis, improves the efficiency of diagnosis and treatment, and reduces the risk of missed diagnosis and misdiagnosis caused by individual experience differences through standardized predictions, providing a scientific reference for the formulation of personalized diagnosis and treatment plans for gastric cancer patients, and promoting the transformation of clinical decision-making from experience-driven to data-driven.
[0069] A specific step of feature screening is:
[0070] Use the LassoCV algorithm to filter out the feature set A with non-zero coefficients;
[0071] Use the RFECV algorithm to construct the feature importance ranking by recursive elimination, and take the top k feature set B;
[0072] Use the Boruta algorithm to compare the importance of the original features and shadow variables, and retain the significant feature set C;
[0073] The overlapping feature index is the intersection of sets A, B, and C, that is, the features that are recognized as important by the three algorithms at the same time.
[0074] The feature selection module uses three machine learning algorithms: LassoCV, recursive feature elimination, and Boruta. These algorithms filter the overlapping features most relevant to gastric cancer metastasis from the preprocessed data and construct feature vectors. The three machine learning algorithms are listed in Table 1.
[0075]
[0076] Table 1
[0077] As an implementation method, the gastric cancer patient data includes imaging data, endoscopy data, and blood test report data;
[0078] Imaging data included: tumor location, tumor size, invasion depth, and lymph node metastasis;
[0079] Endoscopic data included: tumor location, tumor size, pathological type, Lauren classification, degree of differentiation, immunohistochemistry, depth of invasion, and lymph node metastasis;
[0080] Blood test report data includes: tumor markers, routine blood indicators and nutritional indicators.
[0081] The present invention breaks through the information limitations of a single inspection dimension by integrating multimodal data from imaging, endoscopy, and blood tests. Imaging accurately presents the anatomical characteristics of the tumor, endoscopy deeply reveals the pathology and molecular phenotype, and blood indicators reflect the physiological state of the body. The three work together to construct a multidimensional feature map of the tumor, providing richer biological information support for subsequent feature screening and model training, enabling the model to capture the complex associations related to tumor metastasis and improve the comprehensiveness and accuracy of the prediction. From the perspective of clinical implementation, the collected data are all routine preoperative examination items for gastric cancer patients, without the need to increase diagnosis and treatment costs and operational burdens. The system directly connects to the existing inspection process, automatically extracts key information from imaging reports, endoscopic pathology, and blood tests, and fits the actual clinical work scenario. This design not only makes full use of existing medical resources, but also accelerates the conversion process from data collection to prediction output through automated data integration and analysis, helping clinicians to efficiently obtain comprehensive preoperative evaluation basis, and promote the smooth implementation and normalization of the prediction system in clinical scenarios.
[0082] As an embodiment, the blood test report data includes five optimization indexes:
[0083] Tumor marker index TMI, nutritional index PNI, platelet-lymphocyte ratio PLR, lymphocyte-monocyte ratio LMR and neutrophil-lymphocyte ratio NLR.
[0084] See Table 2 below for details.
[0085]
[0086] Table 2
[0087] The five blood optimization indices screened by the present invention break through the observation limitations of a single indicator. Through the composite calculation of tumor markers, nutritional status and immune-inflammatory parameters, they systematically integrate the correlation information between tumor burden, body nutritional reserves and immune response. This multi-dimensional indicator fusion not only compresses data redundancy, but also strengthens the biological signal correlation related to tumor metastasis, provides more targeted hematological characteristics for model construction, and helps the model accurately capture the potential laws of distal metastasis of gastric cancer. From the perspective of clinical implementation, optimization indices such as TMI and PNI are core indicators verified in clinical prognosis studies and have a mature medical cognitive basis.
[0088] Abnormal elevations of the tumor marker TMI are directly correlated with tumor burden and metastatic potential, and combined multi-marker testing offers higher diagnostic efficacy than single markers alone. Preoperative nutritional index (PNI), albumin, and lymphocyte counts reflect nutritional reserves and immune function, respectively. Their combined analysis can assess a patient's resistance to cancer. Elevated platelet-lymphocyte ratios (PLR) and neutrophil-lymphocyte ratios (NLR) indicate activated inflammatory responses, while decreased lymphocytes reflect immunosuppression. Their ratios quantify the inflammatory-immune imbalance within the tumor microenvironment. The lymphocyte-monocyte ratio (LMR) indicates that monocytes participate in tumor-associated macrophage differentiation. A decreased LMR indicates weakened immune surveillance and is associated with increased tumor aggressiveness. The system's use of these indicators aligns the analysis logic of blood data with clinical thinking, facilitating physicians' understanding of the intrinsic relationship between these indicators and metastatic risk. Furthermore, these composite indicators simplify the interpretation of raw data, reduce the cost of analyzing multidimensional blood parameters, and facilitate the seamless integration of the prediction system with the diagnostic and treatment process, enhancing the convenience and acceptability of clinical applications.
[0089] As an implementation method, the model screening module screens the best machine learning model among multiple machine learning models for predicting distant metastasis of gastric cancer according to the following evaluation criteria: the area under the receiver operating characteristic curve is greater than 0.9, accuracy>0.8, sensitivity>0.8, specificity>0.8, F1>0.6, and the area under the precision-recall curve>0.7. Based on the evaluation criteria, the model that meets the criteria and has the largest area under the curve is selected as the best model.
[0090] The model screening module of this invention utilizes a scientific evaluation system based on multi-dimensional quantitative metrics. Its advantage lies in its deep integration of technical performance with clinical needs. An area under the receiver operating characteristic (AUC) score >0.9 ensures the model's core ability to distinguish between metastatic and non-metastatic cases. Accuracy >0.8 ensures overall predictive reliability, mitigating fundamental performance deficiencies. Sensitivity >0.8 and specificity >0.8 address the clinically significant risks of missed diagnosis and misdiagnosis, respectively. High sensitivity reduces the misclassification of metastatic patients as non-metastatic, while high specificity reduces the misclassification of non-metastatic patients as metastatic. F1 scores >0.6 and area under the precision-recall curve >0.7 are specifically tailored to the sample imbalance of gastric cancer metastasis, where metastatic cases are relatively rare. By balancing precision and recall, the model ensures its ability to identify rare metastatic cases and avoids prediction bias caused by uneven data distribution. This multi-metric collaborative screening logic enables the final model to accurately distinguish between the two categories at a macro level while also addressing risk management and rare case identification at a detailed clinical level, achieving a dual optimization of technical performance and medical value.
[0091] First, using standard machine learning tools such as sklearn on an independent test set, pre-set metrics were automatically calculated for multiple models, including logistic regression, random forest, and XGBoost. Models with insufficient baseline performance were first filtered through thresholds. Then, from the eligible candidate models, the model with the highest Area Under the Analyzer (AUC) was selected as the optimal solution. AUC is a core metric that comprehensively measures a model's classification ability at different thresholds. A higher AUC value indicates a stronger model's ability to rank positive and negative samples, thus better meeting the clinical need to prioritize high-risk patients. The thresholds for each metric are determined based on the risk-balancing logic of clinical decision-making. For example, a sensitivity > 0.8 means the model can detect at least 80% of actual metastatic cases, keeping the risk of missed diagnoses within an acceptable range; a specificity > 0.8 ensures that no more than 20% of non-metastatic cases are misclassified, avoiding excessive intervention. The advantage of PR-AUC for imbalanced data lies in its greater focus on predicting the accuracy of positive samples. When clinical practice prioritizes identifying any potential metastases, this metric can effectively avoid the bias of traditional AUC in extreme data distributions. The design of one-to-one correspondence between machine learning indicators and clinical risk control goals makes the screening process not only a comparison of technical parameters, but also a concrete realization of patient-centered diagnosis and treatment logic at the algorithm level.
[0092] As an embodiment, drawing the correlation curve includes: drawing a calibration curve to evaluate the consistency between the model predicted probability and the actual probability;
[0093] Draw a net benefit decision curve to evaluate the clinical practicality of the model at different decision thresholds and help select the optimal decision threshold;
[0094] Draw a learning curve to show how model performance changes with the amount of training data or training time, which is used to diagnose whether the model is underfitting or overfitting;
[0095] Draw the KS curve to evaluate the ability of the classification model to distinguish between positive and negative samples;
[0096] A SHAP diagram was drawn to quantify the contribution of each feature to the model prediction and improve the clinical practicality of the diagnostic model.
[0097] The calibration curve addresses the core question of whether probabilities are trustworthy in clinical decision-making by comparing the consistency of the model's predicted probabilities with the actual probabilities of transition. For example, if the model outputs a 70% probability of transition for a particular patient, the calibration curve verifies whether the actual number of patients who transition within that probability interval is close to 70%. If the curve lies close to the diagonal, the model's probability output has clinical reference value, enabling doctors to formulate precise treatment plans. Conversely, if the predicted probability deviates significantly from reality, it can lead to misjudgments in treatment strategies. The introduction of this curve enables the model to advance from simply distinguishing good from bad to outputting reliable quantitative risk values, providing a solid probabilistic basis for clinical decision-making.
[0098] Traditional model evaluation focuses on differentiation ability, but the DCA test set decision curve incorporates the risk-benefit of clinical decision-making into the evaluation system for the first time. By simulating the net benefits under different decision thresholds, DCA can intuitively present the practical value of the model in real diagnosis and treatment scenarios. For example, when the threshold is set to 50%, the model may increase the missed diagnosis rate and decrease the net benefit due to excessive pursuit of specificity; when the threshold is set to 20%, although the sensitivity is improved, the increase in misdiagnosis may lead to overtreatment. By finding the threshold corresponding to the peak of net benefit, DCA helps doctors find the optimal balance between the risk of missed diagnosis and the cost of misdiagnosis, so that the model output can be directly converted into a feasible diagnosis and treatment strategy.
[0099] The learning curve helps identify whether the model is underfitting or overfitting by showing how model performance changes with the amount of training data. If the AUC of the training set is much higher than that of the validation set and there is no convergence as the amount of data increases, it indicates overfitting and requires optimization through regularization or data augmentation. If both AUCs are low and increase significantly with the amount of data, it indicates underfitting and requires adjustment of feature engineering or model complexity. In clinical scenarios where gastric cancer data collection is limited, the learning curve can clearly determine how much data can meet the model performance requirements, ensure that the model operates stably under the distribution of real clinical data, and avoid prediction failures due to insufficient data or overfitting.
[0100] The KS curve measures the model's extreme performance at the optimal discrimination threshold by calculating the maximum difference between the cumulative distributions of positive and negative samples. In imbalanced data with a low proportion of positive gastric cancer metastasis samples, higher KS values indicate that the model is able to prioritize patients with metastases in the high-probability zone and patients without metastases in the low-probability zone, thereby achieving accurate identification among high-risk populations of clinical focus. For example, a KS value of 0.7 indicates that at a certain threshold, 85% of patients with metastases are correctly identified, while only 15% of patients without metastases are misclassified. This provides quantitative support for clinical strategies that prioritize extremely high-risk patients and avoid model evaluation bias caused by data imbalance.
[0101] SHAP plots quantify the contribution of each feature to the prediction, transforming black-box models into interpretable clinical knowledge. For example, for a high-risk patient, a SHAP plot can reveal that advanced T stage and a high TMI index are the primary drivers, while a high degree of differentiation partially offsets the risk. This helps physicians quickly verify whether model decisions align with pathological logic or identify potential abnormal associations. This interpretability not only enhances physicians' trust in the model but also aids in the discovery of new clinical patterns, enabling the model to evolve from a predictive tool to an intelligent assistant for clinical decision-making, integrating it into the diagnosis and treatment process.
[0102] As an implementation method, a variety of machine learning algorithms include: Logistic classification, XGBoost classification, LightGBM classification, random forest classification, AdaBoost classification, decision tree classification, GBDT classification, Gaussian naive Bayes classification, neural network classification, support vector machine classification, and k-nearest neighbor classification.
[0103] The present invention adopts a multivariate machine learning algorithm system, which comprehensively captures multi-dimensional characteristic laws such as linear association, nonlinear interaction, and complex patterns in the data by covering algorithms of different paradigms such as linear, tree models, ensemble learning, and neural networks. The differences in the learning mechanisms of different algorithms complement each other, avoiding the inductive bias and performance bottlenecks of a single algorithm, and ensuring that the model can stably mine potential associations in complex clinical scenarios for gastric cancer metastasis prediction. Subsequently, the optimal model is screened through multi-dimensional indicators, which not only gives play to the collective advantages of multi-dimensional algorithms, but also avoids the risk of algorithm dependence, so that the system has both deep analysis capabilities for data and cross-scenario generalization performance, providing technical redundancy and performance guarantees for accurately capturing the complex biological characteristics of gastric cancer distal metastasis, and ultimately forming a prediction model with strong reliability and wide adaptability.
[0104] The present invention discloses a method for predicting distal metastasis based on multiple examinations of gastric cancer patients, using any one of the aforementioned distal metastasis prediction systems based on multiple examinations of gastric cancer patients, comprising:
[0105] S1: Collect preoperative imaging data, endoscopic data, and blood test report data of gastric cancer patients through the data acquisition module;
[0106] S2: Use the data preprocessing module to clean and normalize the collected data and remove outliers and missing values;
[0107] S3: The feature selection module uses three machine learning algorithms, LassoCV, recursive feature elimination, and Boruta, to screen the overlapping feature indicators most relevant to gastric cancer distant metastasis from the preprocessed data and construct a feature vector.
[0108] S4: Using the model training module, the screened feature vectors are used as input and the patient's actual metastasis status is used as output. Through automatic grid parameter search and hierarchical nested cross-validation, a prediction model is established based on multiple machine learning algorithms such as logistic classification, XGBoost classification, and random forest classification.
[0109] S5: Through the model screening module, the best model is selected according to the preset evaluation criteria, and the model is evaluated and optimized in multiple dimensions by drawing calibration curves, decision curves, learning curves, KS curves and SHAP graphs;
[0110] S6: Develop the optimal model into an online prediction website or connect it to the hospital intranet using the decision support module to automatically generate metastasis risk prediction results based on the input feature variables.
[0111] This invention builds a standardized prediction system for the entire process through the fusion of multimodal data and the collaboration of multiple algorithms. It integrates multidimensional data from imaging, endoscopy, and blood tests, breaking through the information silos of a single examination and providing the model with more comprehensive pathological, anatomical, and physiological characteristics. It relies on LassoCV, recursive feature elimination, and cross-validation of the Boruta algorithm to accurately screen core features, eliminate redundant information, and ensure a strong correlation between input variables and metastasis risk. The model training phase uses technologies such as automatic grid parameter search and hierarchical nested cross-validation, combined with parallel modeling of multiple algorithms such as Logistic, XGBoost, and Random Forest. This not only leverages the advantages of different algorithms, but also achieves dual optimization of model performance and clinical practicality through multidimensional curve evaluation, ensuring the accuracy, stability, and reliability of the prediction results from a technical perspective.
[0112] From a clinical perspective, this method transforms complex machine learning technology into a practical diagnostic and treatment tool. By integrating the decision support module with existing hospital systems, it automates the entire process from data collection to risk prediction, significantly shortening the preoperative evaluation cycle and reducing manual analysis costs. The standardized prediction process avoids diagnostic bias caused by individual experience differences, providing a unified, quantitative basis for determining metastasis risk. This approach offers significant advantages in identifying high-risk patients early and minimizing the risk of missed or misdiagnosis. Furthermore, the multi-dimensional evaluation curves provide the model with transparent decision logic, enabling physicians to quickly understand and verify predictions.
[0113] As an implementation method, after S5 selects the best model through the model screening module, it also performs internal and external test set verification;
[0114] Internal test set validation includes: inputting a test set from the same center, independent of the training and validation sets, into the best model, and verifying the generalization ability under the same data distribution by calculating indicators such as AUC and accuracy;
[0115] External test set validation involves inputting the model into a test set from another center and comparing the AUC values of data from different centers to evaluate the model's practical application in an unknown data distribution. Center data refers to clinical data collected independently by different medical institutions, such as hospitals, medical centers, or research institutions.
[0116] The value of internal test set verification: Through performance evaluation of an independent test set in the same center, it is possible to accurately identify whether the model is overfitting or data-dependent. If indicators such as AUC and accuracy of the internal test set are consistent with those of the training or validation set, it means that the model has truly captured the inherent laws of data distribution rather than memorizing the noise of the training data. This provides direct evidence for the stable application of the model in the same hospital, with the same equipment, and with the same patient population, and addresses clinical concerns about whether the model is reliable for use in this hospital.
[0117] The core significance of external validation: Introducing heterogeneous, multi-center data for validation is key to preventing single-center bias and assessing the model's real-world generalization ability. Comparing the AUC values of data from different centers essentially quantifies the model's adaptability to differences in data distribution. If the external-center AUC remains high, such as ≥0.8, and differs from the internal-center AUC within a reasonable range, such as <0.1, this indicates that the model is universally applicable to the core characteristics of the heterogeneous data, rather than relying on the unique data patterns of a single center. If the confidence interval of the external validation is narrow and does not include 0.5, it indicates that the improvement in model performance is statistically significant rather than due to chance fluctuations, reinforcing the conclusion that the model can effectively predict across different clinical settings. In clinical practice, gastric cancer patient data exhibit significant multi-center heterogeneity, and internal validation alone cannot demonstrate the cross-scenario applicability of the model. Through the dual verification of internal and external validation, the reliability of the model in homogeneous data is ensured, and the comparison of multi-center AUC and confidence intervals demonstrates its robustness to heterogeneous data.
[0118] As an implementation method, the data collected in S1 includes:
[0119] Imaging data: tumor location, tumor size, invasion depth, and lymph node metastasis;
[0120] Endoscopic data: tumor location, tumor size, pathological type, Lauren classification, degree of differentiation, immunohistochemistry, invasion depth, and lymph node metastasis;
[0121] Blood test report data: tumor markers, routine blood indicators and nutritional indicators, and the blood test report data is optimized into five indices: tumor marker index, nutritional index, platelet-lymphocyte ratio, lymphocyte-monocyte ratio and neutrophil-lymphocyte ratio.
[0122] As an implementation method, the specific steps of screening overlapping feature indicators in S3 are: performing feature screening using the three algorithms LassoCV, RFECV, and Boruta respectively, and taking the intersection features of the results of the three algorithms;
[0123] S4 includes various machine learning algorithms, including Logistic classification, XGBoost classification, LightGBM classification, random forest classification, AdaBoost classification, and support vector machine classification. Parameter optimization uses automatic grid parameter search, and the verification method is stratified nested cross-validation.
[0124] Example:
[0125] Let’s take the data of 800 gastric cancer patients in the hospital as an example:
[0126] 1. Data collection: Imaging, endoscopy, and blood test report data of these 800 patients were collected, and the blood test indicators were optimized into TMI, PNI, PLR, LMR, and NLR.
[0127] 2. Data preprocessing: The data were cleaned, 31 samples with a large number of missing values or outliers were removed, and the data of the remaining 769 samples were normalized.
[0128] 3. Feature selection: Four overlapping feature variables related to distant metastasis of gastric cancer were screened out using three feature screening methods: LassoCV, RFECV, and Boruta, namely, degree of differentiation, clinical T stage, clinical N stage, and TMI.
[0129] The results of visualizing the LassoCV algorithm for screening each characteristic coefficient are shown in the attached figure. Figure 1 The lasso model coefficients in Figure 1 As shown;
[0130] The recursive feature elimination cross validation curve is obtained by RFECV algorithm as follows Figure 2 As shown;
[0131] The Boruta feature importance evaluation diagram obtained by Boruta is as follows Figure 3 As shown in the figure, the gray columns represent the importance distribution of the shadow variables, and the black columns represent the importance ranking of the original features. When the importance of the original feature exceeds the maximum value of the shadow variable, it is determined to be a feature related to gastric cancer metastasis.
[0132] The Venn diagram was used to screen out the overlapping feature indicators most relevant to gastric cancer metastasis from the preprocessed data. There were 4 overlapping variables in the Venn diagram. Figure 4 shown.
[0133] 4. Model training: Six machine learning algorithms (logistic classification, XGBoost classification, random forest classification, AdaBoost classification, Gaussian naive Bayes classification, and support vector machine classification) were used. The selected feature vectors were used as input, and the patient's actual metastasis status was used as output. Automatic grid parameter search was used for parameters, and validation was performed using 5-fold nested cross-validation with a random seed of 42. Six prediction models for distant metastasis in gastric cancer patients were established.
[0134] The performance evaluation visualization results of the machine learning model are as follows Figure 5 、 Figure 6 、 Figure 7 and Figure 8 As shown in;
[0135] The model training and validation module trains a variety of machine learning models and evaluates their performance on the training set. The training set ROC curve (receiver operating curve) is as follows: Figure 5 As shown in;
[0136] The ROC curve of the validation set is as follows Figure 6 As shown in;
[0137] The training set precision-recall curve is the training set PR curve. Figure 7 As shown in;
[0138] The validation set precision-recall curve is the validation set PR curve. Figure 8 As shown in .
[0139] The receiver operating curve is Figure 5 and Figure 6 :
[0140] Figure 5 This is the ROC curve of the training set, which belongs to the receiver operating curve, training set version;
[0141] Figure 6 This is the validation set ROC curve, which belongs to the receiver operating characteristic curve, validation set version;
[0142] Figure 5 The ROC curve of the training set is a visualization tool for evaluating the performance of the binary classification model, including whether gastric cancer has metastasized.
[0143] For each model, traverse the probability threshold and calculate:
[0144] Vertical axis: sensitivity, true positive rate = number of samples with true metastasis and correct prediction / total number of true metastasis samples, TP / (TP+FN);
[0145] Horizontal axis: 1-specificity, false positive rate = number of false metastasis and wrongly predicted samples / all true non-metastasis samples, FP / (FP+TN).
[0146] Visually demonstrate the performance differences between models like XGBoost and Logistic on the training set, supporting subsequent optimal model selection. Compare this with the ROC curve on the subsequent test set to determine whether the model is overfitting. If the test set AUC drops significantly, feature or model optimization is necessary. A high AUC, such as >0.9, demonstrates high prediction accuracy.
[0147] Figure 6 The validation set ROC curve is a core tool for evaluating the performance of a binary classification model.
[0148] The horizontal axis is 1-specificity, false positive rate, and the vertical axis is sensitivity, true positive rate. The closer the curve is to the upper left corner, the stronger the model's ability to distinguish.
[0149] The validation dataset is independent of the training set and is used to verify the model's generalization ability and avoid overfitting. This graph is used to test the model's performance in predicting gastric cancer distant metastasis using new data that was not used in training.
[0150] The figure overlays the ROC curves of models like XGBoost, Logistic Classification, and Random Forest. Model performance is compared using the AUC (area under the curve). When compared with the training set ROC curve, a small difference in AUC indicates that the model is not overfitting.
[0151] Figure 7 The training set precision-recall curve is the training set PR curve.
[0152] Horizontal axis, Recall, recall rate: equivalent to the previous sensitivity, the formula is TP / (TP+FN), the proportion of true transfer samples that are correctly predicted;
[0153] Vertical axis, Precision, precision: the formula is TP / (TP+FP), the proportion of samples predicted to be transferred that are actually transferred.
[0154] Training indicates that the curve is generated based on the training set data and is used to evaluate the model's ability to predict gastric cancer distant metastasis, positive examples, and minority classes during the training phase.
[0155] Gastric cancer metastasis is a minority event with few positive examples and many negative examples. The ROC curve will mask the problem of positive example prediction due to the large number of negative examples, while the PR curve focuses more on the prediction accuracy of positive examples. If the accuracy is low, it means that the model will misjudge a large number of non-metastatic patients as metastatic, increasing clinical anxiety.
[0156] The graph uses AP, AveragePrecision, and Area Under the Curve (AUC) to quantify model performance. For example, RandomForest achieves AP = 0.846, while XGBoost achieves AP = 0.834, visually demonstrating the ability of different models to identify positive examples. When compared with the subsequent PR curve for the test set, a significant difference in AP between the two indicates possible overfitting, with good performance on the training set and poor performance on the test set.
[0157] Figure 8 The validation set precision-recall curve is the validation set PR curve.
[0158] The horizontal axis is the recall rate (Recall = True Positive Rate), which measures the proportion of patients with actual metastases who are correctly predicted, and the vertical axis is the precision rate (Precision), which measures the proportion of patients predicted to have metastases who actually have metastases.
[0159] Because gastric cancer metastasis is a minority event with few positive examples and many negative examples, the PR curve is more sensitive to the quality of positive example prediction than the ROC curve, avoiding the problem of a large number of negative examples masking the positive example prediction problem.
[0160] Validation set: The curve is generated based on validation data independent of the training set and is used to test the generalization ability of the model, that is, whether the accuracy of the model in predicting gastric cancer metastasis on unseen data is stable.
[0161] The AP value, such as the logistic model AP = 0.719 in the figure, quantifies the model's prediction reliability for metastatic cases. The higher the AP, the more accurate the positive case prediction.
[0162] If the AP of the validation set PR curve is slightly different from that of the training set, it indicates that the model is not overfitting. In clinical applications, the proportion of patients with predicted metastasis who actually have metastasis is high, which reduces the anxiety of misdiagnosis.
[0163] 5. Best model screening:
[0164] 1. The best model was the logistic regularized L2 classification model with the largest area under the receiver operating characteristic (ROC) curve on the validation set. The model parameters were: C (regularization factor): 0.1; max-iter (number of iterations): 100; penalty (regularization type): L2; and tol (convergence measure): 1e-06. The validation set results showed that this model performed well in predicting distant metastasis of gastric cancer, achieving an AUC of 0.906, an accuracy of 81.8%, a sensitivity of 84.8%, a specificity of 81.1%, an F1 score of 0.637, and an AP value of 0.719.
[0165] Second, the logistic classification model was further evaluated. Its calibration curve demonstrated good consistency between the model's predicted and actual probabilities. The DCA curve was used to assess the model's clinical utility at different decision thresholds, assisting in selecting the optimal decision threshold. The learning curve demonstrated that the diagnostic model was neither underfit nor overfit. The KS curve demonstrated the classification model's strong ability to distinguish between positive and negative samples. Finally, a SHAP plot was generated to quantify the contribution of each feature to the model's predictions, enhancing the model's interpretability and clinical utility.
[0166] The learning curve of the logistic regression model shows the trend of model performance changing with the number of training samples. Figure 9 As shown, we focus on the ROC-AUC indicator.
[0167] Learning curves are used to analyze the impact of data size on the model, such as whether it is underfitting or overfitting, and how much data is needed to saturate the performance. The vertical axis, Roc-Auct, is the core performance metric of the binary classification model. The area under the curve (AUC) measures the ability to distinguish between metastases and non-metastases. The labeled gastric cancer patient data is divided into training sets with gradually increasing sample sizes, such as 100, 200...500 samples, and an independent validation set is retained. For each training set size, a logistic regression model is trained to learn features, such as the four overlapping variables: degree of differentiation, clinical T / N stage, and TMI, and their relationship with metastasis. For the same-sized training set, the training set AUC (red line) and the fixed validation set AUC (blue line) are calculated, and the Roc-Auc-score is called to calculate the performance. The training / validation AUCs for different sample sizes are connected into curves to compare the trends of the two.
[0168] If the training set AUC and validation set AUC are close, as shown in the figure, and the difference is small, the model is not overfitting and its generalization ability is stable. If the training set AUC is much larger than the validation set AUC, it indicates overfitting and requires regularization or additional data. Observe whether the curves converge. If the validation set AUC stops improving after the sample size reaches 300 and enters a plateau, 500 samples are sufficient to train the model and no additional data collection is needed.
[0169] Comparing multiple model learning curves, such as the XGBoost learning curve, demonstrates the robustness of the baseline model if the validation set AUC for logistic regression remains stable, helping to illustrate the performance advantages of subsequent optimized models. A learning curve is a graph showing the ROC-AUC variation of a logistic regression model under varying amounts of training data. Its core value lies in quantifying the impact of data volume on the model, validating model generalization capabilities, and guiding data strategies and model optimization, directly contributing to the goal of efficient and stable gastric cancer metastasis prediction technology.
[0170] Figure 10 It is the calibration curve of the logistic regression model and the calibration curve.
[0171] Horizontal axis: binned average of the model-predicted probability of gastric cancer metastasis. For example, the predicted probability is divided into intervals of 0-0.1, 0.1-0.2…1.0, and the average predicted probability of each interval is calculated.
[0172] Vertical axis: the proportion of samples that actually transfer within each probability interval.
[0173] Perfect calibration line, black dashed line: If the model predictions are completely reliable, the curve should coincide with the y=x diagonal line, and the predicted probability = the actual probability.
[0174] Use the trained logistic regression model to output the transition probability for the validation set samples.
[0175] Probability binning: Divide the predicted probability into several equal-width intervals, such as 10 intervals: 0-0.1, 0.1-0.2, ..., 0.9-1.0.
[0176] For each interval, the number of samples actually transferred is divided by the total number of samples in the interval to obtain the Fraction of positives on the vertical axis.
[0177] The average predicted probability and the actual positive rate of each interval are connected into a curve (blue line), which is compared with the perfect calibration line (dashed line). The calibration error is also calculated to quantify the deviation between the prediction and the actual.
[0178] The calibration curve can verify the credibility of the probability prediction. If the model predicts an 80% probability of metastasis, but only 50% of patients actually metastasize, the curve will be below the dotted line, which will lead to overtreatment; if the actual probability is 90%, the curve will be above the dotted line, which will result in missed diagnosis. The logistic regression curve in the figure is close to the dotted line trend, with a deviation of 0.100, which is within an acceptable range, indicating that the probability predicted by the model is clinically applicable. If the curve deviates seriously from the dotted line, such as the predicted probability is high but the actual probability is low, the model needs to be corrected through calibration methods such as PlattScaling and IsotonicRegression to make the predicted probability more credible. The calibration curve is a core tool for evaluating whether the predicted probability of the logistic regression model is reliable. It is directly related to whether doctors can make clinical decisions based on the model probability and is a key verification link for the clinical practicality of the prediction system.
[0179] Figure 11 It is the test set decision curve, that is, the DCA curve.
[0180] Horizontal axis: ThresholdProbability, decision threshold. If the predicted transfer probability is ≥ 30%, it is determined to be a transfer.
[0181] Vertical axis: MeanNetBenefit, average net benefit, measures the clinical value of model-guided decision-making, the benefit of correct treatment minus the risk of incorrect treatment.
[0182] Logistic, red solid line: net benefit of the logistic regression model at different thresholds;
[0183] TreatNone, red dashed line: extreme strategy, never treat, the net benefit is always 0, because there is no intervention and no wrong treatment;
[0184] TreatAll, black dashed line: extreme strategy, always treat everyone, the net benefit decreases with the threshold because the lower the threshold, the more misdiagnoses and treatment of non-transferred patients, the higher the risk.
[0185] Use the trained logistic regression model to output the transition probability for the test set samples.
[0186] Set the decision threshold range, and for each threshold:
[0187] Determined to be transferred but not transferred;
[0188] Calculate true positive (TP) and false positive (FP);
[0189] Substitute into the net benefit formula: NetBenefit = total number of samples TP − FP × treatment benefit intervention risk;
[0190] Connect the net returns at each threshold into a curve to compare the performance of the model with extreme strategies.
[0191] If the model's logistic curve consistently exceeds both TreatNone and TreatAll, it indicates that model-guided decision-making is generating positive net benefits. As shown in the figure, when the threshold is >10%, the logistic curve is higher than TreatAll and significantly higher than TreatNone, demonstrating that using model predictions to guide treatment is more beneficial than blindly treating or not treating. Find the threshold range with the highest model net benefit. As shown in the figure, the net benefit peaks significantly between 20% and 40%, providing clinical guidance. If clinicians are more focused on reducing missed diagnoses, a lower threshold may be appropriate; if they are more focused on reducing misdiagnoses, a higher threshold may be appropriate.
[0192] From predictive accuracy to clinical applicability, ROC and AUC demonstrate the model's strong discriminatory capabilities, but DCA demonstrates its ability to guide physician decision-making, corresponding to the clinical practicality of the gastric cancer distant metastasis prediction system. The test set decision curve, or DCA curve, is a decision curve analysis on the test set. It quantifies the clinical value of the model at different decision thresholds, addressing the key issue of how model predictions can be translated into treatment decisions. It is a core validation tool for transitioning from laboratory accuracy to clinical practicality.
[0193] Figure 12 It is the KS statistic graph of the test set.
[0194] KS value: the maximum vertical distance between the two curves. In the figure, 0.747at0.163 means that when the threshold is 16.3%, the cumulative distribution difference between positive and negative samples reaches a maximum value of 0.747.
[0195] Class0, blue line: cumulative distribution of negative cases, such as patients without metastasis, the proportion of samples below the threshold, the lower the threshold, the higher the proportion;
[0196] Class1, green line: cumulative distribution of positive examples, such as metastatic patients;
[0197] The greater the gap between the two curves, the stronger the model's ability to distinguish between positive and negative samples, with positive samples more concentrated in the high-probability region and negative samples more concentrated in the low-probability region. The KS statistic plot on the test set more intuitively demonstrates the threshold at which the model most effectively distinguishes between positive and negative samples, compared to the ROC-AUC. This threshold can be used as a reference for DCA curve analysis, and the final decision threshold can be selected in conjunction with net benefit to implement discriminatory ability into treatment decisions. The KS statistic plot on the test set evaluates the model's ability to distinguish between gastric cancer metastases and non-metastases on the test set. A high KS value directly demonstrates the model's strong classification ability and is a key quantitative indicator of predictive accuracy.
[0198] Figure 13 It is the SHAP summary graph.
[0199] SHAP, which stands for SHapley Additive exPlanations, is a model interpretability method that quantifies the contribution of each feature to the model's prediction results through the Shapley value. Positive contributions increase the transition probability, while negative contributions reduce the probability.
[0200] Vertical axis: features of the model input (such as Nstage (clinical N stage), TMI (tumor metabolic index), Tstage (clinical T stage), Differentiationgrade (degree of differentiation)).
[0201] Horizontal axis: SHAPvalue, the impact of the feature on the model output. A positive value indicates that the feature increases the predicted probability of gastric cancer metastasis, and a negative value indicates a reduced probability.
[0202] Color: Red represents high eigenvalues, and blue represents low eigenvalues.
[0203] Based on the selected features, a prediction model is trained. SHAP values are calculated using the SHAP library for each feature in the test set. The SHAP values of all samples are grouped by feature and a scatter plot is plotted, with each dot representing the SHAP value of that feature for a sample. Colors are used to map high and low feature values, demonstrating how feature values influence contribution.
[0204] A high clinical N stage corresponds to a positive SHAP value—the higher the N stage, the more likely the model is to predict metastasis. A high Differentiation Grade (DG) value corresponds to a negative SHAP value—the higher the differentiation grade, the lower the malignancy, and the more likely the model is to predict no metastasis, which aligns with pathological logic. This addresses the unexplainable pain point of machine learning models and supports clinical practicality.
[0205] Four features were selected through Boruta and univariate analysis. The SHAP plot visually demonstrates that the SHAP values for these four features are widely distributed, demonstrating the effectiveness of feature selection. The relationship between the features and predictions aligns with medical logic, confirming the physical validity of the model. The SHAP summary plot is a core tool for model interpretability. By quantifying feature contributions and associating them with medical logic, it transforms black-box models into clinically understandable decision-making.
[0206] Figure 14 and Figure 15 is a SHAP force-directed graph.
[0207] The SHAP force-directed graph is a model interpretability tool. Its core is to decompose how the predicted value of a single sample is pushed away from the global mean by each feature:
[0208] basevalue: the average predicted value of the model for all patients;
[0209] f(x): predicted value of the current patient;
[0210] Colors and stripes:
[0211] Red band: Feature values push up the predicted value;
[0212] Blue strips: Feature values pull down the predicted value;
[0213] The length of the bar represents the contribution size.
[0214] Select a specific patient from the test set. Using the trained model, calculate the contribution of each feature for that patient using the SHAP library. Stretching the feature contributions in a strip from basevalue to f(x) in the direction of push or pull, visually demonstrates which features contribute to a patient's risk being higher or lower than average.
[0215] Doctors can directly interpret:
[0216] Case 1, first two figures, red dominates: TMI is high, 6.42-16.22, clinical T stage = 4.0 advanced stage, clinical N stage = 2.0 lymph node metastasis - the red strip pushes up the predicted value - the patient's metastasis risk is significantly higher than the average, such as 22%-25%, and intensive intervention is required.
[0217] Case 2, third figure, red and blue balance: TMI = 1.088 low metabolism, Differentiationgrade = 1.0 high differentiation, low malignancy - red band is limited, blue band is pulled down - patient risk is close to the mean 7%, and prognosis is good.
[0218] High TMI, late stage - red - pushes up the risk. Medically, active metabolism, late stage = high risk of metastasis.
[0219] High degree of differentiation - blue - lower risk. In medicine, high differentiation = low malignancy = low risk of metastasis.
[0220] For high-risk patients, red is dominant: more intensive follow-up and more aggressive treatment may be recommended.
[0221] For low-risk patients, blue is dominant: conservative observation may be recommended to avoid overtreatment.
[0222] 3. Internal and external validation: Internal and external validation showed that the model has good diagnostic efficacy, strong generalization ability, and practical application value.
[0223] Internal test set:
[0224] Figure 16 It is the ROC curve of the internal test set, which is used to verify the performance of the internal test set in internal and external validation.
[0225] The internal test set is an independent subset of the same study cohort, used to test the model's generalization ability to new data within the study. The internal test set was split proportionally from the gastric cancer patient data from the study. The model trained on the training set outputs transition probabilities for the internal test set, and sensitivity and 1-specificity are calculated at each threshold. The points are connected to form a receiver operating characteristic (ROC) curve, and the area under the curve (AUC) is calculated to quantify the diagnostic efficacy on the internal test set.
[0226] If the model performs well in the training set but has a low ROC / AUC on the internal test set, it indicates overfitting. The AUC in the figure is 0.918, which proves that the model can still make stable predictions on new data within the study and has not overfitted. The internal test set is the first step in clinical implementation verification. It simulates new patients in the same hospital or study. The high AUC shows that the model is reliable in predicting metastasis of patients of the same type and has practical application potential in internal scenarios. It lays the foundation for subsequent external validation. If the internal validation is unstable, it will be even more difficult to generalize externally. The shape of the ROC curve and the high AUC indicate that it can still maintain high sensitivity at a low misdiagnosis rate and has high clinical practical value. The excellent performance of the internal test set is a prerequisite for demonstrating the strong generalization ability of the model. Only when the internal validation is stable can further external validation be carried out to verify the adaptability of the model to heterogeneous patients. If the internal validation fails, the significance of the external validation will be greatly reduced.
[0227] Figure 17 is the receiver operating curve of the external test set.
[0228] The model training data of the external test set is completely independent of heterogeneous data, which is the key to testing the model's ability to generalize across scenarios. Collect gastric cancer patient data from multi-center or heterogeneous scenarios to ensure that there is no data leakage with the training set. Use the trained model to output the transition probability for the external test set; traverse the probability thresholds, calculate the sensitivity and 1-specificity of each threshold; draw the ROC curve, and calculate the AUC. The AUC = 0.840 and narrow confidence interval in the figure indicate that the performance of the model in the external scenario is stable and effective; the hospital has introduced this model, and even if the data comes from different devices / cohorts, it can still make reliable predictions. If the AUC of the internal test set is higher, the external AUC = 0.840 is still considerable, indicating that the model is more accurate for homologous data and can also adapt to heterogeneous data; forming an evidence chain of internal stability and external adaptability.
[0229] For newly admitted gastric cancer patients, their imaging results, gastroscopy results, and test report data were collected and input into the trained model. The predicted probability of distant metastasis was 42.1%, which was higher than the threshold of 16%. Doctors formulated corresponding examination and treatment plans based on the predicted results. The relevant figures are shown in the figure below. Figure 18 shown.
[0230] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all such improvements and changes should fall within the scope of protection of the appended claims of the present invention.
Claims
1. A system for predicting distant metastasis based on multiple examinations of gastric cancer patients, characterized by: include: Data collection module: used to collect preoperative data of gastric cancer patients; Data preprocessing module: used to clean and normalize the collected gastric cancer patient data; Feature selection module: Using three machine learning algorithms, LassoCV, recursive feature elimination, and Boruta, the preprocessed data is filtered to identify the overlapping feature indicators most relevant to gastric cancer metastasis and construct a feature vector. Model training module: This module uses multiple machine learning algorithms, with the selected feature vectors as input and the patient's actual metastasis status as output. Parameters are automatically grid-searched and validated through stratified nested cross-validation. Standardized random initialization is used to establish multiple machine learning models for predicting distant metastasis of gastric cancer based on the selected feature vectors. Model screening module: used to screen the best machine learning model among multiple machine learning models for predicting distant metastasis of gastric cancer, and conduct multi-dimensional evaluation and optimization of the machine learning model by drawing relevant curves; Decision support module: Develop an online prediction website for the generated machine learning model or connect to the clinical data database on the hospital intranet to automatically generate prediction results based on feature variables.
2. The system for predicting distal metastasis based on multiple examinations of gastric cancer patients according to claim 1, characterized in that: Gastric cancer patient data include imaging data, endoscopic data, and blood test report data; Imaging data included: tumor location, tumor size, invasion depth, and lymph node metastasis; Endoscopic data included: tumor location, tumor size, pathological type, Lauren classification, degree of differentiation, immunohistochemistry, depth of invasion, and lymph node metastasis; Blood test report data includes: tumor markers, routine blood indicators and nutritional indicators.
3. The system for predicting distal metastasis based on multiple examinations of gastric cancer patients according to claim 2, characterized in that: The blood test report data includes five optimization indices: Tumor marker index TMI, nutritional index PNI, platelet-lymphocyte ratio PLR, lymphocyte-monocyte ratio LMR and neutrophil-lymphocyte ratio NLR.
4. The system for predicting distal metastasis based on multiple examinations of gastric cancer patients according to claim 1, characterized in that: The model screening module screens the best machine learning model among multiple machine learning models for predicting distant metastasis of gastric cancer. The evaluation criteria are as follows: the area under the receiver operating characteristic curve is greater than 0.9, accuracy>0.8, sensitivity>0.8, specificity>0.8, F1>0.6, and the area under the precision-recall curve>0.
7. Based on the evaluation criteria, the model that meets the criteria and has the largest area under the curve is selected as the best model.
5. The system for predicting distal metastasis based on multiple examinations of gastric cancer patients according to claim 1, characterized in that: Drawing the relevant curve includes: drawing the calibration curve to evaluate the consistency between the model predicted probability and the actual probability; Draw a net benefit decision curve to evaluate the clinical practicality of the model at different decision thresholds and help select the optimal decision threshold; Draw a learning curve to show how model performance changes with the amount of training data or training time, which is used to diagnose whether the model is underfitting or overfitting; Draw the KS curve to evaluate the ability of the classification model to distinguish between positive and negative samples; A SHAP diagram was drawn to quantify the contribution of each feature to the model prediction and improve the clinical practicality of the diagnostic model.
6. The system for predicting distal metastasis based on multiple examinations of gastric cancer patients according to claim 1, characterized in that: Various machine learning algorithms include: Logistic classification, XGBoost classification, LightGBM classification, random forest classification, AdaBoost classification, decision tree classification, GBDT classification, Gaussian naive Bayes classification, neural network classification, support vector machine classification, and k-nearest neighbor classification.
7. A method for predicting distal metastasis based on multiple examinations of gastric cancer patients, using a system for predicting distal metastasis based on multiple examinations of gastric cancer patients as claimed in any one of claims 1 to 6, characterized in that: include: S1: Collect preoperative imaging data, endoscopic data, and blood test report data of gastric cancer patients through the data acquisition module; S2: Use the data preprocessing module to clean and normalize the collected data and remove outliers and missing values; S3: The feature selection module uses three machine learning algorithms, LassoCV, recursive feature elimination, and Boruta, to screen the overlapping feature indicators most relevant to gastric cancer distant metastasis from the preprocessed data and construct a feature vector. S4: Using the model training module, the screened feature vectors are used as input and the patient's actual metastasis status is used as output. Through automatic grid parameter search and hierarchical nested cross-validation, a prediction model is established based on multiple machine learning algorithms such as logistic classification, XGBoost classification, and random forest classification. S5: Through the model screening module, the best model is selected according to the preset evaluation criteria, and the model is evaluated and optimized in multiple dimensions by drawing calibration curves, decision curves, learning curves, KS curves and SHAP graphs; S6: Develop the optimal model into an online prediction website or connect it to the hospital intranet using the decision support module to automatically generate metastasis risk prediction results based on the input feature variables.
8. The method for predicting distant metastasis based on multiple examinations of gastric cancer patients according to claim 7, characterized in that: After S5 selects the best model through the model screening module, it also conducts internal and external test set verification; Internal test set validation includes: inputting a test set from the same center, independent of the training and validation sets, into the best model, and verifying the generalization ability under the same data distribution by calculating indicators such as AUC and accuracy; External test set validation includes: inputting the test set of another center into the model, comparing the AUC values of data from different centers, and evaluating the actual application effect of the model in unknown data distribution.
9. The method for predicting distant metastasis based on multiple examinations of gastric cancer patients according to claim 7, characterized in that: The data collected in S1 include: Imaging data: tumor location, tumor size, invasion depth, and lymph node metastasis; Endoscopic data: tumor location, tumor size, pathological type, Lauren classification, degree of differentiation, immunohistochemistry, invasion depth, and lymph node metastasis; Blood test report data: tumor markers, routine blood indicators and nutritional indicators, and the blood test report data is optimized into five indices: tumor marker index, nutritional index, platelet-lymphocyte ratio, lymphocyte-monocyte ratio and neutrophil-lymphocyte ratio.
10. The method for predicting distant metastasis based on multiple examinations of gastric cancer patients according to claim 8, characterized in that: The specific steps for screening overlapping feature indicators in S3 are: perform feature screening using the three algorithms LassoCV, RFECV, and Boruta respectively, and take the intersection features of the results of the three algorithms; S4 includes various machine learning algorithms, including Logistic classification, XGBoost classification, LightGBM classification, random forest classification, AdaBoost classification, and support vector machine classification. Parameter optimization uses automatic grid parameter search, and the verification method is stratified nested cross-validation.
Citation Information
Patent Citations
Machine learning-based gastric cancer lymph node metastasis risk assessment system and device, and storage medium
CN114420291A
Fault prediction method and device for power distribution network in wildfire spreading environment and medium
CN117332275A
Construction method of early gastric cancer risk prediction model and electronic equipment
CN117457195A
Characteristic fusion and data enhancement-based energetic material bond dissociation energy prediction method
CN117497095A
Inplanatable ML method and system for predicting death risk of ICU stroke patient in hospital
CN119230130A
Cited By
Method, system and device for constructing temporal-mandibular joint disc anterior displacement screening model
CN120998532A
Model and system for screening temporal-mandibular joint disc anterior displacement and medical instrument
CN121075681A