Devices, Equipment and Storage Media Applied to Gastric Cancer Diagnosis and Prognosis

By combining traditional biomarkers and plasma exosome-related biomarkers, using a variety of machine learning algorithms to build gastric cancer diagnosis and prognosis models, the shortcomings of early screening methods for gastric cancer in the existing technology are solved and higher diagnostic and prognostic accuracy are achieved.

CN118748078BActive Publication Date: 2025-06-17THE SEVENTH AFFILIATED HOSPITAL SUN YAT SEN UNIV SHENZHEN
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410752622.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2025-06-17
Estimated Expiration
2044-06-12

AI Technical Summary

Technical Problem

The existing technology lacks ideal non-invasive and low-cost early screening methods for gastric cancer. Traditional serum tumor markers have low sensitivity and specificity, making it difficult to effectively diagnose early gastric cancer.

Method used

By obtaining test data from healthy people in multicenter cohorts and patients with gastric cancer, combining traditional biomarkers and plasma exosome-related biomarkers, a variety of machine learning algorithms are used to build diagnostic and prognostic models, and the best-performing model is determined to assist in diagnosis and prognostic evaluation.

Benefits of technology

It achieves higher accuracy in diagnosis and prognosis prediction than traditional methods, providing stronger auxiliary medical diagnosis effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118748078B_ABST
    Figure CN118748078B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a device, equipment and storage medium for gastric cancer diagnosis and prognosis. The device includes a data acquisition module configured to acquire test data of healthy populations and gastric cancer patients in a multi-center cohort; a diagnostic model construction module configured to construct multiple diagnostic models by combining multiple different machine learning algorithms based on the test data of the healthy populations and gastric cancer patients in the multi-center cohort, and perform performance evaluation on the multiple diagnostic models to determine the diagnostic model with the optimal performance as the optimal diagnostic model; a prognostic model construction module configured to construct multiple prognostic models by combining multiple different machine learning algorithms based on the test data and clinical prognostic information of gastric cancer patients in the multi-center cohort, and perform performance evaluation on the multiple prognostic models to determine the prognostic model with the optimal performance as the optimal prognostic model, and assist in the diagnosis and prognosis of gastric cancer through this device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a device, equipment and storage medium for gastric cancer diagnosis and prognosis, belonging to the technical field of medical auxiliary diagnosis equipment. Background Art

[0002] The global incidence and mortality of gastric cancer rank fifth among all malignant tumors. Most patients with gastric cancer (GC) are in the advanced stage (Advanced gastric cancer, AGC) at the time of diagnosis, and the survival rate is only 20%. In contrast, if early gastric cancer (EGC) is detected, the survival rate of patients can be increased to 70% through surgical treatment. Imaging techniques, gastroscopy, serological Helicobacter pylori, serum pepsinogen, and traditional biomarker detection are the most commonly used GC screening methods in clinical practice. In high-incidence countries of GC in East Asia, including China, current routine screening mainly relies on gastroscopy. However, this screening method is often limited by poor compliance of subjects, potential complication risks, and high economic costs. Therefore, there is an urgent clinical need to develop non-invasive and low-cost methods for diagnosing ECG. Exosomes are small lipid bilayer extracellular vesicles (30 - 150 nm) that are present in almost all body fluids. Exosomes released by tumor cells carry substances reflecting their characteristics. A large number of studies have shown that exosomal proteins can be used as tumor diagnostic markers.

[0003] Artificial Intelligence (AI) and its machine learning technologies have recently attracted wide attention in the medical field. Through data mining techniques, valuable information can be extracted from large-scale medical databases to develop machine learning models for diagnosis and prognosis prediction. The unique features of machine learning, including non-linearity, fault tolerance, and the ability to be retrained with updated databases, make it suitable for disease diagnosis and prognosis prediction.

[0004] In recent years, with the progress of gastroscopy technology, more and more early-stage GCs have been discovered and treated in a timely manner. However, gastroscopy is an invasive examination and is difficult to be popularized on a large scale. In addition, due to relatively low sensitivity and specificity, traditional serum tumor markers such as alpha-fetoprotein (AFP), carcinoembryonic antigen (CEA), carbohydrate antigen 199 (CA199), carbohydrate antigen 125 (CA125), and carbohydrate antigen 724 (CA724) are mainly used for the screening and treatment monitoring of digestive tract tumors including gastric cancer, rather than early diagnosis. In addition, other serum biomarkers (such as pepsinogen and anti-Helicobacter pylori IgG antibody) can detect pre-GC lesions, but only have a certain sensitivity for the diagnosis of GC. Therefore, there is still a lack of an ideal serum or plasma GC screening method, and new biomarkers need to be explored. Therefore, in the past few decades, more and more studies have begun to explore non-invasive and effective GC diagnostic biomarkers and identify effective early detection biomarkers for GC. The lipid bilayer structure of exosomes protects their contents from degradation and has a relatively stable content in various body fluids. The non-invasiveness and stability of plasma-derived exosomes make them ideal biomarkers. More and more studies have proven that plasma-derived exosomes have great development potential and application prospects in the early diagnosis and prognosis evaluation of GC. In addition, machine learning has been applied to the detection of many types of cancers, such as colorectal cancer and breast cancer, with high accuracy. In the field of gastric cancer, machine learning is mainly used to analyze endoscopic images obtained through invasive procedures. In contrast, the detection of routine blood, biochemical, and tumor markers is non-invasive and inexpensive. Therefore, how to use machine learning to construct a model for the diagnosis and prognosis evaluation of gastric cancer based on these non-invasive features is of great significance to the diagnostic efficiency. Summary of the Invention

[0005] To solve the above technical problems, embodiments of the present application respectively provide a device, a device and a storage medium applied to the diagnosis and prognosis of gastric cancer.

[0006] Other features and advantages of the present application will become apparent through the following detailed description, or be learned in part through the practice of the present application.

[0007] According to one aspect of the embodiments of the present application, a device applied to the diagnosis and prognosis of gastric cancer is provided, and the device includes:

[0008] A data acquisition module, configured to acquire test data of healthy people and gastric cancer patients in a multi-center cohort, where the test data includes traditional biomarkers and plasma exosome-related biomarkers, the traditional biomarkers include AFP, CEA, CA199, CA125, and CA724, and the plasma exosome-related biomarkers include total exosome protein concentration and exosome ALDOA protein concentration;

[0009] A diagnostic model construction module, configured to construct multiple diagnostic models by combining multiple different machine learning algorithms based on the test data of healthy people and gastric cancer patients in the multi-center cohort, and perform performance evaluation on the multiple diagnostic models to determine the diagnostic model with the best performance as the optimal diagnostic model;

[0010] A prognostic model construction module, configured to construct multiple prognostic models by combining multiple different machine learning algorithms based on the test data and clinical prognostic information of gastric cancer patients in the multi-center cohort, and perform performance evaluation on the multiple prognostic models to determine the prognostic model with the best performance as the optimal prognostic model.

[0011] Further, the machine learning algorithms used to construct multiple diagnostic models include Least Absolute Shrinkage and Selection Operator, Ridge Regression, Elastic Net, Generalized Linear Model, Support Vector Machine, Gradient Boosting with Component Linear Model, Linear Discriminant Analysis, Cox's Partial Least Squares Regression, Random Forest, Generalized Boosted Regression Model, Extreme Gradient Boosting Model, and Naive Bayes Model.

[0012] Further, the diagnostic model construction module is further configured to divide the test data of healthy people and gastric cancer patients in the multi-center cohort into a training set and a validation set. The training set is used to construct multiple diagnostic models. For each diagnostic model, calculate the ROC-AUC of the validation set, perform algorithm combination on tumor markers, fit the prediction model based on 10-fold cross-validation in the training cohort to obtain multiple models, and select the model with the highest AUC in the validation set as the optimal diagnostic model.

[0013] Further, the diagnostic model construction module is further configured to add a Scikit-learn module to the optimal diagnostic model to construct a new optimal diagnostic model. The Scikit-learn module is used to preprocess the input data, and use accuracy, AUC, recall rate, precision, F1 score, Kappa coefficient, and Matthews correlation coefficient for the new optimal diagnostic model, and use Shapley additive explanation and local explanation model to visualize the features affecting the new optimal diagnostic model, so as to analyze the importance of a single feature for affecting the prediction result. The features of the new optimal diagnostic model include one or a combination of AFP, CEA, CA199, CA125, CA724, total exosome protein concentration, and exosome ALDOA protein concentration.

[0014] Further, the machine learning algorithms used to construct multiple prognostic models include Lasso, Ridge, Enet, stepwise Cox regression, survival support vector machine, proportional hazards Cox boosting model, supervised principal component analysis, plsRcox, random survival forest, and GBM.

[0015] Further, the prognostic model construction module is further configured to divide the test data and clinical prognostic information of gastric cancer patients in the multi-center cohort into a training data set and a validation data set. Based on the training data set, use a combination of multiple machine learning algorithms to construct a prognostic model. Based on 10-fold cross-validation of the training cohort, perform algorithm combinations on tumor markers to obtain multiple prognostic models. Calculate the C-index of each prognostic model based on the validation data set, select the model with the highest C-index in the validation set as the optimal prognostic model, and visualize the optimal prognostic model.

[0016] According to one aspect of the embodiments of the present application, an electronic device is provided, including: a controller; a memory for storing one or more programs, and when the one or more programs are executed by the controller, the controller is caused to implement the following steps:

[0017] Obtain the test data of healthy people and gastric cancer patients in the multi-center cohort. The test data includes traditional biomarkers and plasma exosome-related biomarkers. The traditional biomarkers include AFP, CEA, CA199, CA125, and CA724. The plasma exosome-related biomarkers include total exosome protein concentration and exosome ALDOA protein concentration;

[0018] Based on the test data of healthy people and gastric cancer patients in the multi-center cohort, use a combination of multiple different machine learning algorithms to construct multiple diagnostic models, and perform performance evaluation on the multiple diagnostic models to determine the optimal diagnostic model with the best performance as the optimal diagnostic model;

[0019] Based on the test data and clinical prognosis information of gastric cancer patients in a multi - center cohort, a variety of different machine learning algorithms are combined to construct multiple prognostic models, and the performance of the multiple prognostic models is evaluated to determine the prognostic model with the best performance as the optimal prognostic model.

[0020] According to one aspect of the embodiments of the present application, there is also provided a computer - readable storage medium, on which computer - readable instructions are stored. When the computer - readable instructions are executed by a processor of a computer, the computer is caused to perform the following steps:

[0021] Obtain the test data of healthy people and gastric cancer patients in a multi - center cohort. The test data includes traditional biomarkers and plasma exosome - related biomarkers. The traditional biomarkers include AFP, CEA, CA199, CA125, and CA724, and the plasma exosome - related biomarkers include the total protein concentration of exosomes and the ALDOA protein concentration of exosomes;

[0022] Based on the test data of healthy people and gastric cancer patients in the multi - center cohort, a variety of different machine learning algorithms are combined to construct multiple diagnostic models, and the performance of the multiple diagnostic models is evaluated to determine the diagnostic model with the best performance as the optimal diagnostic model;

[0023] Based on the test data and clinical prognosis information of gastric cancer patients in a multi - center cohort, a variety of different machine learning algorithms are combined to construct multiple prognostic models, and the performance of the multiple prognostic models is evaluated to determine the prognostic model with the best performance as the optimal prognostic model.

[0024] In the technical solutions provided by the embodiments of the present application, there are at least the following advantages:

[0025] For the first time, the embodiments of the present application construct a GC diagnostic model and a prognostic model that combine traditional biomarkers and PDEV - related biomarkers based on machine learning. Compared with traditional diagnostic and prognostic models, they have higher prediction accuracy, stronger model performance, and better play the role of assisting medical diagnosis.

[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings

[0027] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the accompanying drawings:

[0028] Figure 1 is a structural diagram of a device applied to gastric cancer diagnosis and prognosis shown in an exemplary embodiment of the present application.

[0029] Figure 2 is a working flowchart of a device applied to gastric cancer diagnosis and prognosis shown in an exemplary embodiment of the present application.

[0030] Figure 3 is another working flowchart of a device applied to gastric cancer diagnosis and prognosis shown in an exemplary embodiment of the present application.

[0031] Figure 4 is a schematic diagram of the development and validation of a machine learning-based GC diagnostic model shown in an exemplary embodiment of the present application. (A) A total of 113 prediction models were established, and the AUC of each model in all datasets was calculated. (B) The performance of the diagnostic model based on Enet [alpha = 0.1] in predicting GC and healthy patients in the validation set. (C) The confusion matrix of the diagnostic model based on Enet [alpha = 0.1] in predicting GC and healthy patients in the validation set. (D) The diagnostic efficacy of the diagnostic model based on Enet [alpha = 0.1] for EGC. (E) The attributes of the features in the black box model. Each line represents a feature, and the abscissa is the SHAP value. Red dots indicate higher feature values, and blue dots indicate lower feature values. (F) The feature importance ranking represented by SHAP. The bar chart describes the importance of each feature in the development of the diagnostic prediction model.

[0032] Figure 5It is a schematic diagram of the interpretability of the diagnostic model shown in an exemplary embodiment of the present application. The SHAP analysis (A) and LIME algorithm (B) are used to show a healthy case in the diagnostic model. The SHAP force analysis (C) and LIME algorithm (D) are used to show a GC case. In (A) and (C), the red bars and blue bars represent risk factors and protective factors respectively; the longer bars indicate more important features. On the left side of the graphs in (B) and (D) are the results predicted using LIME. The middle part shows 7 variables that affect the determination of being healthy or having GC from top to bottom. The length of each feature bar represents the importance (weight) of the feature in making the prediction. Longer bars indicate a greater contribution to being healthy or having GC. The right panel shows the actual values of these 7 variables, and different colors indicate their different effects on being healthy or having GC. (E) Overall interpretability of all samples in the validation cohort based on the SHAP heatmap.

[0033] Figure 6 It is a schematic diagram of constructing and validating a GC prognosis model based on machine learning shown in an exemplary embodiment of the present application. (A) A total of 101 prognosis models were established, and the C-index of each model in the training set and validation set was calculated. The model with the highest C-index in the validation set was taken as the final model. Validation of the StepCox[forward]+GBM prognosis model in the training cohort (B) and validation cohort (C) (Kaplan-Meier survival analysis based on the median risk score). The upper left part is the distribution diagram of the relationship between the risk score and the survival status; the lower left part is the heatmap of the distribution of 7 different biomarkers in the GC patients in the cohort; the upper right part is the ROC curve of the risk scores in different cohorts; the lower right part is the survival curve between the high and low risk score groups.

[0034] Figure 7 It is a schematic diagram of the clinicopathological features of different risk subgroups in the StepCox[forward]+GBM prognosis model shown in an exemplary embodiment of the present application: (A) Training cohort; (B) Validation cohort. Detailed implementation manners

[0035] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0036] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0037] The flowcharts shown in the accompanying drawings are only illustrative descriptions and do not necessarily include all content and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0038] In this application, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0039] Please refer to Figure 1 , Figure 1 which is a structural diagram of a device for gastric cancer diagnosis and prognosis shown in an exemplary embodiment of this application. One aspect of this application provides a device for gastric cancer diagnosis and prognosis. The device 100 includes a data acquisition module 101, a diagnostic model construction module 102, and a prognostic model construction module 103.

[0040] Please further refer to Figure 2 , Figure 2 which is a working flowchart of a device for gastric cancer diagnosis and prognosis shown in an exemplary embodiment of this application. The data acquisition module 101, the diagnostic model construction module 102, and the prognostic model construction module 103 are respectively configured to execute steps S201 to S203 as shown in Figure 2 below.

[0041] Step S201: Obtain the test data of healthy people and gastric cancer patients in a multi - center cohort. The test data includes traditional biomarkers and plasma exosome - related biomarkers. The traditional biomarkers include AFP, CEA, CA199, CA125, and CA724. The plasma exosome - related biomarkers include the total protein concentration of exosomes and the ALDOA protein concentration of exosomes.

[0042] Step S201 is specifically executed by the data acquisition module 101. The data acquisition module 101 can obtain the above data by keyboard input. The data acquisition module 101 can be implemented as a module with data storage function, which stores the data required for the subsequent diagnostic model construction module 102 and prognosis model construction module 103 to construct the corresponding models. When data is needed, the diagnostic model construction module 102 and the prognosis model construction module 103 extract data from the data acquisition module 101.

[0043] This embodiment provides the specific sources of the test data of the multi-center cohort of healthy people and gastric cancer patients here.

[0044] GC patients and healthy people undergoing physical examinations who visited the First Affiliated Hospital of Zhengzhou University, the Seventh Affiliated Hospital of Sun Yat-sen University, and the People's Hospital of Fengqing County, Yunnan Province from January 2019 to December 2023 were included. The First Affiliated Hospital of Zhengzhou University was used as the training cohort, and the Seventh Affiliated Hospital of Sun Yat-sen University and the People's Hospital of Fengqing County, Yunnan Province were used as the validation cohorts. Inclusion criteria for GC patients: (1) Complete basic information, clinicopathological, and follow-up data; (2) Patients diagnosed as GC according to the pathological diagnosis guidelines (preoperative or postoperative); (3) Preoperative gastric cancer patients were those scheduled to undergo GC radical resection at the multi-center points of this example; Exclusion criteria: (1) Unwilling to participate in this example; (2) Complicated with hematological and immune system diseases; (3) Complicated with undifferentiated carcinoma, sarcoma, gastric lymphoma, stromal tumor, and other tumors; (4) Those who have received anti-tumor treatments such as preoperative neoadjuvant radiotherapy and chemotherapy; (5) Patients who were lost to follow-up or died within 1 month after radical surgery. Inclusion criteria for healthy people: (1) Normal physical examination; (2) Routine biochemistry: no obvious abnormalities in the three major routine tests, liver and kidney functions, erythrocyte sedimentation rate, blood biochemistry, blood glucose, and blood lipids; (3) No obvious abnormalities were found in electrocardiogram, chest radiograph, abdominal color Doppler ultrasound, or CT examination; (4) No obvious abnormalities were found in gastroscopy examination. Exclusion criteria for healthy people: (1) Unwilling to participate in this example; (2) History of combined hematological and immune system diseases; This example was carried out in accordance with the Declaration of Helsinki and approved by the ethics committees of the above three units. All selected patients and normal physical examination healthy people agreed to participate in this study and signed the informed consent form.

[0045] All GC patients and healthy subjects participating in this example were asked to fast for more than 8 hours, and then 5 ml of fasting venous blood was collected. The blood samples were placed in vacuum clot activator tubes and the basic information of all participants was labeled in sequence. Hemolysis was avoided as much as possible during this process. Then, the specimens were placed at room temperature for 1 hour and centrifuged at 2000 rpm for 10 minutes to obtain the upper plasma. Further centrifugation was performed at 3000 rpm for 10 minutes to remove substances such as red blood cells, peripheral blood mononuclear cells, and cell debris. The plasma was transferred to a new EP tube and stored in an -80°C refrigerator for subsequent use. The basic information and clinicopathological data included gender, age, height, smoking history, drinking history, family history of malignant tumors, weight, body mass index (BMI), alpha-fetoprotein (AFP), CEA, CA125, CA724, CA199, date of first diagnosis, tumor location, Borrmann classification, Lauren classification, and TNM stage.

[0046] All subjects in this example were followed up through outpatient reexamination, inpatient reexamination, and telephone follow-up. The follow-up investigation content included treatment status, efficacy, time of death, and other follow-up information. The definition of the overall survival (OS) of GC patients included the following aspects: (1) the time from radical surgery to the time of death of GC patients due to any cause; (2) if the patient was still alive at the end of the follow-up in this example, the last follow-up time was used as the censoring date of their overall survival (February 1, 2024).

[0047] Subject plasma-derived exosomes were extracted by differential centrifugation, and the absolute expression level of ALDOA was detected by enzyme-linked immunosorbent assay (ELISA), and the total protein concentration was detected by bicinchoninic acid assay (BCA).

[0048] Step S202: Based on the test data of the multi-center cohort of healthy people and gastric cancer patients, multiple different machine learning algorithms were combined to construct multiple diagnostic models, and the performance of the multiple diagnostic models was evaluated to determine the diagnostic model with the best performance as the optimal diagnostic model.

[0049] Step S202 was executed by the diagnostic model construction module 102.

[0050] In an exemplary embodiment, the diagnostic model construction module 102 was further configured as:

[0051] Sort out the test data of healthy people and gastric cancer patients in a multi-center cohort (5 traditional biomarkers: AFP, CEA, CA199, CA125, and CA724, and 2 plasma exosome-related biomarkers: total exosome protein concentration and exosome ALDOA protein concentration), and use 12 machine learning algorithms to construct a diagnostic model, including algorithms such as Least Absolute Shrinkage and Selection Operator (Lasso), Ridge Regression, Elastic Net (Enet), Generalized Linear Model (GLM), Support Vector Machine (SVM), gradient boosting with component-wise linear models (GlmBoost), linear discriminant analysis (LDA), partial least squares regression for Cox (plsRcox), random forest (RF), generalized boosted regression modeling (GBM), eXtreme Gradient Boosting (XGBoost), and Naive Bayes model. For each combined model, calculate the ROC-AUC of the validation set, combine the algorithms for tumor markers, and fit the prediction model based on 10-fold cross-validation in the training cohort to obtain a total of 112 models. Select the model with the highest AUC in the validation set as the optimal model, and further preprocess the data and construct a diagnostic model using the Scikit-learn module in Python. Use Accuracy, AUC, Recall, Precision, F1 Score, Kappa coefficient, and Matthews Correlation Coefficient (MCC) to evaluate the model. Subsequently, use Shapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) to visualize the features (tumor markers) that affect the diagnostic model, so as to analyze the importance of individual features for influencing the prediction results.

[0052] In step S203, based on the test data and clinical prognosis information of gastric cancer patients in the multi-center cohort, a combination of multiple different machine learning algorithms is used to construct multiple prognosis models, and the performance of the multiple prognosis models is evaluated to determine the prognosis model with the best performance as the optimal prognosis model.

[0053] Step S203 is executed by the prognosis model construction module 103.

[0054] In an exemplary embodiment, the prognosis model construction module 103 is further configured to:

[0055] Sort out the test data and clinical prognosis information of gastric cancer patients in the multi-center cohort, and further use a combination of 10 machine learning algorithms such as Lasso, Ridge, Enet, Stepwise Cox Regression (StepCox), Survival Support Vector Machine (survivalSVM), Cox Proportional Hazards Boosting (CoxBoost), Supervised Principal Component Analysis (SuperPC), plsRcox, Random Survival Forests (RSF), and GBM to construct a prognosis model. Subsequently, based on 10-fold cross-validation of the training cohort, a combination of algorithms is performed on tumor markers, and a total of 97 models are obtained. Calculate the Harrell’s concordance index (C-index) for the validation set of each model, select the model with the highest C-index in the validation set as the optimal model, and visualize the model.

[0056] In an exemplary embodiment, please refer to Figure 3 , which is another workflow diagram of the device applied to the diagnosis and prognosis of gastric cancer. The steps S301 to S302 included in this flowchart are respectively executed by Figure 1 the data acquisition module 101, the diagnosis model construction module 102, and the prognosis model construction module 103 in the device structure shown.

[0057] In step S301, obtain the characteristics of gastric cancer patients and healthy people in multi-center medical units.

[0058] The research subjects included in three hospitals were screened, and finally a total of 539 eligible research subjects were included. There were 270 in the training cohort, including 143 healthy people, 92 preoperative GC patients, and 35 postoperative GC patients. There were 269 in the validation cohort, including 143 healthy people, 92 preoperative GC patients, and 34 postoperative GC patients. Their detailed basic information and clinicopathological characteristics are shown in Table 1.

[0059]

[0060]

[0061]

[0062]

[0063] Table Notes: Healthy donor (HD); Preoperative patient with gastric cancer (Pre-op GC); Postoperative patients with gastric cancer (Post-op GC); Body Mass Index (BMI); Tumor location: upper (U); middle (M); low (L); The tumor invaded more than one-third of the stomach (UML); Borrmann classification: type I polypoid, type II - ulcerative type (II-Ulcerative Type), type III - infiltrative ulcerative type (III-Infiltrative Ulcerative Type), type IV - diffuse infiltrative type (IV-Diffuse Infiltrative Type); Lauren classification: diffuse subtype, intestinal subtype, mixed subtype; American Joint Committee on Cancer stage (AJCC stage); Alpha-fetoprotein (AFP); Carcinoembryonic Antigen (CEA); Carbohydrate Antigen 125 (CA125); Carbohydrate antigen 72-4 (CA72-4); Carbohydrate antigen 19-9 (CA19-9); Not Available (N.A.).

[0064] Step S302, establish and interpret a gastric cancer diagnosis model.

[0065] Process the data of the training set and the validation set through machine learning algorithms to construct a GC diagnosis model that combines traditional biomarkers and PDEV-related biomarkers. By fitting 112 prediction models, calculate the AUC of each model in the training set and the validation set (Figure 4 In A). The optimal model we finally selected is the diagnostic model with the highest AUC value in the training set, namely Enet[alpha = 0.1]. This algorithm incorporates 5 traditional biomarkers and 2 PDEV-related biomarkers into the model, and uses the Enet[alpha = 0.1] algorithm to fit these features into the model. Figure 4 In B shows that the AUC value of the diagnostic model constructed by these 7 diagnostic markers in the training set is 0.928 (accuracy: 0.843; precision: 0.848), while the AUC value in the validation set is 0.850 (accuracy: 0.804; precision: 0.795). Figure 4 In C is the confusion matrix of the diagnostic model in the validation set, Figure 4 In D shows the diagnostic efficacy of the diagnostic model for EGC. Its AUC value in the training set is 0.77, 95% CI [0.69 - 0.84], while the AUC value in the validation set is 0.72, 95% CI [0.63 - 0.81], showing good diagnostic ability for EGC. We further used SHAP to demonstrate the importance and direction of each predictive feature for each outcome ( Figure 4 In E - F). Each predictive variable is assigned to the y-axis according to its relative importance, with the most significant predictive variable at the top. In the figure, the position of the points on the x-axis represents the percentage contribution of each subject to the overall SHAP value, and the highly positive contributions are located on the far right of the figure, where ALDOA in PDEV is the most important predictive feature for GC.

[0066] We further used SHAP analysis and the LIME algorithm to explain healthy and GC patients by sampling two samples from the validation set. Figure 5 In A - B is an example of a healthy person using SHAP analysis and the LIME algorithm. In this example, the Enet[alpha = 0.1] model detected ALDOA in PDEV as 34.68, total protein in PDEV as 137.60, CA199 as 1.50, CEA as 0.96, AFP as 1.00, CA724 as 1.53, and CA125 as 9.32, and there is no risk of having GC. The result predicted by the Enet[alpha = 0.1] model is that this example is a healthy population, which is consistent with the actual result. Similarly, Figure 5Among them, cases C-D are GC cases using SHAP force analysis and LIME algorithm. In this example, the Enet[alpha = 0.1] model detected that the ALDOA in PDEV was 246.56, the total protein in PDEV was 485.23, CA199 was 1499.71, CEA was 2.33, AFP was 1.41, CA724 was 9.30, and CA125 was 50.17. Among them, ALDOA in PDEV, total protein in PDEV, CA199, CA724, and CA125 all increased the risk of GC, while CEA and AFP reduced the risk of GC. The result predicted by the Enet[alpha = 0.1] model was GC patients, and the actual result was also GC. Figure 5 Figure E depicts the overall interpretation heatmap of all healthy individuals and GC samples in the validation cohort.

[0067] Step S303, construction and validation of the gastric cancer patient prognosis model.

[0068] Seven biomarkers were used to construct a GC prognosis model through machine learning algorithms. In the training set, we fitted prognosis models of 97 algorithm types ( Figure 6 Figure A), and calculated the C-index of each model on the validation set. The optimal model was StepCox[forward]+GBM, which had the highest C-index of 0.788 in the validation set. This algorithm was selected for subsequent analysis. In the training set and the validation set, risk scores were calculated for all GC patients and divided into high and low groups according to the median value ( Figure 6 Figures B-C). The heatmap results showed that as the risk score increased, the expression of the seven biomarkers also gradually increased, and Kaplan-Meier survival analysis showed that the OS of GC patients in the low-risk score group was significantly better than that of patients with higher risk scores (p < 0.0001). Further, ROC analysis was used to measure the performance of the optimal model. The AUCs at 1 year, 2 years, and 3 years in the training cohort were 0.85, 0.9, and 0.98 respectively; the AUCs at 1 year, 2 years, and 3 years in the validation cohort were 0.83, 0.71, and 0.92 respectively. Therefore, this risk model has a certain predictive value for the OS of GC patients, indicating its application potential in translational medicine.

[0069] To further evaluate the clinicopathological characteristics of patients in different risk subgroups of the prognosis risk model. Heatmaps were used to show the distribution of different patient basic information and clinicopathological characteristics in the high and low risk groups in the two cohorts. The results showed that in the training set ( Figure 7In (A), the high-risk group showed deeper tumor invasion, more lymph node metastases, distant metastases, and later AJCC and clinical stages, while there were no significant differences in gender, age, smoking history, drinking history, family history of cancer, tumor location, Borrmann classification, and Lauren classification. In the validation set ( Figure 7 In (B), the high-risk group showed similar clinicopathological features. Notably, in patients in the high-risk group, the body mass index seemed to be smaller, suggesting that the high tumor burden in AGC patients consumes the body. These results reflect the consistency of clinicopathological features in subgroups based on the prognostic risk model.

[0070] Another aspect of the present application also provides an electronic device, including: a controller; a memory for storing one or more programs, which when executed by the controller, perform the following steps: obtaining test data of multi-center cohort healthy populations and gastric cancer patients, the test data including traditional biomarkers and plasma exosome-related biomarkers, the traditional biomarkers including AFP, CEA, CA199, CA125, and CA724, and the plasma exosome-related biomarkers including total exosome protein concentration and exosome ALDOA protein concentration; based on the test data of the multi-center cohort healthy populations and gastric cancer patients, using a combination of multiple different machine learning algorithms to construct multiple diagnostic models, and performing performance evaluation on the multiple diagnostic models to determine the diagnostic model with the optimal performance as the optimal diagnostic model; based on the test data and clinical prognostic information of gastric cancer patients in the multi-center cohort, using a combination of multiple different machine learning algorithms to construct multiple prognostic models, and performing performance evaluation on the multiple prognostic models to determine the prognostic model with the optimal performance as the optimal prognostic model.

[0071] In particular, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it executes various functions defined in the system of the present application.

[0072] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0074] The modules involved in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0075] On the other hand, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining test data of healthy people and gastric cancer patients in a multi-center cohort, the test data including traditional biomarkers and plasma exosome-related biomarkers, the traditional biomarkers including AFP, CEA, CA199, CA125, and CA724, and the plasma exosome-related biomarkers including total exosome protein concentration and exosome ALDOA protein concentration; based on the test data of healthy people and gastric cancer patients in the multi-center cohort, using a combination of multiple different machine learning algorithms to construct multiple diagnostic models, and performing performance evaluation on the multiple diagnostic models to determine the diagnostic model with the best performance as the optimal diagnostic model; based on the test data and clinical prognosis information of gastric cancer patients in the multi-center cohort, using a combination of multiple different machine learning algorithms to construct multiple prognosis models, and performing performance evaluation on the multiple prognosis models to determine the prognosis model with the best performance as the optimal prognosis model.

[0076] The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device.

[0077] Another aspect of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the following steps: obtaining test data of healthy people and gastric cancer patients in a multi-center cohort, where the test data includes traditional biomarkers and plasma exosome-related biomarkers, the traditional biomarkers include AFP, CEA, CA199, CA125, and CA724, and the plasma exosome-related biomarkers include total exosome protein concentration and exosome ALDOA protein concentration; based on the test data of healthy people and gastric cancer patients in the multi-center cohort, using a combination of various different machine learning algorithms to construct multiple diagnostic models, and performing performance evaluation on the multiple diagnostic models to determine the diagnostic model with the best performance as the optimal diagnostic model; based on the test data and clinical prognosis information of gastric cancer patients in the multi-center cohort, using a combination of various different machine learning algorithms to construct multiple prognosis models, and performing performance evaluation on the multiple prognosis models to determine the prognosis model with the best performance as the optimal prognosis model.

[0078] According to one aspect of the embodiments of the present application, a computer system is also provided, including a central processing unit (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage part into a random access memory (RAM), such as performing the steps executed by the above computer device. In the RAM, various programs and data required for system operation are also stored. The CPU, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0079] The following components are connected to the I / O interface: an input section including a keyboard, a mouse, etc.; an output section including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section including a hard disk, etc.; and a communication section including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the I / O interface as required. A removable medium such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive as required so that a computer program read from it is installed into the storage section as required.

[0080] The above content is only a preferred exemplary embodiment of the present application and is not used to limit the implementation of the present application. Those of ordinary skill in the art can easily make corresponding adaptations or modifications according to the main concept and spirit of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope required by the claims.

Claims

1. A device for gastric cancer diagnosis and prognosis, characterized in that: The device comprises: A data acquisition module is configured to acquire test data of a multi-center cohort of healthy people and gastric cancer patients, wherein the test data includes traditional biomarkers and plasma exosome-related biomarkers, wherein the traditional biomarkers include AFP, CEA, CA199, CA125 and CA724, and the plasma exosome-related biomarkers include exosome total protein concentration and exosome ALDOA protein concentration; A diagnostic model building module is configured to build multiple diagnostic models based on the test data of healthy people and gastric cancer patients in the multi-center cohort by combining multiple different machine learning algorithms, and to perform performance evaluation on the multiple diagnostic models to determine the diagnostic model with the best performance as the optimal diagnostic model; A prognostic model building module is configured to construct multiple prognostic models based on the test data and clinical prognostic information of gastric cancer patients in a multi-center cohort using a combination of multiple different machine learning algorithms, and to perform performance evaluation on the multiple prognostic models to determine the prognostic model with the best performance as the optimal prognostic model; the clinical prognostic information includes basic patient information and clinical pathological characteristics; The diagnostic model building module is further configured to add a Scikit-learn module to the optimal diagnostic model to build a new optimal diagnostic model, the Scikit-learn module is used to preprocess the input data, use accuracy, AUC, recall rate, precision, F1 score, Kappa coefficient and Matthews correlation coefficient to evaluate the new optimal diagnostic model, and use Shapley summation interpretation and local interpretation model to visualize the features affecting the new optimal diagnostic model to analyze the importance of a single feature in affecting the prediction result, the features of the new optimal diagnostic model include one or more combinations of AFP, CEA, CA199, CA125, CA724, exosome total protein concentration and exosome ALDOA protein concentration; The prognostic model building module is further configured to divide the test data and clinical prognostic information of gastric cancer patients in a multicenter cohort into a training data set and a validation data set, build a prognostic model based on the training data set using a combination of multiple machine learning algorithms, perform multiple algorithm combinations on tumor markers based on 10-fold cross-validation of the training cohort to obtain multiple prognostic models, calculate the C-index of each prognostic model based on the validation data set, select the model with the highest C-index in the validation set as the optimal prognostic model, and visualize the optimal prognostic model.

2. The device for gastric cancer diagnosis and prognosis according to claim 1, characterized in that: The machine learning algorithms used to build multiple diagnostic models include least absolute shrinkage and selection operators, ridge regression, elastic net, generalized linear models, support vector machines, gradient boosting with component linear models, linear discriminant analysis, Cox's partial least squares regression, random forest, generalized boosted regression model, extreme gradient boosting model, and naive Bayes model.

3. The device for gastric cancer diagnosis and prognosis according to claim 1, characterized in that: The diagnostic model building module is further configured to divide the test data of healthy people and gastric cancer patients in the multi-center cohort into a training set and a validation set, the training set is used to build multiple diagnostic models, for each diagnostic model, the ROC-AUC of the validation set is calculated, the algorithm is combined for tumor markers, and the prediction model based on 10-fold cross-validation in the training cohort is fitted to obtain multiple models, and the model with the highest AUC in the validation set is selected as the optimal diagnostic model.

4. The device for gastric cancer diagnosis and prognosis according to claim 1, characterized in that: The machine learning algorithms used to build multiple prognostic models included Lasso, Ridge, Enet, stepwise Cox regression, survival support vector machine, proportional hazards Cox boosting model, supervised principal component analysis, plsRcox, random survival forest, and GBM.

5. An electronic device, characterized in that: include: A controller; a memory for storing one or more programs, when the one or more programs are executed by the controller, the controller implements the following steps: Obtaining test data of healthy people and gastric cancer patients in a multicenter cohort, the test data including traditional biomarkers and plasma exosome-related biomarkers, the traditional biomarkers including AFP, CEA, CA199, CA125 and CA724, and the plasma exosome-related biomarkers including exosome total protein concentration and exosome ALDOA protein concentration; Based on the test data of healthy people and gastric cancer patients in the multi-center cohort, a plurality of different machine learning algorithms are combined to construct a plurality of diagnostic models, and the performance of the plurality of diagnostic models is evaluated to determine the diagnostic model with the best performance as the optimal diagnostic model; Based on the test data and clinical prognostic information of gastric cancer patients in a multi-center cohort, a plurality of different machine learning algorithms are combined to construct a plurality of prognostic models, and the performance of the plurality of prognostic models is evaluated to determine the prognostic model with the best performance as the optimal prognostic model; the clinical prognostic information includes basic information of the patient and clinical pathological characteristics; Based on the test data of healthy people and gastric cancer patients in the multi-center cohort, a plurality of different machine learning algorithms are combined to construct a plurality of diagnostic models, and the performance of the plurality of diagnostic models is evaluated to determine the diagnostic model with the best performance as the optimal diagnostic model, including: A Scikit-learn module is added to the optimal diagnostic model to construct a new optimal diagnostic model, wherein the Scikit-learn module is used to preprocess the input data, and the new optimal diagnostic model is evaluated using accuracy, AUC, recall rate, precision, F1 score, Kappa coefficient and Matthews correlation coefficient, and the features affecting the new optimal diagnostic model are visualized using Shapley summation interpretation and local interpretation models to analyze the importance of a single feature in affecting the prediction results, wherein the features of the new optimal diagnostic model include one or more combinations of AFP, CEA, CA199, CA125, CA724, exosome total protein concentration and exosome ALDOA protein concentration; Based on the test data and clinical prognosis information of gastric cancer patients in a multi-center cohort, a plurality of different machine learning algorithms are combined to construct multiple prognostic models, and the performance of the multiple prognostic models is evaluated to determine the prognostic model with the best performance as the optimal prognostic model, including: The test data and clinical prognosis information of gastric cancer patients in a multicenter cohort were divided into a training data set and a validation data set. A prognosis model was constructed based on the training data set using a combination of multiple machine learning algorithms. Based on 10-fold cross-validation of the training cohort, multiple algorithm combinations were performed on tumor markers to obtain multiple prognosis models. The C-index of each prognosis model was calculated based on the validation data set, and the model with the highest C-index in the validation set was selected as the optimal prognosis model, and the optimal prognosis model was visualized.

6. A computer-readable storage medium, characterized in that: Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the following steps: Obtaining test data of healthy people and gastric cancer patients in a multicenter cohort, the test data including traditional biomarkers and plasma exosome-related biomarkers, the traditional biomarkers including AFP, CEA, CA199, CA125 and CA724, and the plasma exosome-related biomarkers including exosome total protein concentration and exosome ALDOA protein concentration; Based on the test data of healthy people and gastric cancer patients in the multi-center cohort, a plurality of different machine learning algorithms are combined to construct a plurality of diagnostic models, and the performance of the plurality of diagnostic models is evaluated to determine the diagnostic model with the best performance as the optimal diagnostic model; Based on the test data and clinical prognostic information of gastric cancer patients in a multi-center cohort, a plurality of different machine learning algorithms are combined to construct a plurality of prognostic models, and the performance of the plurality of prognostic models is evaluated to determine the prognostic model with the best performance as the optimal prognostic model; the clinical prognostic information includes basic information of the patient and clinical pathological characteristics; Based on the test data of healthy people and gastric cancer patients in the multi-center cohort, a plurality of different machine learning algorithms are combined to construct a plurality of diagnostic models, and the performance of the plurality of diagnostic models is evaluated to determine the diagnostic model with the best performance as the optimal diagnostic model, including: A Scikit-learn module is added to the optimal diagnostic model to construct a new optimal diagnostic model, wherein the Scikit-learn module is used to preprocess the input data, and the new optimal diagnostic model is evaluated using accuracy, AUC, recall rate, precision, F1 score, Kappa coefficient and Matthews correlation coefficient, and the features affecting the new optimal diagnostic model are visualized using Shapley summation interpretation and local interpretation models to analyze the importance of a single feature in affecting the prediction results, wherein the features of the new optimal diagnostic model include one or more combinations of AFP, CEA, CA199, CA125, CA724, exosome total protein concentration and exosome ALDOA protein concentration; Based on the test data and clinical prognosis information of gastric cancer patients in a multi-center cohort, a plurality of different machine learning algorithms are combined to construct multiple prognostic models, and the performance of the multiple prognostic models is evaluated to determine the prognostic model with the best performance as the optimal prognostic model, including: The test data and clinical prognosis information of gastric cancer patients in a multicenter cohort were divided into a training data set and a validation data set. A prognosis model was constructed based on the training data set using a combination of multiple machine learning algorithms. Based on 10-fold cross-validation of the training cohort, multiple algorithm combinations were performed on tumor markers to obtain multiple prognosis models. The C-index of each prognosis model was calculated based on the validation data set, and the model with the highest C-index in the validation set was selected as the optimal prognosis model, and the optimal prognosis model was visualized.

Citation Information

Patent Citations

  • Improvements on vehicle springs

    CA13760A

  • Fertilizer

    CA48523A

  • Lung cancer diagnosis system based on multiple machine learning algorithms

    CN112259221A

  • Gastric cancer screening serum biomarker group and application thereof

    CN114354933A