Metabolism-related fatty liver disease intelligent prediction method and system and storage medium
By integrating intelligent tongue imaging parameters with clinical indicators, a predictive model for metabolic-related fatty liver disease was constructed, which solved the problems of high early diagnosis costs or low specificity in existing technologies and achieved early and accurate prediction of metabolic-related fatty liver disease.
Patent Information
- Application Number
- CN202511159666.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies for the early diagnosis of metabolic-related fatty liver disease are either costly or lack specificity, making it difficult to achieve accurate early diagnosis.
By integrating intelligent tongue imaging parameters with clinical indicators, and collecting multi-source data to extract quantitative and qualitative tongue imaging parameters, combined with random forest classification and LASSO regression analysis, key predictive variables were screened to construct a predictive model for metabolic-related fatty liver disease.
It enables early and accurate prediction of metabolic-related fatty liver disease, improving diagnostic specificity and efficiency while reducing costs.
Smart Images

Figure CN120977548A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of liver disease prediction technology, and in particular to an intelligent prediction method, system and storage medium for metabolic-related fatty liver disease. Background Technology
[0002] Metabolic fatty liver disease (MAFLD) has become the leading cause of chronic liver disease worldwide, seriously endangering human health. Because MAFLD often presents with no obvious symptoms in its early stages, most patients are diagnosed at a relatively advanced stage. Therefore, early diagnosis and intervention are crucial for slowing disease progression.
[0003] Currently, although various diagnostic methods exist, they suffer from problems such as high cost, complex operation, or low specificity. Traditional Chinese medicine tongue diagnosis, as one of the important bases for disease diagnosis, combined with modern machine learning technology, holds promise for providing new insights into the early diagnosis of MAFLD.
[0004] In conclusion, it is essential to propose an intelligent prediction method for metabolic-related fatty liver disease that integrates intelligent tongue image parameters with clinical indicators to achieve early prediction of metabolic-related fatty liver disease. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent prediction method, system, and storage medium for metabolic-related fatty liver disease, aiming to solve the technical problem of early prediction of metabolic-related fatty liver disease by integrating intelligent tongue image parameters and clinical indicators.
[0006] To achieve the above objectives, the present invention employs an intelligent prediction method for metabolic-related fatty liver disease, comprising the following steps:
[0007] Collect multi-source data to obtain basic demographic information, laboratory indicators, predictive indicators, and tongue images, and extract quantitative and qualitative tongue image parameters;
[0008] Variable screening was conducted on tongue appearance parameters and clinical indicators to identify key predictive variables;
[0009] Based on the key predictor variables, obtain the values of key variables, obtain the prediction results of the risk of developing metabolic-related fatty liver disease, and output them.
[0010] Among them, in the steps of collecting multi-source data, obtaining basic demographic information, laboratory indicators, predictive indicators, and tongue images, and extracting quantitative and qualitative tongue image parameters:
[0011] Collect basic information, experimental indicators, and fatty liver prediction indicators of the subjects;
[0012] The system collects images of the subject's tongue, extracts the color features of the tongue body and tongue coating, and identifies conditions such as cracks, punctures, and teeth marks, outputting qualitative and quantitative tongue image parameters.
[0013] Among the steps, in collecting images of the subject's tongue, extracting the color features of the tongue body and tongue coating, identifying cracks, punctures, teeth marks, etc., and outputting qualitative and quantitative tongue image parameters:
[0014] The tongue image region is cropped from the original tongue image and segmented into a tongue body image and a tongue coating image;
[0015] Extract color features of the tongue body and tongue coating;
[0016] Identify cracks on the tongue and calculate a crack score for the tongue body.
[0017] Identify tongue shape features and calculate tongue shape scores;
[0018] Identify tongue punctures and teeth marks, and calculate the scores for tongue punctures and teeth marks respectively;
[0019] Identify tongue coating texture features and calculate tongue coating stickiness score and tongue coating thickness score;
[0020] Output qualitative tongue image parameters, including tongue color, cracked tongue, punctate tongue, teeth-marked tongue, yellow coating, and thick coating.
[0021] Before the step of cropping the tongue image region from the original tongue image and segmenting it into tongue body images and tongue coating images:
[0022] The original tongue images were quality-assessed to ensure that they met the requirements in terms of brightness, sharpness, and target distance.
[0023] Among the steps involved in screening tongue image parameters and clinical indicators to determine key predictive variables:
[0024] Preliminary screening of tongue image parameters and clinical indicators was conducted, and statistically significant tongue image parameters and clinical indicators were selected and output as the first variable.
[0025] Random forest classification was performed on the results of the first variable of tongue image parameters and clinical indicators, and bar charts of feature importance were drawn. Variables were initially screened according to the ranking of feature importance, and the second variable was output.
[0026] LASSO regression analysis was performed on the first variable results of tongue appearance parameters and clinical indicators. The coefficient profile and cross-validation curve of LASSO regression were plotted to screen out key variables and output the third variable.
[0027] The process involves performing LASSO regression analysis on the first variables of tongue image parameters and clinical indicators, plotting the coefficient profile and cross-validation curves of the LASSO regression, screening out key variables, and outputting the third variable.
[0028] Find the intersection of the second and third variables to determine the final key predictor variables.
[0029] Among the steps, the following steps are involved: obtaining key variable values based on key predictor variables, acquiring the predicted risk of metabolic-related fatty liver disease, and outputting the results:
[0030] Extract specific values for key predictor variables from the collected data;
[0031] The probability of developing metabolic-related fatty liver disease is calculated based on specific numerical values.
[0032] After calculating the probability of developing metabolic-related fatty liver disease based on specific numerical values:
[0033] Output the predicted risk of developing metabolic-related fatty liver disease, including the probability value of the risk and the corresponding risk level.
[0034] This invention also provides an intelligent prediction system for metabolic-related fatty liver disease, comprising a multi-source data acquisition module, a variable screening module, and a prediction result acquisition module; wherein:
[0035] The multi-source data acquisition module is used to collect multi-source data, obtain basic statistical information, experimental indicators, predictive indicators, and tongue images, and extract quantitative and qualitative tongue image parameters.
[0036] The variable screening module is used to screen tongue image parameters and clinical indicators to determine key predictive variables.
[0037] The prediction result acquisition module is used to obtain key variable values based on key predictor variables, obtain prediction results of the risk of metabolic-related fatty liver disease, and output them.
[0038] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, implements the intelligent prediction method for metabolic-related fatty liver disease.
[0039] This invention discloses an intelligent prediction method, system, and storage medium for metabolic-related fatty liver disease. The method comprises the following steps using a multi-source data acquisition module, a variable screening module, and a prediction result acquisition module: acquiring multi-source data, including basic demographic information, laboratory indicators, prediction indicators, and tongue images, and extracting quantitative and qualitative tongue image parameters; screening the tongue image parameters and clinical indicators to determine key predictive variables; obtaining key variable values based on the key predictive variables to obtain and output the prediction result of the risk of developing metabolic-related fatty liver disease; and achieving early prediction of metabolic-related fatty liver disease by fusing intelligent tongue image parameters and clinical indicators. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart of the steps in the intelligent prediction method for metabolic-related fatty liver disease of the present invention.
[0042] Figure 2 This is a flowchart of steps S100 of the present invention.
[0043] Figure 3 This is a flowchart of the intelligent recognition process for tongue image parameters according to the present invention.
[0044] Figure 4 This is a schematic diagram of the qualitative results of tongue image recognition according to the present invention.
[0045] Figure 5 This is a flowchart of steps S200 of the present invention.
[0046] Figure 6 This is the logistic regression forest graph of the tongue image parameters in the training set of this invention.
[0047] Figure 7 This invention uses a logistic regression forest plot of clinical indicators from the training set.
[0048] Figure 8 This is a schematic diagram of variable selection for the random forest and LASSO regression models of this invention.
[0049] Figure 9 This is a flowchart of steps S300 of the present invention.
[0050] Figure 10 This invention compares the performance of nine MAFLD machine learning prediction models.
[0051] Figure 11 This is the internal and external validation of the random forest model of this invention.
[0052] Figure 12 The ROC curves and Spearman correlation heatmaps for MAFLD prediction using the random forest model and eight fatty liver prediction indicators of this invention are shown.
[0053] Figure 13 This is a schematic diagram of the structural principle of the intelligent prediction system for metabolic-related fatty liver disease of the present invention.
[0054] Figure 14 This is a schematic diagram of the electronic device of the present invention.
[0055] 401 - Multi-source data acquisition module, 402 - Variable screening module, 403 - Prediction result acquisition module. Detailed Implementation
[0056] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0057] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0058] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0059] Please see Figures 1-12 This invention provides an intelligent prediction method for metabolic-related fatty liver disease, comprising the following steps:
[0060] S100: Collects multi-source data, obtains basic demographic information, laboratory indicators, predictive indicators, and tongue images, and extracts quantitative and qualitative tongue image parameters.
[0061] In this embodiment, multi-source data is collected to obtain basic statistical information, experimental indicators, predictive indicators, and tongue images, and quantitative and qualitative tongue image parameters are extracted. The specific process is as follows:
[0062] S101: Collect basic information of subjects, experimental indicators, and fatty liver prediction indicators;
[0063] S102: Collect images of the subject's tongue, extract the color features of the tongue body and tongue coating, and identify cracks, punctures, teeth marks, etc., and output qualitative and quantitative tongue image parameter results.
[0064] Furthermore, in the steps of collecting images of the subject's tongue, extracting the color features of the tongue body and tongue coating, identifying cracks, punctures, teeth marks, etc., and outputting qualitative and quantitative tongue image parameters:
[0065] The original tongue images were quality assessed to ensure that they met the requirements in terms of brightness, sharpness, and target distance.
[0066] The tongue image region is cropped from the original tongue image and segmented into a tongue body image and a tongue coating image;
[0067] Extract color features of the tongue body and tongue coating;
[0068] Identify cracks on the tongue and calculate a crack score for the tongue body.
[0069] Identify tongue shape features and calculate tongue shape scores;
[0070] Identify tongue punctures and teeth marks, and calculate the scores for tongue punctures and teeth marks respectively;
[0071] Identify tongue coating texture features and calculate tongue coating stickiness score and tongue coating thickness score;
[0072] Output qualitative tongue image parameters, including tongue color, cracked tongue, punctate tongue, teeth-marked tongue, yellow coating, and thick coating.
[0073] During the above process, basic demographic information of the subjects was collected, including age, sex, pulse, systolic blood pressure (SBP), diastolic blood pressure (DBP), height, weight, waist circumference, hip circumference, smoking and drinking history.
[0074] Fasting blood samples were collected from the subjects in the morning, and the following tests were performed: ① Complete blood count indicators: red blood cell count (RBC), white blood cell count (WBC), neutrophil count (NC), lymphocyte count (LC), mean corpuscular volume (MCV), hemoglobin (HGB), platelet count (PLT); ② Liver function indicators: alanine aminotransferase (ALT), aspartate aminotransferase (AST), total bilirubin (TBIL), direct bilirubin (DBIL), albumin (ALB), alkaline phosphatase (ALP), gamma-glutamyl transferase (GGT); ③ Kidney function indicators: blood urea nitrogen (BUN), creatinine (CRE), and uric acid (UA); ④ Blood lipid indicators: total cholesterol (TC), triglycerides (TG), high-density lipoprotein cholesterol (HDL-C), and low-density lipoprotein cholesterol (LDL-C); ⑤ Blood glucose indicator: fasting blood glucose (FBG).
[0075] Based on the above indicators, eight predictive indicators related to fatty liver were further calculated, including FLI, ZJU, LAP (male), LAP (female), CMI, AIP, VAI (male), VAI (female), CVAI (male), CVAI (female), and TyG. In addition, other derived indicators are defined as follows: BMI = weight (kg) / height. 2 (m), obesity is defined as BMI > 28 kg / m 2 Elevated ALT and AST are defined as: ALT or AST > 40 IU / L; high UA is defined as: UA > 420 umol / L (male) or > 360 umol / L (female); high TC is defined as: TC ≥ 5.2 mmol / L; high TG is defined as: TG ≥ 1.7 mmol / L; low HDL-C is defined as: HDL-C < 1.04 mmol / L; high LDL-C is defined as: LDL-C ≥ 3.4 mmol / L.
[0076] Standardized acquisition of tongue images: Tongue images were collected from the subjects. The collection time was uniformly from 8:00 am to 11:00 am. All patients underwent the collection in an empty stomach and the tongue coating was not scraped or stained.
[0077] Intelligent recognition of tongue image parameters: The standardized tongue images are intelligently processed using the nahefa cloud system V2.0 to obtain quantitative and qualitative tongue image parameters.
[0078] First, the image quality of the original tongue image is evaluated from three aspects: brightness (whether it is overexposed or underexposed), sharpness (whether the image is clear or blurry), and target distance (whether the tongue is too far or too close). A qualified tongue image is obtained when all three aspects are met.
[0079] Then, the tongue image is automatically processed by the tongue detection and segmentation algorithms to crop out the tongue image area from the original tongue image photo and intelligently segment it into tongue body image and tongue coating image.
[0080] Next, based on the "LAB color space", the color features of the tongue and tongue coating are extracted. The image is converted from the RGB color space to the LAB color space, and its L, A, and B values are calculated to obtain the tongue color score -L (TBCL), tongue color score -A (TBCA), tongue color score -B (TBCB), tongue coating color score -L (TCCL), tongue coating color score -A (TCCA), and tongue coating color score -B (TCCB).
[0081] Crack segmentation models are used to identify cracks on the tongue. These models can determine whether there are cracks in a tongue image. If cracks are present, the corresponding crack areas are drawn and the tongue crack score (TCS) is automatically calculated.
[0082] The fine classification model is trained to determine tongue shape features and automatically calculate the tongue shape score (TSS1).
[0083] The detection model is trained to identify the presence or absence of tongue punctures and teeth marks, detect the coordinates and confidence (probability value or score) of punctures or teeth marks, and automatically calculate the tongue puncture score (TSS2) and tongue teeth mark score (TTMS).
[0084] Based on the classification model, the tongue coating texture features are identified, and the tongue coating stickiness score (TGCS) and tongue coating thickness score (TCTS) are automatically calculated.
[0085] Finally, the system further outputs qualitative tongue appearance parameters such as tongue color, cracked tongue, spotted tongue, toothmarked tongue, yellow coating, and thick coating. A schematic diagram of the qualitative identification results of tongue appearance parameters can be found below. Figure 3 .
[0086] S200: Screening of tongue appearance parameters and clinical indicators to identify key predictive variables.
[0087] In this implementation, tongue image parameters and clinical indicators are screened to identify key predictive variables. The specific process is as follows:
[0088] S201: Perform preliminary screening of tongue image parameters and clinical indicators, and select the tongue image parameters and clinical indicators with statistical significance, and output the first variable;
[0089] S202: Perform random forest classification on the results of the first variable of tongue image parameters and clinical indicators, draw feature importance bar charts, preliminarily screen variables according to feature importance ranking, and output the second variable;
[0090] S203: Perform LASSO regression analysis on the first variable results of tongue image parameters and clinical indicators, plot the coefficient profile and cross-validation curve of LASSO regression, screen out key variables, and output the third variable;
[0091] S204: Obtain the intersection of the second and third variables to determine the final key predictor variables.
[0092] In the above process, firstly, baseline data from the training and test sets are compared and analyzed. Categorical variables are represented by n (%), and the chi-square test is used for comparison; continuous variables that conform to a normal distribution are represented by mean ± standard deviation, and the t-test is used for comparison and analysis; those that do not conform to a normal distribution are represented by median [IQR], and the Mann-Whitney U test is used for comparison and analysis.
[0093] Then, logistic regression was used in the training set to screen for variables that influence the occurrence of MAFLD. Random forest and LASSO regression models were further used to screen model variables, and Venn diagrams were plotted to find their intersection. Finally, two tongue appearance parameters and six clinical indicators were determined as the final model variables. Nine machine learning models were then trained and subjected to 10-fold nested cross-validation, with Random Forest selected as the best predictive model.
[0094] Next, in the test set data, ROC curves, DCA curves, confusion matrices, and KS curves were used to verify the predictive performance of the best model for MAFLD. The predictive ability of the model with previously known fatty liver-related serum predictive indicators was compared using Delong test and IDI analysis. SHAP interpretability analysis, sensitivity analysis, and subgroup analysis of the best model were also performed, as well as correlation analysis with MAFLD-related indicators of liver inflammation, fibrosis, and glucose and lipid metabolism disorders. The predictive performance and value of the model were systematically and comprehensively analyzed.
[0095] The data analysis in this invention uses tools such as SPSS (version 25.0), R (version 3.6.2), MedCal (version 23.2.1), and Python.
[0096] Comparison of baseline data between training and test sets:
[0097] The baseline data included 18 tongue appearance parameters and 42 clinical indicators. There were no statistically significant differences between the indicators in the training set and the test set (all P values were > 0.05), indicating that the two datasets are representative and comparable.
[0098] Selection of model variables in the training set:
[0099] Logistic regression analysis was used to explore the impact of each indicator on the risk of MAFLD and a forest plot was drawn. Twelve tongue appearance parameters and 35 clinical indicators with statistical significance were initially screened out (all P values < 0.05).
[0100] Random forest classification was used to analyze the importance of variables for 12 tongue image parameters, and a feature importance bar chart was plotted. The top 10 variables, from most important to least important, were "TBCA", "TCCL", "TCCA", "TBCL", "TCCB", "TCTS", "TGCS", "Crackedtongue", "Tonguecolour", and "Spottedtongue". To prevent overfitting and avoid collinearity, LASSO regression analysis was also performed on the 12 tongue image parameters, and coefficient profiles and cross-validation curves were plotted to select the variables "Yellowcoating", "TBCA", and "TCCL". Then, the intersection of the variables selected by the random forest and LASSO regression methods was taken, and a Venn diagram was plotted.
[0101] Random forest classification was used to analyze the importance of variables for 35 clinical indicators, and characteristic importance bar charts were plotted. The results showed that the top 10 variables in terms of importance, from highest to lowest, were "Waist", "TG", "ALT", "HyperTG", "BMI", "GGT", "HDLC", "AST", "Age", and "AST / ALT". LASSO regression analysis was also performed, and coefficient profile plots of the LASSO regressions were plotted. Figure 8 E) and cross-validation curves ( Figure 8 F) Filter out the 7 variables: “ElevatedALT”, “HyperTG”, “Age”, “Waist”, “BMI”, “ALT”, and “TG”. Similarly, take the intersection of the variables filtered by the two methods and draw a Venn diagram.
[0102] In summary, the final model variables were selected from two tongue appearance parameters, “TBCA” and “TCCL”, and six clinical indicators, “Waist”, “BMI”, “ALT”, “TG”, “HyperTG”, and “Age”.
[0103] S300: Obtain key variable values based on key predictor variables, obtain the prediction results of the risk of metabolic-related fatty liver disease, and output them.
[0104] In this implementation, key variable values are obtained based on key predictor variables to obtain the predicted risk of metabolic-related fatty liver disease, and then the results are output. The specific process is as follows:
[0105] S301: Extract specific values for key predictor variables from the collected data;
[0106] S302: Calculate the probability of developing metabolic-related fatty liver disease based on specific numerical values;
[0107] S303: Outputs the predicted risk of developing metabolic-related fatty liver disease, including the probability value of the risk and the corresponding risk level.
[0108] In the above process, nine machine learning methods were used in the training set: Extreme Gradient Boosting (XGBoost), Logistic Regression (LR), Light Gradient Boosting Machine (LightGBM), RandomForest (RF), Gradient Boosting Decision Tree (GBDT), Support Vector Machine (SVM), k-Nearest Neighbors (kNN), Gaussian Naive Bayes (GNB), and Multilayer Perceptron (MLP) to construct a MAFLD classification model based on the eight variables selected above. The parameter values for each model are shown below:
[0109] (1) XGBoost Classifier
[0110] AUC = 0.9913623246931342;
[0111] Model parameters:
[0112] objective (optimization objective function): binary: logistic
[0113] learning_rate: None
[0114] max_depth (maximum tree depth): None
[0115] min_child_weight(minimum fork weight sum): None
[0116] reg_lambda(L2 regularization coefficient): None
[0117] (2) Logistic Regression
[0118] AUC = 0.8656996761486804;
[0119] Model parameters:
[0120] C (regularization factor): 1.0
[0121] max_iter (number of iterations): 100
[0122] penalty (regularization type): l2
[0123] TOL (convergence metric): 0.0001
[0124] (3) LightGBM Classifier
[0125] AUC = 0.9912228465722827;
[0126] Model parameters:
[0127] boosting_type (algorithm type): gbdt learning_rate (learning rate): 0.1 max_depth (maximum tree depth): -1 n_estimators (maximum number of trees): 100 num_leaves (maximum number of leaves): 31
[0128] (4) Random Forest Classifier
[0129] AUC = 0.9915705101915462;
[0130] Model parameters:
[0131] criterion (metric): gini
[0132] max_depth (maximum tree depth): None min_impurity_decrease (minimum branch purity gain): 0.0 n_estimators (number of trees): 20
[0133] (5) GBDT Classifier
[0134] AUC = 0.9756953064630812;
[0135] Model parameters:
[0136] learning_rate: 0.1 loss: log_loss max_depth: 3 min_samples_leaf: 1 min_samples_split: 2 n_estimators: 100
[0137] (6) SVM Classifier
[0138] AUC = 0.851099507668681;
[0139] Model parameters:
[0140] C (regularization factor): 1.0
[0141] kernel (kernel type): rbf
[0142] tol (convergence metric): 0.001
[0143] (7) KNN Classifier
[0144] AUC = 0.8919472669483517;
[0145] Model parameters:
[0146] n_neighbors (number of nearest neighbors): 5
[0147] weights (weight type): uniform
[0148] (8) GNB Classifier
[0149] AUC = 0.872985667992104;
[0150] Model parameters:
[0151] priors (prior probabilities): None var_smoothing (var_smoothing): 1e-09
[0152] (9) MLP Classifier
[0153] AUC = 0.6713087947698381;
[0154] Model parameters:
[0155] activation (nonlinear function): ReLU
[0156] hidden_layer_sizes(hidden layer width): (20, 10)
[0157] max_iter (number of iterations): 20
[0158] Among the nine models, Random Forest performed best on the training set (ranked by AUC), and its corresponding scores on the training set for each evaluation metric were as follows:
[0159] AUC (95% CI): 0.995 (0.990-0.999)
[0160] cutoff (95% CI): 0.5 (0.475-0.525)
[0161] Accuracy (95% CI): 0.982 (0.980-0.984)
[0162] Sensitivity (95% CI): 0.981 (0.975-0.987)
[0163] Specificity (95% CI): 0.983 (0.979-0.986)
[0164] Positive predictive value (95% CI): 0.978 (0.973-0.982)
[0165] Negative predictive value (95% CI): 0.985 (0.981-0.990)
[0166] F1 score (95% CI): 0.979 (0.977-0.982)
[0167] Kappa (95% CI): 0.963 (0.959-0.967)
[0168] In the validation set, the best performer was also Random Forest (ranked by AUC), with the following scores for each evaluation criterion:
[0169] AUC (95% CI): 0.915 (0.848-0.983)
[0170] cutoff (95% CI): 0.5 (0.475-0.525)
[0171] Accuracy (95% CI): 0.857 (0.831-0.882)
[0172] Sensitivity (95% CI): 0.811 (0.738-0.885)
[0173] Specificity (95% CI): 0.891 (0.868-0.914)
[0174] Positive predictive value (95% CI): 0.855 (0.833-0.876)
[0175] Negative predictive value (95% CI): 0.866 (0.824-0.909)
[0176] F1 score (95% CI): 0.827 (0.787-0.868)
[0177] Kappa (95% CI): 0.705 (0.650-0.761)
[0178] The two findings agree that Random Forest is the best model choice for this dataset.
[0179] The specific values of key predictive variables are extracted from the collected data and used as input values for the model. The key variable values are then fed into the trained model, which calculates the probability of developing metabolic-related fatty liver disease based on the specific values.
[0180] The model outputs a prediction of the risk of MAFLD, including the probability value of the risk and the corresponding risk level (such as low risk, medium risk, high risk).
[0181] Visualize the prediction results, such as using SHAP to indicate features that are helpful for decision-making, and help users understand the basis of the prediction results.
[0182] like Figure 10 and Figure 11 As shown, where: Figure 10 Comparison of nine MAFLD machine learning prediction models: A: ROC curves of each model's prediction of MAFLD in the training set; B: ROC curves after 10x nested cross-validation in the training set; C: Forest plot of the ROC results of each model's prediction of MAFLD, with the error bars in the plot representing the ROC mean and SD; D: Net return decision curves of each model's prediction of MAFLD during internal validation. Figure 11 Internal and external validation of the random forest model: A&B: ROC curves of the random forest model on the training set and internal validation set; C: ROC curve of the random forest model on the test set; D: KS analysis curve of the random forest model on the test set; E: Confusion matrix of the random forest model on the test set; F: Decision curve of the random forest model on the test set.
[0183] To further evaluate the predictive value of the random forest model for MAFLD, this invention compares it with eight previously reported predictive indicators related to fatty liver disease, including FLI, ZJU, LAP, CMI, AIP, VAI, CVAI, and TyG. Table 1 shows the predictive performance of the random forest model and the other eight indicators for MAFLD in the total sample, training set, and test set. The results show that the AUC values of the RF model in the total sample, training set, and test set are 0.968 (0.959-0.979), 0.984 (0.975-0.991), and 0.916 (0.883-0.949), respectively, all significantly higher than the AUC values of the other eight indicators (Delong test P < 0.001). We also plotted the ROC curves of the random forest model and these eight indicators for MAFLD prediction in each dataset. Figure 12 A, C, E) and Spearman correlation heatmap ( Figure 12 The results show that the random forest model has a strong positive correlation with all eight indicators, and its predictive ability for MAFLD is significantly better than that of these eight indicators.
[0184] Furthermore, we calculated the Integrated Discriminant Improvement (IDI) index between the random forest model and eight fatty liver prediction indicators to compare the predictive performance differences between the models. The IDI index is widely used in medical research for predicting disease risk and other problems. By comparing the performance of two predictive models on a new dataset, it assesses the degree of improvement in predictive performance of one model relative to the other, thus determining the discriminative ability of the two models. IDI values are typically between 0 and 1, where a value of 0 indicates that the two models have the same predictive performance, and a value closer to 1 indicates that one predictive model has a greater improvement in predictive performance relative to the other, meaning its discriminative ability is stronger. As shown in Table 1, in the test set, the IDI values of the Random Forest model compared to FLI, ZJU, LAP, CMI, AIP, VAI, CVAI, and TyG are 0.376 [0.327-0.425], 0.263 [0.216-0.309], 0.336 [0.290-0.383], 0.369 [0.319-0.420], 0.363 [0.311-0.415], and 0.393 [0.33], respectively. The values of 0.9-0.447, 0.317-0.271, and 0.517-0.464 all represent positive improvements, indicating that the Random Forest model, compared to FLI, ZJU, LAP, CMI, AIP, VAI, CVAI, and TyG, improved the prediction ability of MAFLD in the test set by 37.6%, 26.3%, 33.6%, 36.9%, 36.3%, 39.3%, 31.7%, and 51.7%, respectively. Similarly, in both the total sample and the training set, the Random Forest model's ability to predict MAFLD was significantly improved compared to the other eight indicators.
[0185]
[0186]
[0187] Table 1. Comparison of predictive performance of random forest model with 8 fatty liver prediction indicators
[0188] Note: * indicates P < 0.001 compared to the AUC value of the Random Forest model. # indicates the IDI value of the Random Forest model relative to other predictive indicators. Random Forest: Random Forest; FLI: Fatty Liver Index; LAP: Lipid Accumulation Product; CMI: Cardiometabolic Index; AIP: Plasma Atherosclerosis Index; VAI: Visceral Fat Index; CVAI: Chinese Visceral Fat Index; TyG: Triglyceride Glucose Index.
[0189] Corresponding to the aforementioned embodiments of the intelligent prediction method for metabolic-related fatty liver disease, this application also provides embodiments of an intelligent prediction system for metabolic-related fatty liver disease.
[0190] Figure 13 This is a block diagram illustrating an intelligent prediction system for metabolic-related fatty liver disease according to an exemplary embodiment. (Refer to...) Figure 13 The system may include: a multi-source data acquisition module 401, a variable filtering module 402, and a prediction result acquisition module 403; wherein:
[0191] The multi-source data acquisition module 401 is used to acquire multi-source data, obtain basic statistical information, experimental indicators, predictive indicators, and tongue images respectively, and extract quantitative and qualitative tongue image parameters.
[0192] The variable screening module 402 is used to screen tongue image parameters and clinical indicators to determine key predictive variables.
[0193] The prediction result acquisition module 403 is used to obtain key variable values based on key predictor variables, obtain prediction results of the risk of metabolic-related fatty liver disease, and output them.
[0194] In this embodiment, the multi-source data acquisition module 401 acquires multi-source data, obtaining basic statistical information, experimental indicators, predictive indicators, and tongue images, and extracts quantitative and qualitative tongue image parameters; the variable screening module 402 performs variable screening on the tongue image parameters and clinical indicators to determine key predictive variables; the prediction result acquisition module 403 obtains key variable values based on the key predictive variables, obtains the prediction result of the risk of metabolic-related fatty liver disease, and outputs it; through the above method, the effect of early prediction of metabolic-related fatty liver disease is achieved by fusing intelligent tongue image parameters and clinical indicators.
[0195] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0196] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0197] Accordingly, this application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the intelligent prediction method for metabolic-related fatty liver disease as described above. Figure 14 The diagram shown is a hardware structure diagram of any device with data processing capabilities, which is part of an intelligent prediction system for metabolic-related fatty liver disease provided in an embodiment of the present invention. (Except for...) Figure 14 In addition to the processor, memory, and network interface shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0198] Accordingly, this application also provides a storage medium storing computer instructions, which, when executed by a processor, implement the intelligent prediction method for metabolic-related fatty liver disease as described above. The storage medium can be an internal storage unit of any data-processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the storage medium can include both internal storage units of any data-processing device and external storage devices. The storage medium is used to store the computer program and other programs and data required by the data-processing device, and can also be used to temporarily store data that has been output or will be output.
[0199] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0200] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for intelligent prediction of metabolic-related fatty liver disease, characterized in that, Includes the following steps: Collect multi-source data to obtain basic demographic information, laboratory indicators, predictive indicators, and tongue images, and extract quantitative and qualitative tongue image parameters; Variable screening was conducted on tongue appearance parameters and clinical indicators to identify key predictive variables; Based on the key predictor variables, obtain the values of key variables, obtain the prediction results of the risk of developing metabolic-related fatty liver disease, and output them.
2. The intelligent prediction method for metabolic-related fatty liver disease as described in claim 1, characterized in that, In the steps of collecting multi-source data, obtaining basic demographic information, laboratory indicators, predictive indicators, and tongue images, and extracting quantitative and qualitative tongue image parameters: Collect basic information, experimental indicators, and fatty liver prediction indicators of the subjects; The system collects images of the subject's tongue, extracts the color features of the tongue body and tongue coating, and identifies conditions such as cracks, punctures, and teeth marks, outputting qualitative and quantitative tongue image parameters.
3. The intelligent prediction method for metabolic-related fatty liver disease as described in claim 2, characterized in that, In the steps of collecting tongue images of subjects, extracting color features of the tongue body and tongue coating, identifying cracks, punctures, teeth marks, etc., and outputting qualitative and quantitative tongue image parameter results: The tongue image region is cropped from the original tongue image and segmented into a tongue body image and a tongue coating image; Extract color features of the tongue body and tongue coating; Identify cracks on the tongue and calculate a crack score for the tongue body. Identify tongue shape features and calculate tongue shape scores; Identify tongue punctures and teeth marks, and calculate the scores for tongue punctures and teeth marks respectively; Identify tongue coating texture features and calculate tongue coating stickiness score and tongue coating thickness score; Output qualitative tongue image parameters, including tongue color, cracked tongue, punctate tongue, teeth-marked tongue, yellow coating, and thick coating.
4. The intelligent prediction method for metabolic-related fatty liver disease as described in claim 3, characterized in that, Before the steps of cropping the tongue image region from the original tongue image and segmenting it into tongue body images and tongue coating images: The original tongue images were quality-assessed to ensure that they met the requirements in terms of brightness, sharpness, and target distance.
5. The intelligent prediction method for metabolic-related fatty liver disease as described in claim 1, characterized in that, In the process of screening variables for tongue appearance parameters and clinical indicators to determine key predictive variables: Preliminary screening of tongue image parameters and clinical indicators was conducted, and statistically significant tongue image parameters and clinical indicators were selected and output as the first variable. Random forest classification was performed on the results of the first variable of tongue image parameters and clinical indicators, and bar charts of feature importance were drawn. Variables were initially screened according to the ranking of feature importance, and the second variable was output. LASSO regression analysis was performed on the first variable results of tongue appearance parameters and clinical indicators. The coefficient profile and cross-validation curve of LASSO regression were plotted to screen out key variables and output the third variable.
6. The intelligent prediction method for metabolic-related fatty liver disease as described in claim 5, characterized in that, After performing LASSO regression analysis on the first variables of tongue appearance parameters and clinical indicators, plotting the coefficient profile and cross-validation curve of the LASSO regression, screening out key variables, and outputting the third variable: Find the intersection of the second and third variables to determine the final key predictor variables.
7. The intelligent prediction method for metabolic-related fatty liver disease as described in claim 1, characterized in that, In the steps of obtaining key variable values based on key predictor variables, acquiring the predicted risk of metabolic-related fatty liver disease, and outputting the results: Extract specific values for key predictor variables from the collected data; The probability of developing metabolic-related fatty liver disease is calculated based on specific numerical values.
8. The intelligent prediction method for metabolic-related fatty liver disease as described in claim 7, characterized in that, After calculating the probability of developing metabolic-associated fatty liver disease based on specific numerical values: Output the predicted risk of developing metabolic-related fatty liver disease, including the probability value of the risk and the corresponding risk level.
9. A smart prediction system for metabolic-related fatty liver disease, applied to the smart prediction method for metabolic-related fatty liver disease as described in claim 1, characterized in that, It includes a multi-source data acquisition module, a variable filtering module, and a prediction result acquisition module; among which: The multi-source data acquisition module is used to collect multi-source data, obtain basic demographic information, laboratory indicators, predictive indicators, and tongue images, and extract quantitative and qualitative tongue image parameters. The variable screening module is used to screen tongue image parameters and clinical indicators to determine key predictive variables. The prediction result acquisition module is used to obtain the key variable values based on the key predictor variables, obtain the prediction results of the risk of metabolic-related fatty liver disease, and output them.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the intelligent prediction method for metabolic-related fatty liver disease as described in any one of claims 1 to 8.
Citation Information
Cited By
Reasonable plough layer evaluation index system construction method based on dual-objective optimization
CN121526093A