Non-alcoholic fatty liver risk prediction model and application thereof

By constructing a non-alcoholic fatty liver risk prediction model based on questionnaire and physical examination data, using LASSO regression, elastic network analysis and COX regression to screen variables, the problems of insufficient accuracy and convenience of existing diagnostic methods are solved, and efficient and accurate NAFLD risk prediction and early screening are achieved.

CN120412993APending Publication Date: 2025-08-01TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192578.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing diagnostic methods for non-alcoholic fatty liver (NAFLD) are costly, highly invasive, low patient acceptance, and insufficient accuracy in prediction models, making them difficult to be suitable for large-scale screening.

Method used

A non-alcoholic fatty liver risk prediction model was constructed. Through questionnaire surveys and physical examination data, combined with LASSO regression, elastic network analysis and COX regression, 7 independent predictor variables were screened out, and risk prediction was predicted using machine learning methods, and visualized through nomograms to simplify complex algorithms to improve readability.

Benefits of technology

It improves the accuracy and convenience of NAFLD risk prediction, is suitable for different populations and medical scenarios, supports early screening and intervention, reduces dependence on specialized computing devices and complex algorithms, and fills the technical gap in early prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412993A_ABST
    Figure CN120412993A_ABST
Patent Text Reader

Abstract

The invention discloses a non-alcoholic fatty liver risk prediction model and application thereof, and the risk prediction model comprises seven independent prediction variables, namely gender, body mass index, exercise, shift change, neutrophil count, triglyceride and high density lipoprotein. The risk prediction model is Y = 0.555 * sex + 0.254 * body mass index-0.354 * exercise + 0.495 * shift change + 0.120 * neutrophil count + 0.307 * triglycerides-0.872 * high density lipoprotein; wherein Y is the relative risk value of the non-alcoholic fatty liver disease; when the sex is' male ', the value is 1, and when the sex is' female ', the value is 0; when the motion is' YES ', the value is 1, and when the motion is' NO ', the value is 0; when the shift is' YES ', the value is 1, and when the shift is' NO ', the value is 0; the counting unit of the neutrophil is * 10 < 9 > / L; the unit of the triglyceride is mmol / L; the unit of the high-density lipoprotein is mmol / L. The risk prediction model is accurate and reliable, and an important tool is provided for NAFLD high-risk group screening, precise treatment and health management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biomedicine, and particularly relates to a non-alcoholic fatty liver risk prediction model and its application. Background Art

[0002] Non-alcoholic fatty liver disease (NAFLD) is a chronic disease characterized by macrovesicular steatosis of hepatocytes and has become one of the most common causes of chronic liver disease. Although NAFLD is considered to have a benign manifestation, it is associated with an increased risk of steatohepatitis, cirrhosis, or hepatocellular carcinoma. Currently, the global incidence of NAFLD is as high as 4,613 cases per 100,000 person-years, and the incidence of NAFLD in China is also as high as 56.7 cases per 1,000 person-years. More and more evidence shows that NAFLD is also closely related to the high incidence of type 2 diabetes, coronary heart disease, metabolic syndrome, and its components. With the high incidence of obesity and metabolic syndrome, NAFLD has become one of the important diseases seriously threatening public health in China. Some studies have shown that lifestyle changes or drug interventions can prevent the occurrence of NAFLD in high-risk populations. These intervention measures are particularly effective in the early stage of the disease. Therefore, early screening and intervention for NAFLD can reduce liver function damage and prevent or delay the complications of NAFLD.

[0003] However, current clinical diagnoses mainly rely on imaging examinations and liver biopsies, which have problems such as high cost, strong invasiveness, and low patient acceptance, and are not suitable for large-scale screening. In recent years, the research on biomarker-based prediction models in the early diagnosis of NAFLD has gradually increased. However, existing prediction models generally have problems such as insufficient accuracy, poor model interpretability, and inconvenient application. Summary of the Invention

[0004] In view of the problems existing in predicting the risk of non-alcoholic fatty liver disease (NAFLD), the present invention provides a non-alcoholic fatty liver risk prediction model with accurate and convenient characteristics and its application. The present invention constructs a new non-alcoholic fatty liver risk prediction model based on multi-dimensional data such as questionnaire surveys and physical examination data (laboratory test results and examination results) through prospective cohort studies, LASSO regression, elastic net analysis, and COX regression, and validates the risk prediction model through six machine learning methods, which has excellent prediction ability.

[0005] To achieve the above object, the present application adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a non-alcoholic fatty liver risk prediction model, and the risk prediction model includes 7 independent predictive variables, namely gender, body mass index, exercise, shift work, neutrophil count, triglyceride, and high-density lipoprotein.

[0007] In the above technical solution, the risk prediction model is Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein;

[0008] where: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L.

[0009] In the above technical solution, the risk prediction model is visualized using a nomogram.

[0010] In a second aspect, the present invention provides a non-alcoholic fatty liver risk prediction system, which includes a variable input module, an analysis module, and an output module;

[0011] The variable input module collects information on an individual's gender, body mass index, exercise, shift work, neutrophil count, triglyceride, and high-density lipoprotein;

[0012] The analysis module is used to calculate the relative risk value of non-alcoholic fatty liver and has a built-in non-alcoholic fatty liver risk prediction model, which is Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein;

[0013] where: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L;

[0014] The output module is used to output the relative risk value Y of non-alcoholic fatty liver obtained by the analysis module.

[0015] In the above technical solution, the analysis module further includes a non-alcoholic fatty liver risk prediction nomogram established based on the information collected by the variable input module.

[0016] In a third aspect, the present invention provides a non-alcoholic fatty liver risk prediction device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the following regression equation is run:

[0017] Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein;

[0018] Where: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L.

[0019] In a fourth aspect, the present invention provides a computer-readable storage medium, which includes a stored computer program. Wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the following regression equation:

[0020] Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein;

[0021] Where: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L.

[0022] Fifth aspect, the present invention provides an online tool for non-alcoholic fatty liver risk prediction model. The front end of the online tool uses HTML, CSS, and JavaScript to build and optimize the user interface framework to achieve dynamic interaction; the back end uses Flask of Python to create an API model logic, encapsulating the non-alcoholic fatty liver risk prediction model function into the API interface; the front end calls through HTTP requests, the user inputs variable values, and the back end calculates the non-alcoholic fatty liver relative risk value and returns it to the front end;

[0023] The non-alcoholic fatty liver risk prediction model function is: Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high density lipoprotein;

[0024] Where: Y is the non-alcoholic fatty liver relative risk value; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height^2, with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high density lipoprotein is the concentration of high density lipoprotein, with the unit of mmol / L.

[0025] Sixth aspect, the present invention provides the application of the above risk prediction model in the preparation of products for non-alcoholic fatty liver risk prediction.

[0026] The technical principle of the present invention:

[0027] First, a baseline of a cohort study was established using the physical examinations of the employees of a certain hospital in Wuhan in 2018. Through an average follow-up time of 4 years, it was divided into the Non-NAFLD group and the NAFLD group according to whether they had NAFLD during the physical examination in 2022. The t-test / Mann-Whitney U test / χ 2 test was used to screen out the variables with statistical differences between the two groups, and then LASSO regression and elastic net analysis were used to screen out the variables that had a key impact on the dependent variable. LASSO regression and elastic net introduce a penalty term in the model to compress the variable coefficients, which can not only effectively prevent overfitting but also alleviate the problem of multicollinearity, and can screen out the most relevant indicators from numerous variables to construct an accurate and efficient risk prediction model. In the present invention, LASSO regression and elastic net analysis compressed the 68 variables included to 38 and 45 respectively, and then took the intersection, and finally 37 potential predictive variables were screened out.

[0028] Subsequently, using the presence or absence of NAFLD as the dependent variable in COX regression, and the 37 variables jointly selected by LASSO regression and elastic net analysis as the independent variables, the independent variables with the greatest impact on the dependent variable were selected according to the absolute value of the regression coefficient. The larger the absolute value of the regression coefficient, the greater the impact of the independent variable on the dependent variable, and a P value less than 0.05 was considered to have significant statistical differences.

[0029] Furthermore, 7 predictive variables with relatively high absolute values of regression coefficients and P less than 0.05 were screened out through COX regression to construct the NAFLD risk prediction model Y: Y = 0.555 × gender + 0.254 × body mass index (kg / m2) - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count (*10^9 / L) + 0.307 × triglyceride (mmol / L) - 0.872 × high-density lipoprotein (mmol / L).

[0030] Among them, Y is the relative risk value; when gender is "male", exercise is "yes", and shift work is "yes", the value is 1, and when gender is "female", exercise is "no", and shift work is "no", the value is 0.

[0031] Finally, the risk prediction model was visually presented through a nomogram. The nomogram shows the magnitude of the impact of each key variable on the dependent variable. By assigning scores to each variable and then adding them up, the total score can be calculated. Further, using the mapping relationship between the total score and the probability of the occurrence of the dependent variable, the risk of individual outcomes was predicted. Using the nomogram can make the complex regression equation more intuitive and understandable, and improve the operability of the prediction model.

[0032] The beneficial effects of the present invention are as follows:

[0033] (1) A scientific and rigorous modeling process to ensure the robustness of the model

[0034] Based on a prospective cohort study, the present invention combines LASSO regression and elastic net to screen feature variables, ensuring the rationality and stability of model variable selection. At the same time, using Cox regression to construct the prediction model fully considers the impact of time factors on the risk of NAFLD occurrence, making the model closer to the actual situation.

[0035] (2) Multi-dimensional data integration to improve prediction accuracy

[0036] Compared with the prior art, the present invention combines questionnaire survey data with physical examination data (including laboratory test results and medical examination results), comprehensively considering multiple dimensions such as demographic information, living habits, metabolic indicators, and disease history. This multi-dimensional data integration method significantly improves the comprehensiveness and accuracy of NAFLD occurrence risk prediction.

[0037] (3) Diverse verification methods ensure the reliability of the model

[0038] Compared with the prior art, the present invention comprehensively verifies the model using six mainstream machine learning methods, namely random forest, logistic regression tree, decision tree, gradient boosting decision tree, extreme gradient boosting tree, and support vector machine. These methods verify the prediction ability of the model from different algorithm perspectives, significantly improving the generalization and credibility of the model.

[0039] (4) Simplify complex algorithms and enhance clinical applicability

[0040] The complex regression equation is visualized through a nomogram, making the results of the prediction model more readable and facilitating users to quickly understand and apply it in actual clinical practice. In addition, the simple operation of the nomogram is also convenient for popularization, reducing the dependence on specialized computing devices or knowledge of complex algorithms.

[0041] (5) Strong adaptability, suitable for multi-scenario applications

[0042] The prediction model of the present invention is constructed based on general data and is applicable to the risk prediction of NAFLD in different populations and medical scenarios. Whether it is health screening, physical examination assessment, or disease prevention and management, it can be efficiently applied, with wide applicability.

[0043] (6) Fill the technical gap and promote early disease prevention

[0044] Most of the prior art focuses on intervention programs for diagnosed patients, while there are few efficient prediction models for NAFLD risk. The proposed invention fills this gap in the field, providing a scientific basis for the early prevention and precise management of NAFLD, and helping to reduce the incidence and burden of this disease. Description of the Drawings

[0045] Figure 1 It is the flow chart for constructing the NAFLD prediction model of the present invention.

[0046] Figure 2 It is the screening process of LASSO regression and elastic net analysis of the present invention; wherein A: LASSO coefficient path diagram; B: LASSO cross-validation curve; C: elastic net coefficient path diagram; D: elastic net cross-validation curve; E: Venn diagram of LASSO regression analysis and elastic net analysis.

[0047] Figure 3 It is the COX regression forest plot of the NAFLD prediction model of the present invention.

[0048] Figure 4 It is the nomogram of the NAFLD prediction model of the present invention.

[0049] Figure 5Evaluation of the NAFLD prediction model of the present invention; wherein A: ROC curve; B: calibration curve.

[0050] Figure 6 The partial dependence plot and variable importance ranking plot of the survival model of the NAFLD prediction model of the present invention; wherein A: partial dependence plot of the survival model of gender; B: partial dependence plot of the survival model of exercise; C: partial dependence plot of the survival model of shift work; D: partial dependence plot of the survival model of triglyceride; E: partial dependence plot of the survival model of high-density lipoprotein; F: partial dependence plot of the survival model of body mass index; G: partial dependence plot of the survival model of neutrophil count; H: variable importance ranking plot.

[0051] Figure 7 Verification of six machine learning methods of the present invention for constructing the NAFLD prediction model; wherein A: ROC curves of six machine learning methods; B: precision-recall curves of six machine learning methods; C: DCA curves of six machine learning methods. Detailed implementation manners

[0052] To better illustrate the purpose, technical solution and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments. The present invention can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the present invention to those skilled in the art. The present invention will only be defined by the claims.

[0053] The research idea of the present invention is as follows: First, a prospective cohort study is established based on the physical examination population of employees in a hospital in Wuhan; then the final research population is determined according to strict exclusion and inclusion criteria; then LASSO regression, elastic net analysis and COX regression are used to gradually screen out key variables, and a new non-alcoholic fatty liver risk prediction model is constructed; finally, the prediction model is verified by six machine learning methods.

[0054] (1) The 37 potential predictive variables screened by LASSO regression and elastic net analysis are age, gender, education level, exercise, shift work, sleep duration, having breakfast every day, history of diabetes, history of hypertension, body mass index, diastolic blood pressure, neutrophil count, monocyte count, eosinophil percentage, basophil percentage, red blood cell count, hemoglobin, mean hemoglobin concentration, RBC distribution width SD, platelet count, large platelet ratio, triglyceride, low-density lipoprotein, high-density lipoprotein, alanine aminotransferase, gamma-glutamyl transpeptidase, alkaline phosphatase, aspartate aminotransferase, globulin, direct bilirubin, glucose, creatinine, carcinoembryonic antigen, uric acid, pH, urine red blood cell count, urine white blood cell count.

[0055] (2) Seven key variables determined by COX regression: body mass index, triglyceride, gender, high-density lipoprotein, neutrophil count, shift work, and exercise.

[0056] (3) Prediction model utility tool: visual nomogram.

[0057] (4) Validation results of six machine learning methods for the prediction model.

[0058] Example 1: Construction of a NAFLD risk prediction model based on a prospective cohort study of the physical examination of employees in a hospital in Wuhan

[0059] S1. In 2018, the sociodemographic characteristics, lifestyle, and disease history of the employees undergoing physical examinations in a hospital in Wuhan, as well as the employees' physical examination data (including laboratory test data, physical examination, and imaging examination data), were collected to form the first dataset. In this example, 8,994 employees who participated in the physical examination in 2018 were used as the baseline of the cohort study to construct the first dataset.

[0060] S2. In this example, by applying the exclusion and inclusion criteria to the first dataset, 5,644 study subjects were finally used as the second dataset of this example.

[0061] Population inclusion criteria: Employees undergoing physical examinations in a hospital in Wuhan in 2018

[0062] Population exclusion criteria: 1. Exclude those with NAFLD at the 2018 baseline and those with missing NAFLD diagnoses; 2. Exclude those who drink excessively (men drink no less than 210 grams per week and women drink no less than 140 grams per week); 3. Exclude those with hepatitis; 4. Exclude those with tumors; 5. Exclude those with missing values.

[0063] Definition of NAFLD: According to the imaging diagnosis criteria proposed by the National Symposium on Fatty Liver and Alcoholic Liver Disease, if a participant is confirmed to have hepatic steatosis by imaging and has no history of excessive alcohol consumption in the past 12 months, they are diagnosed with NAFLD.

[0064] Data collection:

[0065] Questionnaire data: Age, gender, smoking, second-hand smoke exposure, alcohol consumption, education level, exercise, shift work, history of diabetes, history of hypertension, history of coronary heart disease, history of stroke, sleep duration, lunch break duration, having breakfast every day, sedentary time per day;

[0066] Physical examination data: body mass index, heart rate, systolic blood pressure, diastolic blood pressure, glomerular filtration rate, white blood cell count, neutrophil count, neutrophil percentage, lymphocyte count, lymphocyte percentage, monocyte count, monocyte percentage, eosinophil count, eosinophil percentage, basophil count, basophil percentage, red blood cell count, hemoglobin, hematocrit, mean RBC volume, mean hemoglobin content, mean hemoglobin concentration, RBC distribution width CV, RBC distribution width SD, platelet count, PLT distribution width, mean PLT volume, large platelet ratio, platelet hematocrit, total cholesterol, triglyceride, low density lipoprotein, high density lipoprotein, alanine aminotransferase, γ-glutamyl transpeptidase, alkaline phosphatase, aspartate aminotransferase, total protein, globulin, albumin, total bilirubin, direct bilirubin, indirect bilirubin, glucose, creatinine, urea, carcinoembryonic antigen, uric acid, pH, cast count, urinary red blood cell count, urinary white blood cell count.

[0067] S3. Exclude 112 subjects lost to follow-up in the 2022 physical examination, and finally include 5532 subjects in the study to establish the third sample set. After an average of 4 years of follow-up, 4919 of the 5532 subjects did not develop non-alcoholic fatty liver (Non-NAFLD group), and 613 developed into NAFLD (NAFLD group). For the 5532 subjects, 68 research variables were included in the subsequent analysis, as shown in Table 1.

[0068] S4. Complete the differential analysis of 68 research variables in the Non-NAFLD group and the NAFLD group in the third sample set through R4.4.2; screen out the variables with differences between the two groups as the variables to be analyzed, and form the fourth sample set with all the variables to be analyzed in the Non-NAFLD group and the NAFLD group.

[0069] In this embodiment, different statistical methods are selected for different variable types. For continuous variables in the third sample set, the Anderson-Darling test is first used for normality test, and P less than 0.05 is considered not to conform to the normal distribution. If the continuous variable conforms to the normal distribution, it is described in the form of x±s, and the independent sample t-test is used for component difference analysis; if the continuous variable does not conform to the normal distribution, it is expressed by the median (M) and the upper and lower quartiles (Q1, Q3), and the Mann-Whitney U test is used to analyze the inter-group differences, and P less than 0.05 is considered to have statistical differences. All classifications in the third sample set in step S3 are expressed in cases (percentages), and χ 2 test is used for inter-group difference analysis, and P less than 0.05 is considered that the difference is statistically significant.

[0070] Table 1. Variables included in the study and their differential analysis

[0071]

[0072]

[0073]

[0074]

[0075] S5. Use LASSO regression and elastic net analysis to analyze the fourth sample set, respectively screen out multiple potential predictive variables, take the intersection through a Venn diagram, and finally screen out the common potential predictive variables. After screening, the fifth sample set is composed of all potential predictive variables of the Non-NAFLD group and the NAFLD group.

[0076] In this embodiment, LASSO regression and elastic net analysis respectively screen out 38 ( Figure 2 A and B in Figure 2 and 45 ( Figure 2 C and D in

[0077] potential predictive variables. Take the common intersection through a Venn diagram, and finally screen out 37 (

[0078] S6. Use the COX regression model to analyze the potential predictive variables of the fifth sample set, and screen out the independent influencing variables for the risk of NAFLD occurrence; the sixth sample set is composed of independent predictive variables in the Non-NAFLD group and the NAFLD group. Further conduct model parameter statistics on the predictive variables of the sixth sample set to construct a predictive model, and at the same time construct a risk prediction model.

[0079] In this example, 7 independent predictive variables were selected through the COX regression model, including gender (HR value: 1.743, 95% CI: 1.443 - 2.105, P < 0.001), body mass index (HR value: 1.289, 95% CI: 1.255 - 1.324, P < 0.001), exercise (HR value: 0.702, 95% CI: 0.596 - 0.827, P < 0.001), shift work (HR value: 1.641, 95% CI: 1.384 - 1.945, P < 0.001), neutrophil count (HR value: 1.128, 95% CI: 1.065 - 1.195, P < 0.001), triglyceride (HR value: 1.359, 95% CI: 1.258 - 1.469, P < 0.001), and high-density lipoprotein (HR value: 0.418, 95% CI: 0.304 - 0.575, P < 0.001). For details, see the COX regression forest plot ( Figure 3 ).

[0080] NAFLD risk prediction model Y: Y = 0.555 × gender + 0.254 × body mass index (kg / m 2 ) - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count (*10^9 / L) + 0.307 × triglyceride (mmol / L) - 0.872 × high-density lipoprotein (mmol / L).

[0081] Where Y is the relative risk value; when gender is "male", exercise is "yes", and shift work is "yes", the value is 1; when gender is "female", exercise is "no", and shift work is "no", the value is 0.

[0082] The NAFLD risk prediction model has been visualized using a nomogram ( Figure 4 ). In this example, the survival package in R 4.4.2 software was used to complete the COX regression analysis, forestplot was used to draw the forest plot, and the rms package was used to draw the nomogram.

[0083] Table 2. 7 independent predictive variables selected by COX regression

[0084] Variable name HR value 95% confidence interval P value Gender Female reference reference - Male 1.743 1.443-2.105 <0.001 <![CDATA[Body mass index, kg / m 2 > 1.289 1.255-1.324 <0.001 Exercise No reference reference - Yes 0.702 0.596-0.827 <0.001 Shift work No reference reference - Yes 1.641 1.384-1.945 <0.001 Neutrophil count, *10^9 / L 1.128 1.065-1.195 <0.001 Triglyceride, mmol / L 1.359 1.258-1.469 <0.001 High density lipoprotein, mmol / L 0.418 0.304-0.575 <0.001

[0085] S7. The predictive ability and accuracy of the prediction model were evaluated using the ROC curve and calibration curve. In this example, the area under the ROC curve was 0.821 ( Figure 5 A), indicating that the prediction model has good predictive ability. The calibration curve ( Figure 5B) Along the shape of y = x, and the mean squared error value is 0.06, indicating that the risk prediction model can be largely consistent with the actual observed values. In this embodiment, the pROC package in the R 4.4.2 software is used for ROC curve plotting. The Bootstrap method is used for internal validation of the prediction model; after resampling 1000 times, the calibration curve is used for evaluation; rms is used to plot the calibration curve.

[0086] S8. Use the partial dependence plots ( Figure 6 A - F) and feature importance plots ( Figure 6 H) of the survival model to visualize the contribution degrees and importance of the 7 independent predictive variables. In this embodiment, from the partial dependence plots of the survival model, it can be seen that male, lack of exercise, shift work, elevated triglycerides, decreased high - density lipoprotein, increased body mass index, and increased neutrophil count will significantly increase the risk of NAFLD. In addition, the feature importance plot shows that the contribution degrees to the risk of NAFLD from large to small are body mass index, triglyceride, gender, high - density lipoprotein, neutrophil count, shift work, and exercise. In this embodiment, the DALEX, ranger, and survival packages are used to plot the partial dependence plots and feature importance plots of the survival model.

[0087] S9. Use 6 machine learning methods such as random forest, logistic regression tree, decision tree, gradient - boosting decision tree, extreme gradient - boosting decision tree, and support vector machine tree to verify the model.

[0088] In this embodiment, the random number 123 is used to split the training set and the test set, and the splitting ratio is 8:2. The prediction model is repeatedly verified by 6 machine learning methods. The verification results of multiple models show that the prediction model has good sensitivity and specificity ( Figure 7 A), high precision ( Figure 7 B), and good net benefit ( Figure 7 C). Overall, the verification results of multiple machine learning models show that the NAFLD risk prediction model has excellent prediction ability.

[0089] The present invention provides an important tool for the screening, precise treatment, and health management of high - risk populations of NAFLD.

[0090] The application of the risk prediction model of the present invention is extensive. Especially in the aspects of early prevention, precise treatment, and health management of NAFLD, it has important clinical and public health significance. The specific applications are as follows:

[0091] (1) Early screening of high-risk populations: The prediction model can quickly calculate the risk value of developing NAFLD based on individual independent risk variables, thus helping to efficiently screen high-risk populations and detect potential patients at an early stage. In addition, the prediction model can optimize screening resources, reduce overexamination of low-risk populations, and concentrate limited medical resources on high-risk populations, improving screening efficiency and cost-effectiveness.

[0092] (2) Assisting clinical decision-making: The NAFLD prediction model can serve as an auxiliary tool for clinicians to help them formulate more scientific diagnosis and treatment plans and conduct hierarchical management. High-risk patients can be arranged for more frequent follow-up and examinations, while low-risk patients can reduce unnecessary medical interventions.

[0093] (3) Guiding health interventions and lifestyle: The prediction model can not only identify high-risk individuals but also guide patients in targeted lifestyle interventions and evaluate the effectiveness of the interventions. Doctors can customize intervention measures for patients, such as adjusting the diet structure, increasing physical exercise, and controlling body weight. The prediction model can also serve as a dynamic tool to judge the effectiveness of intervention measures by regularly evaluating the changes in the risk values of patients, thereby optimizing management strategies.

[0094] (4) Assisting in disease prevention and public health policy formulation: From the perspective of public health, the prediction model is of guiding significance for the prevention and treatment of NAFLD. The prediction model can be used for risk censuses of large populations to identify key populations for prevention and control. It provides a basis for the government and medical institutions to formulate intervention measures and policies, such as designing healthy diet plans or exercise promotion programs.

[0095] (5) As an education tool for health: The risk scores and incidence probabilities of the prediction model can serve as auxiliary tools for health education, providing an understanding of the risk levels and risk factors of NAFLD, thereby enhancing the awareness of non-alcoholic fatty liver.

[0096] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A non-alcoholic fatty liver risk prediction model, characterized in that: The risk prediction model includes 7 independent prediction variables, namely gender, body mass index, exercise, shift work, neutrophil count, triglyceride, and high-density lipoprotein.

2. The risk prediction model according to claim 1, wherein: The risk prediction model is Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein; Wherein: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L.

3. The risk prediction model according to claim 1, wherein: The risk prediction model is visualized using a nomogram.

4. A non-alcoholic fatty liver risk prediction system, characterized in that: The risk prediction system includes a variable input module, an analysis module, and an output module; The variable input module collects information on an individual's gender, body mass index, exercise, shift work, neutrophil count, triglyceride, and high-density lipoprotein; The analysis module is used to calculate the relative risk value of non-alcoholic fatty liver. It has a built-in non-alcoholic fatty liver risk prediction model, and the risk prediction model is Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein; Wherein: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when the exercise is "yes", the value is 1, and when the exercise is "no", the value is 0; when the shift work is "yes", the value is 1, and when the shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L; The output module is used to output the relative risk value Y of non-alcoholic fatty liver obtained by the analysis module.

5. The risk prediction system according to claim 4, wherein: The analysis module also includes a non-alcoholic fatty liver risk prediction nomogram established based on the information collected by the variable input module.

6. A non-alcoholic fatty liver risk prediction device, characterized in that: The risk prediction device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it runs the following regression equation: Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein; Wherein: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the following regression equation: Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein; Wherein: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L.

8. An online tool for a non-alcoholic fatty liver risk prediction model, characterized in that: The front end of the online tool uses HTML, CSS, and JavaScript to build and optimize the user interface framework to achieve dynamic interaction; the back end uses Flask of Python to create an API model logic and encapsulates the non-alcoholic fatty liver risk prediction model function into an API interface; the front end calls through an HTTP request. The user inputs variable values, and the back end calculates the relative risk value of non-alcoholic fatty liver and returns it to the front end; The non-alcoholic fatty liver risk prediction model function is: Y = 0.555 × gender + 0.254 × body mass index - 0.354 × exercise + 0.495 × shift work + 0.120 × neutrophil count + 0.307 × triglyceride - 0.872 × high-density lipoprotein; Wherein: Y is the relative risk value of non-alcoholic fatty liver; when the gender is "male", the value is 1, and when the gender is "female", the value is 0; body mass index = weight / height², with the unit of kg / m 2 ; when exercise is "yes", the value is 1, and when exercise is "no", the value is 0; when shift work is "yes", the value is 1, and when shift work is "no", the value is 0; the unit of neutrophil count is *10^9 / L; triglyceride is the concentration of triglyceride, with the unit of mmol / L; high-density lipoprotein is the concentration of high-density lipoprotein, with the unit of mmol / L.

9. Use of the risk prediction model according to any one of claims 1-3 in the preparation of a product for predicting the risk of non-alcoholic fatty liver disease.