Diabetes risk prediction method and device based on naive Bayesian algorithm, medium and product
Through the diabetes risk prediction method based on the Naive Bayes algorithm, a diabetes risk prediction model is constructed using lifestyle data and postprandial blood sugar change data, which solves the problems of insufficient inclusion of factors and low accuracy of the scoring model in the existing technology, and quantitative and accurate prediction of future diabetes risks are achieved.
Patent Information
- Application Number
- CN202510131308.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems such as insufficient inclusion of factors, low accuracy of scoring models and limited promotion in the prediction of diabetes risk.
A diabetes risk prediction method based on Naive Bayes algorithm was adopted. By obtaining living habit data, age data and postprandial blood sugar change data, a similar group was constructed and a one-variable linear regression model was established to obtain similar group labels, and then a diabetes risk prediction model was constructed using Naive Bayes method.
Quantitative and accurate prediction of future diabetes risks is achieved, and the cost of follow-up and data collection is reduced by correlating the changes in life habits and blood sugar with age.
Smart Images

Figure CN120072297A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of risk prediction, and particularly to a diabetes risk prediction method, device, medium and product based on the Naive Bayes algorithm. Background Art
[0002] Diabetes Mellitus (DM) is a metabolic disease characterized by hyperglycemia due to insulin secretion deficiency or insulin action disorder. Its characteristics are chronic hyperglycemia, accompanied by insufficient insulin secretion or action disorder, leading to disorders in carbohydrate, fat, and protein metabolism, causing chronic damage, dysfunction, and failure of multiple organs. Currently, the global prevalence of diabetes has affected more than 440 million people. Among them, the number of diabetes patients in the Asia-Pacific region is the largest, and the prevalence has increased sharply in recent decades. China has a population of 1.38 billion, and about 110 million people suffer from diabetes, becoming the country with the largest number of diabetes patients in the world, and this number is still increasing, bringing a huge burden to the medical system. In a previous study involving 170,287 participants, the prevalence of diabetes was 10.9%, and 60% of them did not know that they were diagnosed with diabetes. In addition, another 35.7% of the population was found to have abnormal blood glucose homeostasis, highlighting a large number of people at risk of diabetes. The main reasons for missed diagnosis include the lack of disease management awareness among patients themselves. In addition, the test results often rely on fasting blood glucose measurement, and this result is not stable. The ideal method is to check both fasting blood glucose and the blood glucose value 2 hours after an oral glucose tolerance test (OGTT). However, in areas with relatively low medical levels, blood glucose detection still needs to be further popularized.
[0003] Considering that the onset of diabetes has a certain degree of concealment in the early stage, it is possible to effectively prevent the occurrence and development of diabetes by detecting diabetes risk factors as early as possible and intervening in the risk factors at an early stage. According to the data of the World Health Organization (WHO), the main risk factors for type II diabetes include obesity, hypertension, smoking habits, abnormal lipid metabolism, family history, and low intake of fruits and vegetables. Recent studies have shown that diabetes and its complications (eye, cardiovascular, and systemic complications) can be prevented through a healthy diet, regular physical activity, and control of blood glucose, blood pressure, and cholesterol. In addition, through the precise screening of early diabetes patients, the risk rating of diabetes onset can be carried out and individualized prevention guidance can be formulated, thereby effectively reducing the occurrence and development of diabetes and the disease burden it causes to the whole society.
[0004] In recent years, with the rapid development of artificial intelligence technology, in-depth analysis of health big data has gradually received increasing attention. Currently, the basic data of health care is being gradually accumulated and integrated, and upper-layer analysis and application research based on health big data are gradually being carried out. In addition, the prevention and control of chronic diseases, including diabetes prevention and control, have also begun to be closely integrated with health big data and related technologies. Against this background, research on the initial screening of diabetes and personalized diabetes prevention guidance based on health big data is highly necessary and realistically feasible.
[0005] In recent years, a series of studies have explored diabetes risk screening, prediction, and individualized prevention guidance. Lindstrom et al. conducted a study to explore a non-laboratory test method to identify the risk of an individual having type 2 diabetes. They constructed a diabetes scoring model called FINDRISC (Finnish Diabetes Risk Score Model), and the main variables included age, BMI index, waist circumference, use of antihypertensive drugs, and history of hyperglycemia. The verification results showed that this model is a simple, rapid, economical, non-invasive, and reliable tool that can be used to identify high-risk populations of type 2 diabetes. Saafisto et al. studied the performance of FINDRISC and made improvements, and verified the effectiveness of the improved model. The research results showed that this improved model can be used as a self-management test to screen high-risk subjects of type 2 diabetes, and can also be used in the general population and clinical practice to identify undetected type 2 diabetes patients. In addition, the study combined the main risk factors of the population, increased the sample size of the study, provided a stable and reliable model, and developed a scoring system for the risk of type 2 diabetes onset. On this basis, European scholars further improved the constructed model in combination with population characteristics, better promoting the application value of the model. Li Yanyun et al. in China constructed a diabetes risk screening model based on the diabetes epidemiology survey of the Shanghai community population, and the main factors included age, cardiovascular disease, family history, systolic blood pressure, waist-hip ratio, and BMI index. Mi Wei et al. also used the same method to construct a diabetes risk screening model, and lifestyle information was introduced into this model. In addition, some studies have evaluated the application effect of foreign mainstream models in the Chinese population and compared the prediction effects of the models. However, the construction of the above models is based on traditional statistical methods and has the following limitations: (1) fewer factors are included in the analysis, and a large amount of available routine physical examination information is not included in the model; (2) the construction of the scoring model is based on the traditional logistic regression model, with low accuracy; (3) the generalization of the research results is limited by the sample size.
[0006] Therefore, based on the above problems, there is an urgent need to provide a new diabetes risk prediction method or system. Summary of the Invention
[0007] The purpose of this application is to provide a diabetes risk prediction method, device, medium and product based on the Naive Bayes algorithm, which can realize quantitative and accurate prediction of future diabetes risks.
[0008] To achieve the above object, this application provides the following solutions:
[0009] In the first aspect, this application provides a diabetes risk prediction method based on the Naive Bayes algorithm. The diabetes risk prediction method based on the Naive Bayes algorithm includes:
[0010] Obtain sample data; the sample data includes: lifestyle data, age data, and the change data of postprandial blood glucose;
[0011] Determine lifestyle variables according to the lifestyle data; and divide the sample data with the same lifestyle variables into the same similarity group;
[0012] For the age data and the change data of postprandial blood glucose in the same similarity group, establish a univariate linear regression model and obtain a similarity group label; the similarity group label is used to judge whether the change data of postprandial blood glucose rises with the increase of age data;
[0013] Using the lifestyle variables in the similarity group as the input and the similarity group label as the output, adopt the Naive Bayes method to obtain a diabetes risk prediction model;
[0014] Obtain the lifestyle variables of the person to be predicted, and use the diabetes risk prediction model to predict the diabetes risk.
[0015] Optionally, after obtaining the sample data, it further includes:
[0016] Perform missing value processing and outlier processing on the sample data.
[0017] Optionally, the determining the lifestyle variables according to the lifestyle data specifically includes:
[0018] Determine candidate lifestyle variables according to the lifestyle data;
[0019] Adopt the chi-square test method or the t-test method for data dimensionality reduction to obtain lifestyle variables with high correlation with diabetes risk.
[0020] Optionally, after dividing the sample data with the same lifestyle variables into the same similarity group, it further includes:
[0021] Perform preprocessing on the similarity group; the preprocessing includes: similarity group screening.
[0022] Optionally, the method for establishing the univariate linear regression model is the method of linear regression.
[0023] In a second aspect, the present application provides a diabetes risk prediction device based on the Naive Bayes algorithm. The diabetes risk prediction device based on the Naive Bayes algorithm includes:
[0024] A data acquisition module for acquiring sample data; the sample data includes: lifestyle data, age data, and the change data of postprandial blood glucose;
[0025] A similar group construction module for determining lifestyle variables according to the lifestyle data; and dividing the sample data with the same lifestyle variables into the same similar group;
[0026] A similar group label determination module for establishing a unary linear regression model for the age data and the change data of postprandial blood glucose in the same similar group, and obtaining a similar group label; the similar group label is used to determine whether the change data of postprandial blood glucose rises with the increase of age data;
[0027] A diabetes risk prediction model determination module for using the lifestyle variables in the similar group as the input and the similar group label as the output, and adopting the Naive Bayes method to obtain a diabetes risk prediction model;
[0028] A risk prediction module for obtaining the lifestyle variables of the person to be predicted and performing diabetes risk prediction by using the diabetes risk prediction model.
[0029] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the diabetes risk prediction method based on the Naive Bayes algorithm.
[0030] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the diabetes risk prediction method based on the Naive Bayes algorithm.
[0031] In a fifth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the diabetes risk prediction method based on the Naive Bayes algorithm.
[0032] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0033] The present application provides a diabetes risk prediction method, device, medium and product based on the Naive Bayes algorithm. Aiming at quantitatively predicting the future diabetes risk around the lifestyle data of the subjects, a diabetes risk prediction model based on the Naive Bayes algorithm is constructed. The diabetes risk prediction model outputs the probability of a significant increase in future blood glucose according to the lifestyle data of the subjects, as a measure of the future diabetes risk; the data is reconstructed by a method of constructing similar groups according to the sample data with the same lifestyle variables. Each similar group consists of all individuals with the same values on each lifestyle variable. Since there are differences in the postprandial blood glucose and age among individuals within the similar group, the change of postprandial blood glucose corresponding to the change of individual age within the similar group can be examined, so as to establish the association between the lifestyle and the change of blood glucose with age. And the change of blood glucose with age can reflect the future diabetes risk; the present application can realize the quantitative prediction of its future diabetes risk, and construct a prediction model of future diabetes risk through cross-sectional data, greatly saving the costs of follow-up and data collection. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0035] Figure 1 It is an application environment diagram of a diabetes risk prediction method based on the Naive Bayes algorithm in an embodiment of the present application;
[0036] Figure 2 It is a schematic flowchart of a diabetes risk prediction method based on the Naive Bayes algorithm provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0038] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0039] The diabetes risk prediction method based on the Naive Bayes algorithm provided in the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send the lifestyle variables of the person to be predicted to the server 104. After receiving the lifestyle variables of the person to be predicted, for the lifestyle variables of the person to be predicted, the server 104 performs diabetes risk prediction based on the diabetes risk prediction model. The server 104 can feedback the obtained diabetes risk prediction result to the terminal 102. In addition, in some embodiments, the diabetes risk prediction method based on the Naive Bayes algorithm can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly predict the lifestyle variables of the person to be predicted, or the server 104 can obtain the lifestyle variables of the person to be predicted from the data storage system and predict the lifestyle variables of the person to be predicted.
[0040] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0041] In an exemplary embodiment, as Figure 2 shown, a diabetes risk prediction method based on the Naive Bayes algorithm is provided. This method is executed by a computer device, and can be specifically executed separately by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in it as an example for description, it includes the following S201 to S205. Among them:
[0042] S201, obtaining sample data; the sample data includes: lifestyle data, age data, and the change data of postprandial blood glucose;
[0043] According to the inclusion and exclusion criteria, the underage population is removed, and only the population aged 18 or above is retained. After screening, 18,482 samples are obtained. In addition, patients who already have diabetes need to be excluded, so the population with fasting blood glucose exceeding 7 or glycated hemoglobin exceeding 6.5 needs to be removed. After screening again, 17,758 samples are obtained.
[0044] After S201, it further includes:
[0045] Perform missing value handling and outlier handling on the sample data.
[0046] Statistically analyze the missing situation of each variable in the dataset. The results show that the variable with the most serious missing situation is HbA1c, and the missing rate reaches 36%. Considering that this variable of glycated hemoglobin can play a great role in the screening and risk prediction of diabetes. If this variable is removed or the missing values of this variable are filled in some way, it may have a great impact on the subsequent modeling analysis. Therefore, the glycated hemoglobin variable needs to be retained as much as possible according to the original situation. So, all samples with missing glycated hemoglobin are removed. After further data research, it is found that if the glycated hemoglobin variable is not missing, then the number of missing values of the remaining variables is very small. For the convenience of subsequent modeling, for the remaining samples, any sample with a missing value is removed. Therefore, the actual missing value handling is to remove all samples with missing values. After the missing value handling, the total number of samples drops from 17758 to 11428, a decrease of about 36% (mainly caused by deleting samples with missing glycated hemoglobin).
[0047] Abnormal data due to data acquisition, recording, or operation needs to be verified and processed. In this study, only logical verification is carried out. For each variable, combined with the reasonable range of the variable, extreme abnormal data is located. It mainly includes: First, the number of days of high-intensity physical activity per week exceeds 7 days; Second, the body mass index exceeds 40, which is significantly higher than the normal level; Third, the hip circumference is less than 60 cm, which is significantly lower than the normal level; Fourth, the waist circumference is less than 50 cm or higher than 140 cm; Fifth, the heart rate exceeds 50; Sixth, the systolic blood pressure exceeds 1000 mmHg. Statistically analyze the total number of samples with any of the above data abnormal conditions, and the total number is 46, accounting for 0.4% of the current remaining total number of samples, and the proportion is extremely small. Delete the outlier data, and finally 11382 samples remain.
[0048] After data processing, first perform descriptive statistics on the data distribution. For categorical variables, use frequency and composition ratio to describe them, while for continuous variables, first calculate the maximum and minimum values of the variables, then divide the interval from the minimum value to the maximum value into five equal parts, and then calculate the proportion of the number of samples in each equal part interval to the total number. The results show that the distributions of basic information and lifestyle variables are relatively uniform. Among them, cholesterol, triglyceride, hemoglobin, platelet count, white blood cell count, alanine aminotransferase, aspartate aminotransferase, urea nitrogen, and creatinine all have a value interval that contains at least 90% of the samples. However, considering that most people are usually within a normal range for these indicators, and a small number of people will deviate from the normal range a lot, this kind of non-uniformity can be considered a normal phenomenon and conforms to the actual situation.
[0049] Analyze the impact of the collected variables on the screening of diabetes. Considering that the two variables of case number and physical examination date are mainly based on the data preprocessing process and have no potential impact on the diabetes screening model itself, they are excluded during the model construction process. All the remaining variables except the 2-hour postprandial blood glucose are used as input variables. The outcome takes whether the 2-hour postprandial blood glucose is greater than 11.1 as the output variable. Samples with a 2-hour postprandial blood glucose greater than 11.1 are called positive samples, and the rest are negative samples. The number of positive samples is 1308, and the number of negative samples is 10074.
[0050] For all input categorical variables, perform one-hot encoding, that is, encode them in the following way: First, count the number of values k of the variable; then set a k - 1-bit code all of whose bits are 0. If the value of the variable is the i-th value (i = 1, 2,..., k - 1), then the i-th position of the code is set to 1.
[0051] Lifestyle intervention is the first-line treatment recommended by the guidelines for chronic diseases, aiming to provide the fourth level of prevention and mitigate disease progression. However, these recommendations are often underestimated and unable to fulfill their potential to provide health or economic value. The prevalence of individuals entering and progressing through the four stages of the chronic disease model based on glucose disorders (including insulin resistance, prediabetes, type 2 diabetes, and vascular complications) is increasing, which has a significant impact on the economy. For example, the average annual economic cost per new diagnosis of type 2 diabetes is 25 times higher per person than that of prediabetes.
[0052] Although the pathophysiological mechanisms of metabolic syndrome (MetS) and hyperglycemia are considered different, insulin resistance is the fundamental defect in both cases and is significantly associated with type 2 diabetes and coronary heart disease. Insulin resistance, prediabetes, T2DM, and vascular complications including coronary heart disease are the main factors contributing to healthcare costs, disease burden, and impaired quality of life. Studies have shown that more than 75% of people secrete excessive insulin during an oral glucose test. When excessive insulin secretion continues to be triggered, insulin resistance develops. Most individuals with insulin resistance compensate by secreting more insulin. Initially, compensatory excessive insulin secretion is sufficient to prevent obvious disorders in blood glucose control but directly increases the risk of coronary heart disease. Environmental and lifestyle factors including diet trigger this excessive insulin secretion response and are thus insulinogenic.
[0053] Currently, more than 85% of adults are considered to meet at least one of the following diagnostic criteria for metabolic syndrome: triglycerides (TG) ≥ 150 mg / dL, blood pressure (BP) ≥ 120 / 80 mm Hg, fasting plasma glucose > 100 mg / dL or HbA1c > 5.6%, waist circumference > 88 cm in women, high-density lipoprotein cholesterol (HDL-C) < 50 mg / dL in women, waist circumference > 102 cm in men, HDL-C < 40 mg / dL in men, and / or taking related medications. These are all independent risk factors for coronary heart disease, and hyperglycemia increases the risk of developing type 2 diabetes. The presence of multiple independent risk factors for metabolic syndrome leads to the accumulation of the risk of developing coronary heart disease, and the clustering of these abnormalities is the most common metabolic abnormality in patients with known coronary heart disease. Each traditional diagnostic criterion for metabolic syndrome is associated with metabolic syndrome with different manifestations in terms of gender and race / ethnicity. Both aerobic exercise and resistance training can improve health. When these two methods are combined, it has been shown to improve glycemic control. Cardiorespiratory fitness is considered an independent risk factor for morbidity and mortality, yet it is rarely monitored and / or used as a clinical marker in primary care; therefore, improving cardiorespiratory fitness rarely becomes the goal of clinical guidance.
[0054] S202, determining a lifestyle variable according to lifestyle data; and dividing sample data with the same lifestyle variable into the same similarity group;
[0055] A similar group can be defined as a set of individuals where the similarity between any two individuals within the group is less than a certain threshold. However, since lifestyle variables are mainly categorical variables, it is difficult to characterize the similarity between different individuals through quantitative methods. In view of this problem, this study defines a similar group as follows: individuals with all the values of the given variables being equal are considered as a similar group. Considering that the dimension of the behavioral habit data is relatively high, if all variables are used to divide similar groups, the number of individuals in each similar group will be too small to analyze the relationship between age and blood glucose 2 hours after a meal. There are a total of 11 current lifestyle variables, and each variable has at least two values. In theory, the number of similar groups can reach 4096, and on average, each similar group has only about 4 samples. Due to the correlation between lifestyle variables, some similar groups with certain values may not appear. Even so, there will inevitably be a large number of similar groups with only a few samples, and it is difficult to support the analysis of blood glucose changes with age with these similar groups with a small number of samples. To solve this problem, the lifestyle variables can be screened once to reduce the number of lifestyle variables, thereby achieving dimensionality reduction of the lifestyle data. The variable screening can be carried out through the relationship between behavioral habits and diabetes risk in cross-sectional data. If the impact of a certain behavioral habit variable on the current diabetes risk is not statistically significant regardless of its value, it can be considered that this behavioral habit variable is not helpful for predicting future diabetes risk.
[0056] S202 specifically includes:
[0057] S21, determining candidate lifestyle variables according to lifestyle data;
[0058] S22, taking individuals with blood glucose 2 hours after a meal exceeding 11.1 as positive samples and the rest as negative samples. For each behavioral habit variable, data dimensionality reduction is carried out by using the chi-square test method or the t-test method (the chi-square test is used for discrete variables and the t-test is used for continuous variables) to obtain lifestyle variables with a high correlation with diabetes risk. The original hypothesis is that the corresponding lifestyle is not related to diabetes risk, and the alternative hypothesis is that the corresponding lifestyle is related to diabetes risk. The results are shown in Table 1.
[0059] Table 1
[0060]
[0061]
[0062] As can be seen from Table 1, the p-values of the correlation tests between the types of non-staple food choices, the frequency of eating fruits, taste preferences, the frequency of dining out, the intensity of physical activity, the duration of physical activity and whether the sample is positive are greater than 0.05. It is considered that these variables have no significant correlation with the diabetes risk and can be excluded in the subsequent modeling. The remaining variables include: smoking, drinking, the proportion of staple and non-staple foods, foods rich in oil, and the degree of fullness. The p-values of the correlation tests of these 5 variables with the diabetes risk are less than 0.05, and it is considered that they have a significant correlation with the diabetes risk. Therefore, the variables with a screening p-value less than 0.05 are retained, including: for constructing similar groups in the follow-up. For the selected 5 variables, the possible number of values of two of the variables is 2, and the possible number of values of the remaining three variables is 3. Therefore, the upper limit of the number of similar groups is 2×2×3×3×3 = 81. Compared with before variable screening, there are at least 4096 similar groups for 11 variables. It can be seen that through variable screening, the number of possible similar groups has decreased significantly.
[0063] Individuals who are equal in the 5 variables of smoking, drinking, the proportion of staple and non-staple foods, foods rich in oil, and the degree of fullness are grouped into a similar group. For the convenience of subsequent data processing, the values of the variables screened above are numbered with numbers, as shown in Table 2.
[0064] Table 2
[0065]
[0066] After S202, it also includes:
[0067] Preprocessing the similar groups; the preprocessing includes: screening of similar groups.
[0068] The main principles of screening are as follows: (1) The significance of predicting the diabetes risk through lifestyle habits is greater for the current non-diabetic population to guide the prevention of diabetes. Therefore, all samples with a history of diabetes should be removed; (2) When studying the significance of blood glucose increasing with age within a similar group, it should be noted that it is common for blood glucose to continuously increase in people over 60 years old. If the similar group includes people over 60 years old, it may lead to model bias. Therefore, all people over 60 years old are removed. Among the original 11382 samples, after removing people with a history of diabetes and people over 60 years old, 8739 individual samples are obtained. For the remaining samples, similar groups are formally constructed. Individuals with the same values in the 5 variables obtained by screening are recorded in a similar group, and the similar groups with different values are counted. The results are shown in Table 3.
[0069] Table 3
[0070]
[0071]
[0072]
[0073]
[0074] A total of 107 similar groups were obtained. The similar groups in the following situations were deleted: (1) The number of samples within the similar group was too small to judge the correlation between blood glucose and age; (2) If the ages within the similar group were concentrated on only a few values, even if the number of samples was acceptable, the correlation between age and blood glucose could not be analyzed either. Considering these two points, it was found that there were 8 such similar groups, involving a total of 27 samples. After removing these similar groups, there were 99 similar groups left, covering 8712 individual samples, accounting for 99.7% of the total number of screened individual samples. It can be seen that the vast majority of the individual samples that entered the similar groups would be retained and participate in the subsequent modeling.
[0075] Before modeling, it is necessary to consider whether the value distributions of various lifestyle variables are uniform. The results in this regard will affect the stability of the model and introduce biases to the effectiveness of the model. To examine the distribution of lifestyle variables in terms of the number of individuals and the number of different ages, we summarized the basic situations of the similar groups, and the results are shown in Table 4.
[0076] Table 4
[0077]
[0078]
[0079] In Table 4, the sum of the number of individuals refers to the total number of individuals included in the similar groups that meet the corresponding conditions. The sum of the number of different ages refers to the total number of different ages included in the similar groups that meet the corresponding conditions. The proportion of the number of individuals refers to the ratio of the sum of the number of individuals that meet the corresponding conditions to the total sum of the number of individuals under all values of this variable. The proportion of the sum of the number of different ages refers to the ratio of the sum of the number of different ages that meet the corresponding conditions to the total sum of the number of different ages under all values of this variable. As can be seen from Table 4, the vast majority of the proportion of the number of individuals and the proportion of the sum of the number of different ages are both above 20%. The smallest proportion of the number of individuals is the set of similar groups where the variable value of "how full" is 3 (i.e., eating 70-80% full), and its proportion of the number of individuals is 7%. However, the corresponding proportion of the sum of the number of different ages is as high as 17%, close to 20%. In addition, for the variable value of "smoking" being 2 (i.e., the group of people who have quit smoking) and the variable value of "food rich in oil" being 2 (i.e., the group of people who occasionally eat food rich in oil), their proportions of the number of individuals are also less than 20% (10% and 16% respectively), but the corresponding proportions of the sum of the number of different ages are 22% and 24% respectively. In the modeling analysis of this scenario, the uniformity of the proportion of the sum of the number of different ages has a greater impact on the model than the uniformity of the proportion of the number of individuals. Considering the above situation, overall, the distribution of each lifestyle variable is relatively uniform. Therefore, in subsequent modeling, the model bias caused by the non-uniformity of sample values will not be considered.
[0080] S203. Establish a univariate linear regression model for the age data and the change data of postprandial blood glucose in the same similar group, and obtain a similar group label; the similar group label is used to judge whether the change data of postprandial blood glucose rises with the increase of age data.
[0081] The label will be used as the output variable (dependent variable) for subsequent modeling. The method of linear regression is used to determine whether blood glucose increases with age. Specifically, within each similar group, a linear regression is performed on the age of all samples and the blood glucose 2 hours after a meal, where age is the independent variable and the blood glucose 2 hours after a meal is the dependent variable. Then, a t-test is conducted on the coefficient of the age term in the linear regression. If the p-value of the t-test is less than 0.05, it is considered that there is a significant correlation between the age of the individual samples in this similar group and blood glucose. When there is a significant negative correlation between age and blood glucose value, that is, the older the individual, the lower the blood glucose value, this situation is also regarded as no significant increase in blood glucose value with age. Therefore, the similar groups with a significant increase in blood glucose value with age need to meet two conditions: First, the p-value of the parameter significance test of the age coefficient in the age-blood glucose linear regression is less than 0.05; Second, the estimated value of the age coefficient is greater than 0. The similar groups that meet these two conditions are marked with the number 1. For the similar groups that do not fully meet these two conditions, it is considered that the individual samples in this similar group do not have a significant increase in blood glucose with age and are marked with the number 0. Through the above method, the linear regression results and data labels of each similar group are obtained. In the results of the regression of 99 similar groups, 32 are marked as 1 and 67 are marked as 0. The proportion of the sample numbers of the two types of labels is approximately 1:2, indicating that there is no obvious data imbalance. The impact of data imbalance on the model will not be considered in subsequent modeling.
[0082] S204, using the lifestyle variables in the similar group as the input and the similar group label as the output, and adopting the Naive Bayes method to obtain the diabetes risk prediction model;
[0083] Through variable screening, constructing similar groups, and adding data labels, the original individual cross-sectional data is reorganized into structured data with similar groups as tuples, including 5 lifestyle characteristics and a data label (used to reflect whether blood glucose significantly increases with age). Based on this form of data, the problem of predicting an individual's future diabetes risk can be transformed into a supervised binary classification learning problem for 5 characteristics.
[0084] The Naive Bayes method can describe the future diabetes risk as accurately as possible in terms of probability on the premise of making full use of known information; the Naive Bayes method does not rely on various distribution assumptions. The results of the model are directly calculated based on the sample observation information, and there is no need for the dependent variable to follow a normal distribution; the Naive Bayes method does not involve parameter estimation and will not increase the requirement for the sample size as the number of parameters to be estimated increases. Therefore, even in the case of a small sample size, relatively reliable risk probability estimates can be obtained; the Naive Bayes method has strong interpretability and is relatively easy to understand the internal relationship between the model output results and diabetes risk.
[0085] In view of the characteristics of the naive Bayes method described above, the naive Bayes method is used to solve the label binary classification problem based on similar groups to quantitatively predict the future risk of diabetes. The analysis results are shown in Table 5, which contains the probability of an individual's blood sugar significantly increasing with age under all possible combinations of lifestyle variable values. Through the data in the table, the risk of an individual's future blood sugar increase can be directly predicted given the lifestyle information. Considering that the probability of blood sugar increase is significantly positively correlated with the probability of having diabetes, the probability of future blood sugar increase can be used as a measure of future diabetes risk.
[0086] Table 5
[0087]
[0088]
[0089]
[0090] S205, obtaining the life habit variables of the person to be predicted, and using the diabetes risk prediction model to predict diabetes risk.
[0091] This application focuses on the living habit data of the subjects, with the goal of quantitatively predicting the future risk of diabetes, and constructs a diabetes risk prediction model based on the naive Bayes algorithm. The diabetes risk prediction model outputs the probability of a significant increase in blood sugar in the future as a measure of future diabetes risk based on the living habit data of the subjects. In order to solve the problem that this study is an analysis based on cross-sectional data, the method of constructing similar groups is used to reconstruct the data. Each similar group consists of all individuals with the same values on each living habit variable. Since there are differences in postprandial blood sugar and age between each individual in the similar group, the changes in postprandial blood sugar corresponding to the age of the individual in the similar group can be examined, thereby establishing a correlation between living habits and changes in blood sugar with age. Changes in blood sugar with age can reflect future diabetes risks.
[0092] The posterior probability that blood glucose significantly increases with age obtained by the Naive Bayes model is used to measure the future diabetes risk of the subject. The greater the probability of blood glucose increase, the greater the diabetes risk. Any subject only needs to input the values of the corresponding lifestyle variables to achieve quantitative prediction of their future diabetes risk. Through the quantitative prediction results, the size of the future diabetes risk of the subject is prompted, so as to play a warning role for high-risk individuals and strengthen their determination and compliance to take preventive measures in the future; through the quantitative prediction results, the size of the future diabetes risk of the subject is prompted, so as to play a warning role for high-risk individuals and strengthen their determination and compliance to take preventive measures in the future. On the one hand, it can make the recommendation of the preventive measure plan more targeted. On the other hand, it can also guide the subject to gradually improve their lifestyle, that is, first improve the lifestyle with the greatest risk, and after adaptation, then improve the subsequent lifestyle. In this way, the subject's fear of difficulties can be reduced, and the success rate of improvement can be further increased.
[0093] Based on the same inventive concept, an embodiment of the present application also provides a diabetes risk prediction device based on the Naive Bayes algorithm for implementing the above-mentioned diabetes risk prediction method based on the Naive Bayes algorithm. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the diabetes risk prediction device based on the Naive Bayes algorithm provided below can refer to the limitations on the diabetes risk prediction method based on the Naive Bayes algorithm in the above text, and will not be repeated here.
[0094] In an exemplary embodiment, a diabetes risk prediction device based on the Naive Bayes algorithm is provided, including:
[0095] A data acquisition module, configured to acquire sample data; the sample data includes: lifestyle data, age data, and change data of postprandial blood glucose;
[0096] A similar group construction module, configured to determine lifestyle variables according to the lifestyle data; and divide the sample data with the same lifestyle variables into the same similar group;
[0097] A similar group label determination module, configured to establish a unary linear regression model for the age data and the change data of postprandial blood glucose in the same similar group, and obtain a similar group label; the similar group label is used to judge whether the change data of postprandial blood glucose rises with the increase of the age data;
[0098] A diabetes risk prediction model determination module, configured to use the lifestyle variables in the similar group as the input and the similar group label as the output, and adopt the Naive Bayes method to obtain a diabetes risk prediction model;
[0099] A risk prediction module, configured to obtain the lifestyle variables of the person to be predicted, and perform diabetes risk prediction using a diabetes risk prediction model.
[0100] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a diabetes risk prediction method based on the Naive Bayes algorithm.
[0101] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor, implements the steps in the above method embodiments.
[0102] In an exemplary embodiment, a computer program product is provided, including a computer program, which when executed by a processor, implements the steps in the above method embodiments.
[0103] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0104] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAMs), magnetoresistive random access memories (MRAMs), ferroelectric random access memories (FRAMs), phase change memories (PCMs), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0105] The databases involved in the various embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0106] In the present application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.
[0107] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0108] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present application.
Claims
1. A diabetes risk prediction method based on the naive Bayes algorithm, characterized in that: The diabetes risk prediction method based on the naive Bayes algorithm includes: Acquire sample data; the sample data includes: life habit data, age data and postprandial blood sugar change data; Determine lifestyle variables based on lifestyle data; and divide sample data with the same lifestyle variables into the same similar group; A univariate linear regression model is established for the age data and the postprandial blood glucose change data in the same similar group, and a similar group label is obtained; the similar group label is used to determine whether the postprandial blood glucose change data increases with the increase of age data; Using the lifestyle variables in the similar groups as input and the similar group labels as output, the naive Bayes method was used to obtain a diabetes risk prediction model; Obtain the life habit variables of the person to be predicted, and use the diabetes risk prediction model to predict diabetes risk.
2. The diabetes risk prediction method based on the naive Bayes algorithm according to claim 1, characterized in that: The obtaining of sample data further includes: The sample data are processed for missing values and outliers.
3. The diabetes risk prediction method based on the naive Bayes algorithm according to claim 1, characterized in that: Determining the lifestyle variables according to the lifestyle data specifically includes: Determine candidate lifestyle variables based on lifestyle data; Data dimension reduction was performed using the chi-square test or t-test method to obtain lifestyle variables with a high correlation with diabetes risk.
4. The diabetes risk prediction method based on the naive Bayes algorithm according to claim 1, characterized in that: The sample data with the same living habit variables are divided into the same similar group, and then the following is included: Preprocessing is performed on similar groups; the preprocessing includes: similar group screening.
5. The diabetes risk prediction method based on the naive Bayes algorithm according to claim 1, characterized in that: The method to establish a univariate linear regression model is the linear regression method.
6. A diabetes risk prediction device based on the naive Bayes algorithm, characterized in that: The diabetes risk prediction device based on the naive Bayes algorithm includes: A data acquisition module is used to acquire sample data; the sample data includes: living habit data, age data and postprandial blood sugar change data; A similar group building module is used to determine the lifestyle variables according to the lifestyle data; and divide the sample data with the same lifestyle variables into the same similar group; A similar group label determination module is used to establish a univariate linear regression model with the age data and the postprandial blood glucose change data in the same similar group, and obtain a similar group label; the similar group label is used to determine whether the postprandial blood glucose change data increases with the increase of age data; A diabetes risk prediction model determination module is used to obtain a diabetes risk prediction model using the lifestyle variables in the similar group as input and the similar group labels as output, using the naive Bayes method; The risk prediction module is used to obtain the life habit variables of the person to be predicted and use the diabetes risk prediction model to predict diabetes risk.
7. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the diabetes risk prediction method based on the naive Bayes algorithm described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting diabetes risk based on the naive Bayes algorithm described in any one of claims 1 to 5 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for predicting diabetes risk based on the naive Bayes algorithm described in any one of claims 1 to 5 is implemented.