Device for predicting the probability of impaired fasting glucose in non-diabetics, method for predicting the probability of impaired fasting glucose in non-diabetics and program stored in a recording medium
Patent Information
- Application Number
- KR1020220074200
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-06-17
Smart Images

Figure 112022063553965-PAT00028_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a technology for predicting the probability of impaired fasting glucose, and more specifically, to an apparatus, a method, and a program stored on a recording medium for predicting the probability of impaired fasting glucose in non-diabetics. Background Technology
[0002] With the increase in the elderly population, the number of diabetes patients is also on the rise. Diabetes is a disease characterized by persistently high blood sugar levels and is known to cause damage to microvessels in the retina and kidneys, leading to various complications such as blindness and kidney disease.
[0003] Although the prevalence of diabetes has increased over the past 20 years, many diabetic patients are unaware of their condition because there are almost no noticeable symptoms unless blood sugar levels are excessively high, making early detection difficult.
[0004] Impaired fasting glucose is diagnosed when fasting blood glucose is 100–125 mg / dL, which is intermediate between normal and diabetes levels, or when glycated hemoglobin is 5.7–6.4%; these individuals are defined as prediabetes because they have a high likelihood of progressing to diabetes in the future. [Prior Art Literature] Prior Art Literature 1: Korean Patent Publication No. 10-2018-0079208 (July 10, 2018) Prior Art Literature 2: Mohammed Hussein Al-Sarem et al., Feature Selection and Classification using CatBoost Method for Improving the Performance of Predicting Parkinson's Disease, International conference of advanced computing and informatics, Morocco (October 23, 2020) Prior Art Literature 3: Mohammadreza Bozorgmanesh et al., Diabetes prediction, lipid accumulation product, and adiposity measures; 6-year follow-up: Tehran lipid and glucose study, Lipids in Health and Disease (2010.12.31) The problem to be solved
[0005] In Korea, as of 2021, one in seven adults aged 30 or older (13.8%) suffers from diabetes, and one in four people aged 30 or older (26.9%) are reported to have prediabetes; however, 35% of patients diagnosed with diabetes between 2016 and 2018 had no symptoms. Therefore, to prevent diabetes, it is necessary to screen for impaired fasting glucose, which corresponds to the prediabetic stage, at an early stage and monitor for progression to diabetes.
[0006] The present invention was developed to solve the problems of the prior art described above. The objective of the present invention is to provide a device, a method, and a program for predicting the probability of impaired fasting glucose in non-diabetics, which can predict the probability of impaired fasting glucose in non-diabetics with high accuracy by considering complex influencing factors, while presenting the results in a visual form that medical professionals can understand. means of solving the problem
[0007] As a means to solve the above-mentioned problem, a device for predicting the probability of impaired fasting glucose according to one embodiment of the present invention, which is a device for predicting the probability of impaired fasting glucose in a non-diabetic person using explanatory variables, may include: a measurement value acquisition unit for acquiring a measurement value for multiple explanatory variables including a plurality of explanatory variables related to the subject of prediction; a score calculation unit for calculating an impaired fasting glucose prediction score corresponding to the measurement value acquired for each explanatory variable of the multiple explanatory variables; and an impaired fasting glucose probability calculation unit for calculating the probability of impaired fasting glucose for the subject of prediction from the total sum of the impaired fasting glucose prediction scores calculated for each explanatory variable.
[0008] In one embodiment, the score calculation unit can calculate an impaired fasting glucose prediction score corresponding to the measured value of each explanatory variable by reflecting the weight of each explanatory variable.
[0009] In one embodiment, the score calculation unit and the fasting glucose impairment probability calculation unit calculate a fasting glucose impairment prediction score and a fasting glucose impairment probability based on a nomogram, and the nomogram may include: a prediction point line having a score range between 0 and 100, representing a fasting glucose impairment prediction score assigned to each explanatory variable; a variable line having a length corresponding to the degree of influence on the fasting glucose impairment probability for each explanatory variable, and including a start point and an end point that match at least a part of the score range of the prediction point line; a total point line representing the total sum of the fasting glucose impairment prediction scores calculated for each explanatory variable; and a probability line representing the fasting glucose impairment probability corresponding to the total sum of the total point line.
[0010] In one embodiment, the multiple explanatory variables may be selected based on feature importance using the mean decrease in impurity calculated using the CatBoost algorithm.
[0011] In one embodiment, the multiple explanatory variables may include (1) age, (2) BMI, (3) high cholesterol, (4) marital status, (5) experience of drinking at least one drink per month in the past year; (6) WHtR, (7) smoking status, and (8) high blood pressure.
[0012] In one embodiment, the score calculation unit calculates a fasting glucose impairment prediction score corresponding to the measurement value of each explanatory variable by reflecting the weight of each explanatory variable, and the weight of each explanatory variable may have a larger value in the order of (1) age, (2) BMI, (3) high cholesterol, (4) marital status, (5) experience of drinking at least one cup a month in the past year; (6) WHtR, (7) smoking status, and (8) high blood pressure.
[0013] As another means for solving the above-mentioned problem, a method for predicting the probability of impaired fasting glucose according to one embodiment of the present invention is a method for predicting the probability of impaired fasting glucose performed by an apparatus for predicting the probability of impaired fasting glucose of a non-diabetic person using multiple explanatory variables, and may include: a measurement value acquisition step of acquiring a measurement value for multiple explanatory variables including a plurality of explanatory variables related to the subject to prediction; a score calculation step of calculating an impairment of fasting glucose prediction score corresponding to the measurement value acquired for each explanatory variable of the multiple explanatory variables; and a impairment of fasting glucose probability calculation step of calculating the probability of impairment of fasting glucose for the subject to prediction from the total sum of the impairment of fasting glucose prediction scores calculated for each explanatory variable.
[0014] As another means for solving the above-mentioned problem, a program according to one embodiment of the present invention may include a program stored on a recording medium to perform the above-mentioned fasting glucose disorder probability prediction method by a computer. Effects of the invention
[0015] According to one embodiment of the present invention, a device, method, and program for predicting the probability of fasting glucose in non-diabetics can be provided, which can predict the probability of fasting glucose in non-diabetics with high accuracy by considering complex influencing factors and presenting it in a visual form that medical professionals can understand, by acquiring a measurement value for multiple explanatory variables including a plurality of explanatory variables related to a subject to prediction, calculating a fasting glucose impairment prediction score corresponding to the measurement value acquired for each explanatory variable of the multiple explanatory variables, and calculating the probability of fasting glucose impairment for the subject to prediction from the total sum of the fasting glucose impairment prediction scores calculated for each explanatory variable. Brief explanation of the drawing
[0016] FIG. 1 is a block diagram illustrating the configuration of a device for predicting the probability of impaired fasting glucose in non-diabetics according to one embodiment of the present invention. Figure 2 is a graph showing the feature importance of factors related to the probability of impaired fasting glucose in non-diabetics calculated using CatBoost according to one embodiment of the present invention. FIG. 3 is a diagram illustrating a nomogram for predicting the probability of impaired fasting glucose according to one embodiment of the present invention. Figure 4 is a graph showing the accuracy of a nomogram predicting the probability of impaired fasting glucose according to one embodiment of the present invention. Figure 5 is a graph showing the AUC of a nomogram predicting the probability of impaired fasting glucose according to one embodiment of the present invention. Figure 6 is a graph showing the calibration plot of a nomogram predicting the probability of impaired fasting glucose according to one embodiment of the present invention. FIG. 7 is a flowchart illustrating a method for predicting the probability of fasting impaired blood glucose performed by a device for predicting the probability of fasting impaired blood glucose according to one embodiment of the present invention. Specific details for implementing the invention
[0017] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the attached drawings.
[0018] The embodiments of the present invention are provided to more fully explain the invention to those skilled in the art, and the following embodiments may be modified in various different forms, and the scope of the invention is not limited to the following embodiments. Rather, these embodiments are provided to make the disclosure more faithful and complete and to fully convey the spirit of the invention to those skilled in the art.
[0019] The various embodiments described herein may be implemented, for example, in a recording medium readable by a computer or similar device using software, hardware, or a combination thereof.
[0020] According to hardware implementation, the embodiments described herein may be implemented using at least one of ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), processors, controllers, microcontrollers, microprocessors, and electrical units for performing functions.
[0021] According to the software implementation, embodiments such as procedures or functions may be implemented together with separate software modules that perform at least one function or operation. The software code may be implemented by a software application written in a suitable programming language.
[0023] FIG. 1 is a block diagram illustrating the configuration of a fasting blood glucose disorder probability prediction device (10) according to one embodiment of the present invention.
[0024] Referring to FIG. 1, the fasting blood glucose disorder probability prediction device (10) may include a measurement value acquisition unit (110), a score calculation unit (130), a fasting blood glucose disorder probability calculation unit (150), and a database unit (170).
[0025] The measurement acquisition unit (110) can acquire measurement values for multiple explanatory variables including multiple explanatory variables related to the prediction subject. The fasting glucose disorder probability prediction device (10) can calculate the fasting glucose disorder probability based on the measurement values (response values) for each of the multiple explanatory variables acquired from the prediction subject information.
[0026] Multiple explanatory variables can be selected based on feature importance derived from the mean decrease in impurity calculated using the CatBoost algorithm. A higher value for feature importance based on the mean decrease in impurity calculated using the CatBoost algorithm can be interpreted as indicating a greater influence of the corresponding explanatory variable on impaired fasting glucose.
[0027] The CatBoost algorithm was developed to address the problems of the Gradient Boosting algorithm. When using the Gradient Boosting algorithm, the training data is derived as a residual of the entire prediction data at every boosting step, causing the training data to change into a distribution that converges to the prediction data. This change in distribution leads to problems such as overfitting or prediction shift, which causes inaccurate predictions. Additionally, there is a problem with the difficulty of handling categorical variables when using the Gradient Boosting algorithm. Previously, for categorical variables, a new binary variable had to be added for each category; therefore, as the number of categories increased, a large number of new binary variables had to be created. Consequently, the amount of statistics increased, leading to higher computation time and memory consumption.
[0028] Unlike Gradient Boosting algorithms, Catboost builds a model using Ordered Boosting. While conventional boosting models learn all residual errors sequentially, Catboost calculates the residual error of a portion of the data to build a model, and then uses this model to calculate the residual error for the remaining data. Additionally, Catboost prevents overfitting by shuffling the data order through Random Permutation within Ordered Boosting.
[0029] For preprocessing categorical variables, the Catboost algorithm calculates the average sample value of variables with the same category from a dataset that has undergone Random Permutation, which can be expressed mathematically as [Equation 1].
[0030] [Equation 1]
[0031]
[0032] The multiple explanatory variables selected to predict the probability of impaired fasting glucose are: (1) age (1=30–39 years, 2=40–49 years, 3=50–64 years); (2) BMI (1=underweight, 2=normal weight, 3=pre-obesity, 4=obesity stage 1, 5=obesity stage 2 or higher); (3) high cholesterol (0=none, 1=present); (4) marital status (1=cohabiting with spouse, 2=separated / divorced / widowed, 3=single); (5) experience of drinking at least one cup per month in the past year (0=none, 1=present); (6) WHtR (1=less than 0.5, 2=0.5 or greater); (7) smoking status (1=non-smoker, 2=former smoker, 3=current smoker); (8) May include hypertension (1=normal, 2=prehypertension, 3=hypertension).
[0033] The score calculation unit (130) can calculate a fasting glucose impairment prediction score corresponding to the measurement value of each explanatory variable by reflecting the weight of each explanatory variable. The score calculation unit (130) can calculate a fasting glucose impairment prediction score corresponding to the measurement value obtained for each explanatory variable of multiple explanatory variables based on a nomogram.
[0034] The fasting impaired glucose prediction score corresponding to the measurement value of each explanatory variable can be calculated based on nomogram information stored in the database unit (170). The nomogram information stored in the database unit (170) includes fasting impaired glucose prediction score data corresponding to the measurement value obtained for each explanatory variable as shown in FIG. 3, and each fasting impaired glucose prediction score can be derived based on this nomogram information.
[0035] The fasting glucose disorder probability calculation unit (150) can calculate the probability of a predicted subject having a fasting glucose disorder based on a nomogram from the total sum of the fasting glucose disorder prediction scores calculated for each explanatory variable. Once the total sum of the fasting glucose disorder prediction scores calculated for all explanatory variables is determined, the probability of a predicted subject having a fasting glucose disorder can be calculated based on the nomogram information stored in the database unit (170). The nomogram information stored in the database unit (170) includes fasting glucose disorder probability data corresponding to the total sum of the prediction scores as illustrated in FIG. 3, and the probability of a predicted subject having a fasting glucose disorder can be derived based on this nomogram information.
[0036] Nomogram information for calculating a score for a fasting blood glucose disorder prediction score by the score calculation unit (130) and for calculating the fasting blood glucose disorder probability by the fasting blood glucose disorder probability calculation unit (150) to calculate the probability of a fasting blood glucose disorder of a predicted subject can be stored in the database unit (170). Additionally, the database unit (170) can store measurement values for multiple explanatory variables including multiple explanatory variables related to a predicted subject obtained by the measurement value acquisition unit (110).
[0037] A nomogram may include a prediction point line, a variable line, a total point line, and a probability line. The prediction point line represents the predicted score for impaired fasting glucose assigned to each explanatory variable and has a score range between 0 and 100. The variable line has a length corresponding to the degree to which it influences the probability of impaired fasting glucose for each explanatory variable and includes a start point and an end point that match at least a portion of the score range of the prediction point line. The total point line represents the sum of the predicted scores for impaired fasting glucose calculated for the explanatory variables. The probability line represents the probability of impaired fasting glucose corresponding to the sum of the total point line.
[0039] The features of the device and method for predicting the probability of impaired fasting glucose according to the present invention will be explained below through the description of the development and verification process of the model for predicting impaired fasting glucose in non-diabetics.
[0041] [Source]
[0042] The inventor utilized raw data from the 2020 National Health and Nutrition Survey, which is a national statistical data (Approval No. 1702) administered by the Ministry of Health and Welfare and the Korea Centers for Disease Control and Prevention. The sample for the National Health and Nutrition Survey was selected based on the population residing in Korea as the population, and the subjects of the survey were selected from household members aged 1 year or older using stratified cluster sampling and systematic sampling methods based on data from the Population and Housing Census (complete enumeration).
[0043] The 2020 National Health and Nutrition Survey covered 9,949 individuals across 192 sample districts nationwide, and due to the suspension of the survey caused by the COVID-19 pandemic, 17,359 subjects (participation rate 74.0%) completed the health questionnaire and medical examination. Regarding the survey method, the health questionnaire was conducted by surveyors visiting target households in person through face-to-face interviews and self-administered questionnaires.
[0044] The screening survey consisted of blood pressure measurement, anthropometric measurements, and blood tests. During the survey period, medical professionals (doctors and nurses) visited the survey area using mobile screening vehicles to conduct one-on-one screenings and health questionnaires. Pregnant women and individuals diagnosed with diabetes prior to the survey were excluded. Finally, 3,019 adults aged 30 to under 65 who completed all blood tests, anthropometric measurements, blood pressure measurements, and health questionnaires were analyzed.
[0046] [Measurement and Definition of Variables]
[0047] The dependent variable, impaired fasting glucose, was classified into normal (glycated hemoglobin levels less than 5.7% and fasting blood glucose levels 100 mg / dl or less) and impaired fasting glucose (glycated hemoglobin levels 5.7% to 6.4% and fasting blood glucose levels 100 to 125 mg / dl) based on the diagnosis of a medical professional according to the Clinical Practice Guidelines of the Korean Diabetes Association (2021).
[0048] The explanatory variables included sociodemographic factors, health habit factors, anthropometric factors, dietary habit factors, and cardiovascular disease risk factors. Sociodemographic factors included gender (male, female), marital status (cohabiting with spouse, separated / divorced / widowed, single), age (30-39, 40-49, 50-64), residential area (urban, rural), and average monthly household income (less than 2 million won, 2-4 million won, more than 4 million won). Health habit factors included whether at least one drink was consumed per month in the past year (yes, no), smoking status (non-smoker, former smoker, current smoker), subjective stress level (almost none, moderate, high), average daily moderate physical activity time in leisure activities (activities that cause slight shortness of breath or a slightly faster heart rate, such as jogging or strength training: none, 10 minutes or more but less than 1 hour, 1 hour or more), average daily sedentary time (4 hours or less, 5 hours or more but less than 7 hours, 8 hours or more), number of days walking for 10 minutes or more in the past week (none, 1-2 days, 3-4 days, 5-6 days, every day), and average daily sleep time during the week (5 hours or less, 6-7 hours, 8 hours or more).
[0049] Anthropometric factors included Body Mass Index (BMI: underweight (less than 18.5 kg / m2), normal weight (18.5–23 kg / m2), pre-obesity (23–25 kg / m2), stage 1 obesity (25–30 kg / m2), stage 2 obesity or worse (greater than 30 kg / m2)) and waist-to-height ratio (WHtR: classified as less than 0.5 and greater than 0.5).
[0050] Dietary habits included the average weekly frequency of breakfast over the past year (5–7 times a week, 3–4 times a week, 1–2 times a week, rarely), and the average frequency of eating out including delivery food over the past year (more than once a day, less than once a day).
[0051] Cardiovascular disease risk factors included the prevalence of high cholesterol (total cholesterol ≥ 240 mg / dL: yes, no), high triglycerides (triglycerides ≥ 200 mg / dL: yes, no), and hypertension (normal blood pressure: systolic blood pressure < 120 mmHg and diastolic blood pressure < 80 mmHg; prehypertension: systolic blood pressure 120–140 mmHg and diastolic blood pressure 80–90 mmHg; hypertension: systolic blood pressure > 140 mmHg or diastolic blood pressure > 90 mmHg)).
[0052] Blood pressure was measured on the right upper arm using a mercury sphygmometer (Wall Unit 33, Baumanometer, America) under the supervision of a nurse after the patient rested for 5 minutes before the test.
[0054] [Selection of Variables]
[0055] CatBoost (categorical boosting) is an algorithm that offers superior performance compared to existing XGBoost and LightGBM. Gradient Boosting algorithms such as XGBoost and LightGBM have two limitations: first, as boosting training progresses, changes in the data distribution occur, which ultimately leads to prediction shifts that cause overfitting or inaccurate predictions. Second, there is a problem with the processing of categorical variables being time-consuming.
[0056] For example, in the case of XGBoost or LightGBM, computation time and memory consumption increase because the statistics themselves increase when new binary variables are created. To overcome these limitations, Catboost builds a model using Ordered Boosting. Existing boosting models built a model by calculating residuals on all training data. However, as shown in [Equation 2], Catboost calculates residuals using only a portion of the training data, builds a model with them, and then reuses the residuals from the data to re-predict the values of this model.
[0057] [Equation 2]
[0058]
[0059] CatBoost has the advantage of being easier to use compared to other Gradient Boosting algorithms that require hyperparameter tuning, as it solves hyperparameter optimization with an internal algorithm without the need for special hyperparameter optimization. In one embodiment of the present invention, the number of trees for catboost was set to 100, the regularization Lambda to 3, the learning rate to 0.300, and the limit depth of individual trees to 6.
[0060] The selection of key variables for predicting impaired fasting glucose in the CatBoost algorithm was verified by calculating feature importance using the mean decrease in impurity. In addition, to efficiently interpret risk probabilities in the nomogram for predicting the high-risk group for impaired fasting glucose, a nomogram was developed by selecting only the top 8 variables with high feature importance.
[0062] [Development and Validation of Logistic Nomogram]
[0063] A predictive model for impaired fasting glucose was constructed using logistic regression analysis to identify the independent associations of variables related to impaired fasting glucose in community-dwelling non-diabetics by inputting the top 8 variables with high importance identified in CatBoost.
[0064] In the regression model, the adjusted Odds Ratio (aOR) and 95% CI, adjusted for all confounding factors, were presented, respectively. The developed prediction model included a nomogram to enable clinicians to easily interpret the predicted results (predicted probabilities).
[0065] A nomogram is a two-dimensional diagram representing the relationships between multiple risk factors to calculate the prediction probability of a disease simply and efficiently, and consists of a point line, a risk factor line, a probability line, and a total point line.
[0066] The score line is placed at the top of the nomogram to derive scores corresponding to the class of individual risk factors, and the risk factor line is eight, which is the number of risk factors for fasting impaired glucose in one embodiment of the present invention. The total point line represents the sum of the scores of the individual risk factors. The probability line is placed at the bottom of the nomogram as the probability value of the prediction of fasting impaired glucose in non-diabetics, finally calculated based on the total point line.
[0067] Let the presence of impaired fasting glucose (impaired fasting glucose = 1, not impaired fasting glucose = 0) be the dependent variable. The dependent variable Y follows a Bernoulli distribution, and the probability of having impaired fasting glucose is P(Y = 1) = , the probability of the case without impaired fasting glucose is P(Y = 0) = 1 - is. Therefore, when the explanatory variable X=x, the probability that the dependent variable Y=y is as follows.
[0068] [Equation 3-1]
[0069]
[0070] Probability of success If we represent it as a linear probability model And in this case, it has a structural defect. That is, the left side has a probability between 0 and 1, but the right side has values for the entire real range, and the regression coefficient and When estimating, the least squares estimator (LSE) no longer has minimum variance. Therefore, the probability of success for the explanatory variable x It can be expressed as a non-linear function as follows.
[0071] [Equation 3-2]
[0072]
[0073] Both sides of [Equation 3-2] have real values between 0 and 1. Also, k influencing factors The equation when considering is as follows.
[0074] [Equation 3-3]
[0075]
[0076] [Equation 3-3] can be rearranged so that the right-hand side becomes a linear function as follows.
[0077] [Equation 3-4]
[0078]
[0079] [Equation 3-4] is a logistic regression model.
[0080] Once the coefficients for the logistic regression model are estimated, scores are calculated for each factor using the LP (Linear Predictor) values. For the j-th attribute value of the i-th factor It is calculated, and each factor is assigned a score according to the following formula.
[0081] [Equation 3-5]
[0082]
[0083] Here represents the LP values for the factor with the largest absolute value of the estimated regression coefficient. Therefore, all factors are assigned a score between 0 and 100 based on the point line according to [Equation 3-5], and the most influential attribute value is assigned 100 points, having the largest score.
[0084] Once scores are assigned to the attributes of all factors, each individual has a total score (total points = ) can be calculated. Therefore, the probability (p) corresponding to this value is found using the total points line and the probability line. To draw the probability line corresponding to the total score, the minimum and maximum probabilities corresponding to the minimum and maximum values of the total score must first be calculated, and these are calculated through the following process.
[0085] [Equation 3-6]
[0086]
[0087] [Equation 3-7]
[0088]
[0089] [Equation 3-8]
[0090]
[0091] The maximum and minimum probabilities for cases of impaired fasting glucose in the nomogram are and The value is fixed and depends on the total score, and the probability can be calculated using points per unit of LP as shown in [Equation 3-7] and [Equation 3-8].
[0092] When the maximum probability corresponding to the minimum total score and the maximum total score is determined on the probability line, the total score corresponding to the probability existing between them is calculated through the following formula.
[0093] [Essence 3-9]
[0094]
[0095] Finally, when the total score corresponding to the predicted probability (p) is calculated through [Equation 3-9], the total points line and the probability line can be represented on the logistic nomogram.
[0096] The predictive performance (F1-score, the area under the curve (AUC), precision, recall, general accuracy, calibration plot) of the nomogram predicting impaired fasting glucose in non-diabetics was verified using 10-fold cross-validation.
[0097] [General Characteristics of Subjects with Impaired Fasting Glucose]
[0098] The results of the chi-square test analyzing differences in general characteristics among groups according to the prevalence of impaired fasting glucose are presented in [Table 1-1] and [Table 1-2]. Among the total 3,019 subjects, 29.1% (879 subjects) had impaired fasting glucose. The chi-square test results showed significant differences in marital status, age, average monthly household income, drinking experience in the past year, subjective stress, average number of walking days per week, average weekday sleep duration, BMI, WHtR, average weekly frequency of breakfast in the past year, average weekly frequency of eating out in the past year, high cholesterol, high triglycerides, and hypertension (p<0.05).
[0101] [Table 1-1]
[0102]
[0105] [Table 1-2]
[0106]
[0108] [Predictors of Impaired Fasting Glucose in Community-Residing Non-Diabetic Individuals in Korea]
[0109] The results of calculating the feature importance of factors related to impaired fasting glucose in non-diabetics using CatBoost are presented in Figure 2.
[0110] In Fig. 2, the numbers on the left side of the graph represent: (1) age (1=30–39 years, 2=40–49 years, 3=50–64 years); (2) high cholesterol (0=none, 1=present); (3) WHtR (1=less than 0.5, 2=0.5 or greater); (4) BMI (1=underweight, 2=normal weight, 3=pre-obesity, 4=obesity stage 1, 5=obesity stage 2 or higher); (5) experience of drinking at least one cup per month in the past year (0=none, 1=present); (6) marital status (1=living with spouse, 2=separated / divorced / widowed, 3=single); (7) hypertension (1=normal, 2=pre-hypertension, 3=hypertension); (8) It indicates whether or not you smoke (1=non-smoker, 2=past smoker, 3=current smoker).
[0111] Based on feature importance, the top 8 variables were identified as age, high cholesterol, WHtR, BMI, experience of drinking at least one drink per month in the past year, marital status, hypertension, and smoking status.
[0112] The results of the logistic regression analysis for predicting impaired fasting glucose in Korean non-diabetics using the top 8 variables with high importance in Catboost are presented in [Table 2]. The analysis results of the adjusted model showed that the independent influencing factors of impaired fasting glucose were separation / divorce / widowhood of a spouse (AOR=1.79, 95% CI: 1.26, 2.55), age (40–49 years: AOR=1.59, 50–64 years: AOR=4.09), no alcohol consumption in the past year (AOR=1.47, 95% CI: 1.21, 1.77), current smoker (AOR=1.36, 95% CI: 1.06, 1.74), BMI grade 2 obesity or higher (AOR=3.80, 95% CI: 1.80, 8.01), WHtR 0.5 or higher (AOR=1.37, 95% CI: 1.04, 1.81), and prevalence of hypercholesterol (AOR=2.03, 95% CI: 1.66, 2.49) Hypertension prevalence (prehypertension: AOR=1.34, hypertension: AOR=1.31) was confirmed (p<0.05).
[0114] [Table 2]
[0115]
[0117] [Development and Validation of a Predictive Nomogram for High-Risk Groups of Impaired Fasting Glucose in Non-Diabetic Patients]
[0118] The predictive nomogram for impaired fasting glucose in non-diabetics in Korea is presented in Figure 3. In the predictive nomogram, non-diabetics aged 50 to 64 who live with a spouse, have a BMI of at least level 2 (over 30 kg / m2), have hypertension and high cholesterol, have no history of drinking alcohol in the past year, and are current smokers were found to have a high risk prediction probability of impaired fasting glucose of 87%.
[0119] The predictive performance of the developed impaired fasting glucose prediction nomogram was verified using general accuracy (Fig. 4), AUC (Fig. 5), precision, recall, F1-score, and Calibration plot (Fig. 6).
[0120] The predicted and observed probabilities for the impaired fasting glucose group and the normal blood glucose group were compared using a calibration plot (Fig. 6) and a chi-square test, and there was no significant difference between the predicted and observed probabilities (P<0.05). The results of the 10-Fold Cross Validation showed that the general accuracy of the nomogram predicting impaired fasting glucose in non-diabetics was 0.73, AUC was 0.75, Precision was 0.71, recall was 0.73, and F1-score was 0.70.
[0122] FIG. 7 is a flowchart illustrating a method for predicting the probability of fasting blood glucose impairment performed by a fasting blood glucose impairment probability prediction device (10) according to one embodiment of the present invention.
[0123] Referring to FIG. 7, the method for predicting the probability of impaired fasting glucose may include a measurement acquisition step (S110), a score calculation step (S130), and a fasting glucose impairment probability calculation step (S150).
[0124] In the measurement acquisition step (110), the measurement acquisition unit (110) can acquire a measurement value for multiple explanatory variables including multiple explanatory variables related to the prediction target.
[0125] In the score calculation step (S130), the score calculation unit (130) can calculate a fasting glucose disorder prediction score corresponding to the measurement value obtained for each explanatory variable of multiple explanatory variables.
[0126] In the step of calculating the probability of impaired fasting blood glucose (S150), the impaired fasting blood glucose probability calculation unit (150) can calculate the probability of impaired fasting blood glucose from the total sum of the predicted scores of impaired fasting blood glucose calculated for each explanatory variable.
[0127] In relation to each step (S110, S130, S150) of the fasting blood glucose disorder probability prediction method, the above-described details regarding the measurement value acquisition unit (110), score calculation unit (130), and fasting blood glucose disorder probability calculation unit (150) of the fasting blood glucose disorder probability prediction device (10) may be referenced.
[0128] The steps or processes described above may be executed by hardware components, software components, and / or a combination of hardware components and software components. For example, the steps or processes described in the embodiments may be executed using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.
[0129] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0130] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0131] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0132] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
Claim 1 An apparatus for predicting the probability of impaired fasting glucose in non-diabetics using multiple explanatory variables, comprising: a measurement acquisition unit for acquiring measurement values for multiple explanatory variables including multiple explanatory variables related to the subject to prediction; a score calculation unit for calculating an impaired fasting glucose prediction score corresponding to the measurement values acquired for each explanatory variable of the multiple explanatory variables; and an impaired fasting glucose probability calculation unit for calculating the probability of impaired fasting glucose for the subject to prediction from the total sum of the impaired fasting glucose prediction scores calculated for each explanatory variable, wherein the multiple explanatory variables include (1) age, (2) BMI, (3) high cholesterol, (4) marital status, and (5) experience of drinking at least one cup per month in the past year; (6) WHtR, (7) smoking status, (8) hypertension, and the score calculation unit calculates an impaired fasting glucose prediction score corresponding to the measurement value of each explanatory variable by reflecting the weight of each explanatory variable, and the weights of each explanatory variable are (1) age, (2) BMI, (3) high cholesterol, (4) marital status, (5) experience of drinking at least one cup per month in the past year; (6) WHtR, (7) smoking status, (8) hypertension, having larger values in that order, and the score calculation unit and the fasting glucose impairment probability calculation unit calculate a fasting glucose impairment prediction score and a fasting glucose impairment probability based on a nomogram, wherein the nomogram includes: a prediction point line having a score range between 0 and 100, representing a fasting glucose impairment prediction score assigned to each explanatory variable; a variable line having a length corresponding to the degree of influence on the fasting glucose impairment probability for each explanatory variable, and including a start point and an end point that match at least a part of the score range of the prediction point line; a total point line representing the total sum of the fasting glucose impairment prediction scores calculated for each explanatory variable; and a probability line representing the fasting glucose impairment probability corresponding to the total sum of the total point line, [Drawing] The above nomogram is a fasting glucose disorder probability prediction device that is the nomogram described in the above drawing. Claim 2 delete Claim 3 delete Claim 4 A device for predicting the probability of impaired fasting glucose, wherein the multiple explanatory variables are selected based on feature importance using the mean decrease in impurity calculated using the CatBoost algorithm. Claim 5 delete Claim 6 delete Claim 7 A method for predicting the probability of impaired fasting glucose performed by an impaired fasting glucose probability prediction device for predicting the probability of impaired fasting glucose in non-diabetics using multiple explanatory variables, comprising: a measurement value acquisition step for acquiring measurement values for multiple explanatory variables including multiple explanatory variables related to a subject to prediction; a score calculation step for calculating an impaired fasting glucose prediction score corresponding to the measurement values acquired for each explanatory variable of the multiple explanatory variables; and a impaired fasting glucose probability calculation step for calculating the probability of impaired fasting glucose for a subject to prediction from the total sum of the impaired fasting glucose prediction scores calculated for each explanatory variable, wherein the multiple explanatory variables include (1) age, (2) BMI, (3) high cholesterol, (4) marital status, and (5) experience of drinking at least one cup per month in the past year; (6) WHtR, (7) smoking status, (8) hypertension are included, and in the score calculation step above, a fasting impaired glucose prediction score corresponding to the measurement value of each explanatory variable is calculated by reflecting the weight of each explanatory variable, and the weight of each explanatory variable is (1) age, (2) BMI, (3) high cholesterol, (4) marital status, (5) experience of drinking at least one cup per month in the past year; (6) WHtR, (7) smoking status, (8) hypertension, having larger values in that order, and in the score calculation step and the fasting glucose impairment probability calculation step, the fasting glucose impairment prediction score and the fasting glucose impairment probability are calculated based on a nomogram, and the nomogram comprises: a prediction point line having a score range between 0 and 100, representing the fasting glucose impairment prediction score assigned to each explanatory variable; a variable line having a length corresponding to the degree of influence on the fasting glucose impairment probability for each explanatory variable, and including a start point and an end point that match at least a part of the score range of the prediction point line; and a total point line representing the total sum of the fasting glucose impairment prediction scores calculated for each explanatory variable.Includes a probability line representing the probability of impaired fasting glucose corresponding to the total sum of the above total lines, [Figure]; A method for predicting the probability of impaired fasting glucose, wherein the above nomogram is the nomogram described in the above drawing. Claim 8 A program stored on a recording medium to perform the method for predicting the probability of impaired fasting glucose described in paragraph 7 by a computer.
Citation Information
Patent Citations
Apparatus and method for predicting disease risk of metabolic disease
KR1020180079208A