Method and system for generating personalized biological age prediction model

The personalized biological age prediction model uses binary logistic regression to calculate an 'excess aging factor' based on gender and age bands, addressing inconsistencies in existing models and providing a more accurate biological age assessment.

JP7680803B2Active Publication Date: 2025-05-21ユジン バイオソフト カンパニーリミテッド
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024513366
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-28
Filing Date
2022-02-24
Publication Date
2025-05-21
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

Existing biological age prediction models provide only a single numerical value, lacking objective quantitative and qualitative analysis, and exhibit inconsistencies in predicting biological age across different age groups and genders, leading to overestimation in younger populations and underestimation in older populations.

Method used

A personalized biological age prediction model using binary logistic regression to calculate an 'excess aging factor' relative to chronological age, considering gender and age bands, generating a probability spectrum rather than a single numerical value.

Benefits of technology

Provides a more accurate and objective prediction of biological age by calculating excess age relative to chronological age, accounting for gender and age-specific differences, resulting in a more reliable and personalized biological age assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680803000012
    Figure 0007680803000012
  • Figure 0007680803000013
    Figure 0007680803000013
  • Figure 0007680803000014
    Figure 0007680803000014
Patent Text Reader

Abstract

The present invention relates to a personalized biological age prediction model generating method and system for generating a model capable of predicting an individual's biological age by obtaining excess age relative to chronological age by age based on health checkup data. More specifically, the present invention relates to a personalized biological age prediction model generating method and system for generating a personalized biological age prediction model capable of predicting an individual's biological age according to the biological age prediction model by constructing a biological age prediction model by gender and chronological age band in consideration of the fact that aging mechanisms differ between men and women or chronological age bands.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a model generation method for predicting biological age through personalization, and more particularly to a personalized biological age prediction model generation method and system for generating a model capable of predicting an individual's biological age by calculating excess age relative to age-specific chronological age based on health checkup data. [Background technology]

[0002] Generally, chronological age indicates the difference between the current year and the year of birth, and people born in the same year have no choice but to have the same chronological age, regardless of the individual's current health condition.

[0003] Therefore, since chronological age alone cannot fully indicate "aging" related to an individual's current health status or general decline in physical function, there is a need to develop technology that can predict or measure "biological age," which indicates the decline in physical function associated with aging.

[0004] Biological age is different from chronological age in that it is a numerical representation of the parts of the body that vary depending on the overall health of the body, i.e., it is a numerical representation of the health and aging of the body. Since people of the same chronological age may have different physical health conditions, it can be said that using biological age determined by measuring or estimating the overall physical health condition is a more accurate way of measuring current overall health condition, aging, and even actual life expectancy than chronological age.

[0005] <Existing research on predicting / measuring biological age> Research attempting to measure biological age began with Comfort in 1969 and has continued to the present day.

[0006] The biomarkers used to measure biological age should have the following characteristics: 1) Providing information about bodily functions or metabolic systems 2) Possession of quantitative characteristics that correlate with chronological age 3) Reproducibility, sensitivity, and specific properties 4) It is suitable for application to experimental animals as well as humans.

[0007] Taking these factors into consideration, studies have been conducted to measure biological age using physical, physiological, and biochemical biomarkers.

[0008] Biomarkers commonly used to measure biological age include body mass index (BMI), blood pressure (systolic blood pressure, diastolic blood pressure), waist circumference, lung capacity, muscle mass, albumin, and cholesterol levels. Using these as independent factors, we are researching biological age measurement models using multivariable linear regression analysis and principal component analysis (PCA).

[0009] <Research into predicting mortality risk> Levine and Crimmins conducted a study on predicting 10-year mortality rates using biological age, while Brown and McDaid conducted surveys and studies on the effects of factors such as chronological age, education level, sex, income, marital status, occupation, race, religion, smoking, alcohol consumption, physical activity, and obesity on adult mortality rates.

[0010] On the other hand, there has also been research into a model that assesses the risk of death by constructing a logistic regression model based on nine factors, including gender, smoking status, chronological age, and underwriting class.

[0011] In South Korea, a model was constructed to measure biological age using health examination data from a large Korean population, and then a Cox regression model was used to study the effect that biological age, when measured more frequently than chronological age, had on mortality over a 17-year period.

[0012] Currently, biological age measurement models published in the form of papers or patents present only one numerical value, such as an individual's biological age = 55.7 years. However, the quantitative and qualitative analysis of this value is unclear and not objective, so it is necessary to show an individual's aging state in another form, such as a biological age probability spectrum / distribution, rather than a single numerical value. [Prior art documents] [Patent documents]

[0013] [Patent Document 1] Korean Patent Publication No. 2014-0126229

[0014] <SCI-level thesis on biological age measurement> Currently published models for measuring biological age (a)A new approach to the concept and computation of biological age 2006, Mechanisms of Ageing and Development (for Czechs)

[0015] Nonlinear modeling of biomarker influence (b)A method for identifying biomarkers of aging and constructing an index of biological age in humans 2007, Journal of Gerontology (Kyoto University, Japanese men)

[0016] Modeling using PCA analysis technique (R2=0.52) (c)Development of models for predicting biological age(BA) with physical, biochemical, and hormonal parameters 2008, Arch Gerontol Geriatr (Comprehensive biology, body, biochemistry, hormone age, for Koreans)

[0017] Multiple linear regression modeling (male R2 = 0.62, female R2 = 0.66) (d)Developing a biological age assessment equation using principal component analysis and clinical biomarkers of aging in Korean men 2009, Archives of Gerontology and Geriatrics (classified into normal, abnormal glucose, and diabetic patients by age group, Seoul National University, Korean men)

[0018] Modeling using PCA analysis technique (R2=0.581) (e)Development and Application of Biological Age Prediction Models with Physical Fitness and Physiological Components in Korean Adults 2012, Gerontology (classified into normal and obese patients by age group, Asan Medical Center, Seoul, Koreans)

[0019] Modeling using PCA analysis technique (male R2 = 0.638, female R2 = 0.672) (f) Analysis of the influence of biological age on mortality Biological age as a useful index to predict seventeen-year survival and mortality in Koreans 2017, BMC Geriatrics (Analysis of the influence of biological age on death using data from a 17-year follow-up study of 550,000 Koreans)

[0020] Here, R2 means the coefficient of determination.

[0021] <Multiple Linear Regression Analysis Model: MLR> FIG. 3 shows the linear regression line. The linear regression line in FIG. 3 can be expressed as a linear regression equation such as Y=a+b*X. The points in Figure 3 show the measured coordinates X (health checkup value) and Y (age) for each individual, and the larger the health checkup value, the greater the tendency for chronological age to increase. When this is expressed using a linear regression model, the larger the health checkup value, the greater the influence of increasing age. (The quantitative influence of health check values ​​on age increase is the slope of the linear regression equation)

[0022] In other words, the outline of the biological age prediction model using a linear regression model is to consider biological age, which is estimated to exist somewhere in the increase / decrease relationship between health checkup results and age (more precisely, chronological age), as the Y value in the linear regression equation.

[0023] The multiple linear regression analysis model can be expressed by the following equation 1.

[0024]

number

[0025] The above-mentioned equation 1 indicates the linear influence of the independent variables on chronological age, with the dependent variable (Y) being chronological age and the three variables of BMI, SBP, and HDL being independent variables. Here, a1, a2, and a3 are regression coefficients, which respectively indicate the influence of BMI, SBP, and HDL on chronological age. and a0 is the regression constant (intercept or regression constant).

[0026] Y calculated by the above formula 1 is a value calculated when BMI, SBP, and HDL measurement values ​​are input, and the essence of the MLR model is to consider this value as biological age.

[0027] Such a multiple linear regression model (MLR) has the following problems. In young people, BA (biological age) is predicted to be higher (overestimated) than CA (chronological age), while in older people, BA is predicted to be lower (underestimated).

[0028] This is presumably due to properties that the data possess, but the exact mechanism is unknown.

[0029] FIG. 4 is a graph showing the relationship between chronological age (X) and biological age (Y), illustrating an example of over (under) estimation of a multiple linear regression model. A contradiction exists in that biological age (BA) is dependent on chronological age (CA) (the dependent variable) on health examination items.

[0030] That is, chronological age (CA) is not a medical item but is dependent on calendar time.

[0031] In particular, if the correlation between the health check items and chronological age (CA) is "1", the health check items themselves are useless. (Source: Ingram, 1988) This means that there is a contradiction in the assumptions made when establishing the model.

[0032] Next is a paper that mentions the problems with multiple linear regression models. (A) 2008 Linear Regression Model - MLR Model Development of models for predicting biological age(BA) with physical, biochemical, and hormonal parameters

[0033] (b) 2009 Seoul National University Hospital model - PCA model Developing a biological age assessment equation using principal component analysis and clinical biomarkers of aging in Korean

[0034] (c)2011 Seoul Asan Hospital Model - PCA Model Development and Application of Biological Age Prediction Models with Physical Fitness and Physiological Components in Korean Adults

[0035] (d) 2010 Comparison paper between biological age models An empirical comparative study on biological age estimation algorithms with an application of Work Ability Index(WAI)

[0036] <Explanation of the principal component analysis model; PCA> Principal Component Analysis (PCA) is a method for analyzing data sets. As shown in Figure 5, this method involves analyzing common characteristics exhibited by multiple variables v1 to v5 and finding a small number of independent factors (Factor 1, Factor 2) that can represent them.

[0037] For example, by performing PCA analysis using five variables, SBP, DBP, HDL, LDL, and TG, two independent factors, "blood pressure factors" and "cholesterol factors," can be extracted.

[0038] PCA is applied to a large number of health check variables (BMI, WST, SBP, DBP, AST, ALT, GGTP, HDL, LDL, TG, vital capacity, etc.) to extract a "single factor" that exists in common among these variables.

[0039] In this way, it is analyzed that there is a considerable amount of correlation between a factor extracted through PCA and chronological age. (Pearson's correlation coefficient 0.8)

[0040] Therefore, the core of the PCA biological age prediction model is to determine the "one factor" extracted by the PCA method as the "biological age" that indicates a person's actual aging state. Next is a biological age prediction model using PCA.

[0041] (a) 2009 Seoul National University Hospital model - PCA model Developing a biological age assessment equation using principal component analysis and clinical biomarkers of aging in Korean men

[0042] (b) 2011 Seoul Asan Hospital model - PCA model Development and Application of Biological Age Prediction Models with Physical Fitness and Physiological Components in Korean Adults

[0043] (c)2007 Japanese Model - PCA Model A Method for Identifying Biomarkers of Aging and Constructing an Index of Biological Age in Humans PCA

[0044] Characteristics of biological age prediction models using PCA Unlike multiple regression analysis, PCA analysis does not distinguish between dependent and independent variables. In other words, if there are five health check items, PCA is a method for selecting elements (principal components) that appear in common in the five values.

[0045] In Figure 5, when we look at the positions of the five variables on the coordinate system, we can see that v1 to v3 and v4 to v5 belong to two different clusters, which means that the five variables can be explained as two factors.

[0046] In the end, five variables are entered as input values, but the variables used to predict actual biological age (BA) are factor 1 and factor 2.

[0047] Here, the actual biological age prediction model uses only the single most influential factor.

[0048] Unlike the multiple linear regression analysis (MLR) model, the biological age prediction model using PCA does not use chronological age (CA) as a dependent variable, but the extracted factors showing the greatest influence have units such as age (e.g., 1 year, 2 years), and in order to correct bias in the prediction of biological age (BA), chronological age (CA) is entered into the biological age (BA) prediction model as an independent variable.

[0049] The PCA model can be summarized as follows:

[0050]

number

[0051] Here, BA represents biological age, X1 represents one principal component factor extracted via PCA, CA represents chronological age, F represents a conversion function using X1 as the input variable, and G represents a conversion function using CA as the input variable.

[0052] That is, the biological age means a numerical value calculated by multiplying the PCA principal component factors and the chronological age by their respective weighted values and then adding them together.

[0053] <Disadvantages of the PCA Model> Since the principal components extracted via the PCA model have a very high correlation with chronological age, the claim that this represents the biological age is merely a subjective opinion of the researcher.

[0054] In addition, in order to use the factors extracted via PCA as a variable (biological age) with the unit of "age", a conversion function using "chronological age" as a parameter is introduced, which is not objectively proven and is merely a simple idea of the researcher.

[0055] Another reason for including "chronological age" as a parameter in the biological age model is that before using "chronological age" as a parameter, the same phenomenon of overestimation in the younger population and underestimation in the older population occurs, similar to the MLR model.

[0056] Korean Patent Publication No. 2014 - 0126229, "Method and System for Generating a Biological Age Calculation Model and Method and System for Calculating the Biological Age Thereof", provides a method for calculating biological age using the PCA biological age prediction model.

Summary of the Invention

Problems to be Solved by the Invention

[0057] In a domestic environment where the aging population is rapidly increasing, a method is required to predict the aging state of each individual from a preventive perspective in order to live a healthier life for a long time.

[0058] The present invention aims to provide a method and system for generating a personalized biological age prediction model that takes into account the fact that aging mechanisms differ between men and women or chronological age bands, constructs a biological age prediction model for each gender and chronological age band, and predicts biological age according to the biological age prediction model for each age band.

[0059] In addition, the present invention aims to provide a personalized biological age prediction model and service system that can provide biological age information that can be analyzed more objectively and clearly by showing an individual's aging state in the form of a biological age probability spectrum / distribution, rather than simply presenting a numerical value of biological age (e.g., 55 years old). [Means for solving the problem]

[0060] Currently, biological age measurement models published in papers or patents present only one numerical value, such as an individual's biological age = 55.7 years. However, the quantitative and qualitative analysis of this value is not objective and is unclear. Therefore, it is necessary to show an individual's aging state in the form of a biological age probability spectrum / distribution, rather than a single numerical value.

[0061] The present invention differs from conventional biological age prediction models (MLR, PCA) in that it does not directly predict biological age using health check data, but instead calculates the "excess aging factor (i.e., Δ)" that is not explained by chronological age through health check data.

[0062] The present invention seeks to develop multiple biological age measurement models that behave differently depending on gender and chronological age band, since aging mechanisms are expected to differ between men and women or chronological age bands.

[0063] The present invention seeks to predict biological age using a statistical model that takes into account the distribution of differences in health checkup values ​​measured from an individual when compared to values ​​representative of people of the same chronological age (e.g., average body mass index, average blood pressure, etc.).

[0064] The method for generating a personalized biological age prediction model of the present invention includes the steps of: An age interval setting process for setting an age interval (x to y) to be used as training data to generate a binary logistic regression model; A binary logistic regression model generation process for dividing the training data into two groups, an under-age group (UAGm) and an over-age group (OAGm), for each age unit in the age range set in the age range setting process, and generating a binary logistic regression model (Mx~My) for each age unit; an age prediction probability calculation process for calculating the probability (Pm) of being predicted as being in the over-age group (OAGm) for each individual sample subject according to a binary logistic regression model; A cutoff extraction process in which the underage group (UAGm) and the overage group (OAGm) are set as dichotomous response variables, the probability (Pm) predicted for the overage group (OAGm) is set as a predictor variable, and a cutoff (Cm) is extracted by a receiver operating characteristic curve (ROC) analysis; An age prediction probability correction process in which a cutoff (Cm) is applied (Pm-Cm) to the probability (Pm) predicted to be in the over-age group (OAGm) to calculate the excess probability (Dm) predicted to be in the over-age group (OAGm); an excess age calculation process for calculating an individual's excess aging by calculating a weighted average (Δi) of the over-age group (OAGm) and the predicted excess probability (Dm) obtained by the age prediction probability correction process; and a biological age calculation step of calculating biological age by adding the individual excess age calculated by the excess age calculation step to chronological age.

[0065] The training data in the binary logistic regression model generation process is configured according to health check item information, and the method further includes a health check item information setting process for querying, adding, deleting, and setting the health check item information used as the training data.

[0066] The method may further include a condition information setting step for setting gender condition information for training data in the binary logistic regression model generating step.

[0067] In the process of calculating the overage, the overage for each individual is as follows: It is characterized in that it is calculated by multiplying the Dm (m=26, ..., 75) calculated for each individual by the relevant age (=m) and adding up all the results, then averaging the results.

[0068] The personalized biological age prediction model generation system of the present invention comprises: A health checkup data collection means for collecting health checkup data provided from the health checkup system and storing and managing the data in a data storage means; a training data setting means for determining effective training data from medical examination data provided by a medical examination data collecting means according to the set training data reference age range (x to y) and medical examination item information; a binary logistic regression model generating means for generating a binary logistic regression model (Mx to My) for each age unit within the age interval (x to y) set for the training data set by the training data setting means; an age prediction probability calculation means for calculating a probability (Pm) of being predicted as being in the over-age group for each individual of the training data according to the binary logistic regression model generated by the binary logistic regression model generation means; A cutoff extraction means for extracting a cutoff (Cm) through a receiver operating characteristic curve analysis by setting an underage group (UAGm) and an overage group (OAGm) as dichotomous response variables and setting a probability (Pm) predicted to be in the overage group (OAGm) as a predictor variable; an age prediction probability correction means for calculating an excess probability (Dm) predicted for an individual over-age group (OAGm) by applying a cutoff (Cm) (Pm-Cm) to the probability (Pm) predicted for the over-age group (OAGm) calculated by the age prediction probability calculation means, and correcting the probability (Pm) predicted for the over-age group (OAGm) calculated by the age prediction probability calculation means; an excess age calculation means for calculating an individual's excess aging by calculating a weighted average (Δi) of the over-age group (OAGm) and the predicted excess probability (Dm) calculated through the age prediction probability correction means; a biological age calculation means for calculating a biological age from a chronological age using the individual excess age obtained by the excess age calculation means; The medical examination data collection means collects medical examination data, and the training data set by the training data setting means stores and manages the medical examination data.

[0069] The training data setting means further includes a user setting means for providing a process for a user to inquire about and set the age range and health check item information of the training data setting means.

[0070] The training data setting means further includes a user setting means for providing a process for enabling a user to set condition information for determining training data, the condition information being gender information.

[0071] The medical examination item information of the training data setting means is It is characterized by being composed of health insurance checkup item data, including physical examination indicators such as body mass index, waist circumference, systolic blood pressure, and diastolic blood pressure, as well as blood test indicators such as three liver values ​​(AST, ALT, γ-GTP), creatinine, three cholesterol values ​​(HDL, LDL, TG), fasting blood sugar, and hemoglobin. Effect of the Invention

[0072] The present invention develops a biological age prediction model by utilizing high-quality, large-scale health checkup data already accumulated by the National Health Insurance Service, thereby reducing the cost and time required to separately build and research data to develop a biological age prediction model.

[0073] In addition, taking into account that the degree of aging differs depending on gender and age group, the present invention uses health checkup data to calculate each individual's excess age using relative values ​​for each individual based on gender and age group, and uses this as weighted value information to predict biological age, thereby generating a more reliable personalized biological prediction model. [Brief description of the drawings]

[0074] [Figure 1] FIG. 1 is a diagram showing an example of a data distribution showing the correlation between chronological age and systolic blood pressure. [Diagram 2] FIG. 1 illustrates an example of a data distribution showing the correlation between chronological age and hemoglobin. [Diagram 3] FIG. 1 is a diagram showing linear regression lines in a multiple linear regression analysis model (MLR). [Figure 4] 1 is a graph showing the relationship between chronological age (X) and biological age (Y). [Diagram 5] FIG. 1 is a diagram showing a biological age prediction model using Principal Component Analysis (PCA). [Figure 6] 1 is a flowchart showing the steps of a method for generating a personalized biological age prediction model according to the present invention. [Figure 7] FIG. 2 is a diagram illustrating a process of generating a binary logistic regression model in the present invention. [Figure 8] 1 is a graph showing probability values ​​(Pm) obtained according to a binary logistic regression model in the present invention. [Figure 9] 1 is a chart showing cutoff values ​​extracted through a cutoff extraction process in the present invention. [Figure 10] 1 is a graph showing the overage group (OAGm) and the predicted exceedance probability (Dm) obtained through an age prediction probability correction process in the present invention. [Figure 11] FIG. 2 is a diagram showing an example of an individual overage profile in the present invention. [Figure 12] 1 is a flow chart illustrating an embodiment of a process for generating a model for predicting biological age in accordance with the present invention. [Figure 13] FIG. 1 is a block diagram showing a configuration of a personalized biological age model generation system according to the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0075] The technical feature of the method for generating a personalized biological age prediction model of the present invention is to calculate an "excess aging factor (Δ)" that is not explained by chronological age through health check data, and to use this to predict biological age.

[0076] The process of generating a personalized biological age model of the present invention is performed as follows. an age interval setting process for setting age intervals (x to y) to be used as training data for generating a binary logistic regression model; and a binary logistic regression model generating process for dividing the training data into two groups, an under-age group (UAGm) and an over-age group (OAGm), for each age unit, with each age unit being set as one unit in the age interval set in the age interval setting process, and generating a binary logistic regression model (Mx to My) for each age unit; an age prediction probability calculation process for calculating the probability (Pm) of being predicted as being in the over-age group (OAGm) for each individual sample subject according to a binary logistic regression model; A cutoff extraction process in which the underage group (UAGm) and the overage group (OAGm) are set as dichotomous response variables, the probability (Pm) predicted for the overage group (OAGm) is set as a predictor variable, and a cutoff (Cm) is extracted through ROC curve analysis; An age prediction probability correction process in which a cutoff (Cm) is applied (Pm-Cm) to the probability (Pm) predicted to be in the over-age group (OAGm) to calculate the excess probability (Dm) predicted to be in the over-age group (OAGm); an excess age calculation process for calculating an individual's excess aging by calculating a weighted average (Δi) of the over-age group (OAGm) and the predicted excess probability (Dm) obtained through the age prediction probability correction process; and a biological age calculation step of calculating biological age by adding the individual excess age calculated through the excess age calculation step to chronological age.

[0077] The biological age prediction model of the present invention can be defined as a multivariable binary logistic regression (MBLR), and its characteristics can be simplified and expressed as follows:

[0078] The biological age prediction model (MBLR) of the present invention; Biological age (BA) = chronological age (CA) + Δ Δ=f(BMI, SBP, ....., CA) Here, f(BMI, SBP, . . . ) represents an excess aging factor calculation function based on a binary logistic regression model using health checkup values ​​as input variables.

[0079] In contrast, the conventional MLR model and PCA model can be expressed as follows: MLR model: BA=a0+a1×BMI+a2×SBP+… PCA model: BA = F(BMI, SBP, …) + G(CA) The present invention thus configured is The technical feature of this system is that it is possible to calculate excess age (Δi) relative to chronological age (CA) when calculating biological age (BA). As shown in FIG. 6, (a) Age range setting process; (b) Binary logistic regression model generation process; (c) Age prediction probability calculation process; (d) Cut-off extraction process; (e) Age prediction probability correction process; (f) excess age calculation process; (g) comprising a biological age calculation process.

[0080] The age range setting process includes: A process for setting subjects of health insurance medical checkup data for using training data to calculate biological age, in which an age interval (x to y) is set to be used to calculate a binary logistic regression model.

[0081] In the embodiment of the present invention, the age range of 26 (x) to 75 (y) is set as the subjects of the health insurance medical checkup data.

[0082] The ages 26 and 75 are values ​​used for the characteristics of health insurance data. If the data is not health insurance data, x (26 years old) and y (75 years old) can be changed.

[0083] The binary logistic regression model generation process is a process for generating a binary logistic regression model for determining the probability (Pm) of predicting overage from two groups, which divides "chronological age" into two groups and generates a model capable of predicting one of the two groups (OAGm).

[0084] There are 50 age units that can be set within the set 26 to 75 years old range, and for each unit, training data for each health check item is divided into two groups: under age group (UAGm) and over age group (OAGm).

[0085] FIG. 7 illustrates the process of generating a binary logistic regression model. As shown in Figure 7, each unit age is divided into a group below that age (UAGm) and a group above that age (OAGm), and one of the two groups is selected as training data for each unit to generate a total of 50 binary logistic regression models.

[0086] For example, in age units of 26 years, a group under 26 and a group over 26 are set, and the under 26 group and the over 26 group are divided (0, 1) in units of the health check item data set as training data to generate a binary logistic regression model (M26) for predicting those over 26 in the age prediction probability calculation process. A binary logistic regression model (M26) is generated by dividing specific values ​​for each health check item into "0" for people under 26 and "1" for people over 26.

[0087] A binary logistic regression model (M26) is generated for each health check data item of health insurance checkup items, such as physical examination indicators such as body mass index, waist circumference, systolic blood pressure, and diastolic blood pressure, and blood test indicators such as three liver values ​​(AST, ALT, γ-GTP), creatinine, three cholesterol values ​​(HDL, LDL, TG), fasting blood sugar, and hemoglobin, by dividing the data into people under 26 years old and people over 26 years old.

[0088] That is, a binary logistic regression model is generated according to two groups, the under-age group (UAGm) and the over-age group (OAGm), which are the response variables on the Y-axis, and the training data (health check data) which are the predictor variables on the X-axis.

[0089] The method may further include a health check-up item information setting step so that the health check-up items used as training data can be added, deleted, and set as health check-up item information in accordance with inquiries.

[0090] Also, the method may further include a condition information setting step for setting condition information for the training data, and the condition information may be configured with gender information.

[0091] This allows separate biological age prediction models to be constructed for male and female genders. This process is repeated for ages 26 to 76 to generate a total of 50 binary logistic regression models (M26 to M75).

[0092] The age prediction probability calculation process is a process of calculating the probability (Pm) of being predicted as being in the over-age group (OAGm) for each individual according to the binary logistic regression models (M26 to M75) generated as described above.

[0093] The following Equation 3 shows the age prediction probability calculation process using the binary logistic regression model.

[0094]

number

[0095] Where: Y: Individual's aging status P(Y=OAGm): Probability to be predicted as OAGm Yi: ith individual's aging status i=1, 2, ..., : sample number m = 26(x), 27, …, 75(y); ages used in training data (chronological age observed in the traning data) CA: Chronological age Xk: kth independent variable βk: regression coefficient of kth independent variable p: number of independent variables

[0096] FIG. 8 is a chart showing the probability values ​​(Pm) determined according to the binary logistic regression model. In the chart of FIG. 8, the probability value "P45" is a probability value calculated using a binary logistic regression model (M45), and means a probability value predicted for those aged 45 or older.

[0097] For example, for a person with sample ID=1, the predicted probability of being 45 years or older (P45) is 0.655, and the predicted probability of being 75 years or older is 0.211.

[0098] In the age prediction probability calculation process, these probability values ​​are calculated for all people (samples) for all ages, 50 each (P26 to P75), to generate a graph as shown in FIG.

[0099] In other words, probability (Pm) values ​​are calculated for each individual for all age units. Here, as shown in FIG. 8, when the probability (P26) of being predicted as being in the over-age group (OAG26) is considered, it is found that the probability is 0.998, which is close to 1.

[0100] This is an absolute value, and due to the inaccuracies in predicting biological age relative to the probabilities (Pm) mentioned above, a more accurate prediction of biological age must be made using relative values.

[0101] Therefore, a cutoff value (Cm), which is a standard value for determining biological age, is necessary.

[0102] The cutoff extraction process is a process for obtaining a reference value for determining biological age through ROC (Receiver Operating Characteristic and Area Under the Curve) analysis of the probability values ​​(Pm) obtained for 50 models (M26 to M75) for all people aged 26 to 75. The underage group (UAGm) and overage group (OAGm) are set as dichotomous response variables, and the probability (Pm) predicted for the overage group (OAGm) is set as a predictor variable to perform ROC curve analysis to extract the cutoff (Cm).

[0103] This cutoff extraction process involves extracting the cutoff (Cm) at the point that maximizes Youden's J statistic, which means the cutoff extraction result that maximizes the sum of sensitivity and specificity.

[0104] FIG. 9 is a chart showing the cutoff values ​​extracted through the cutoff extraction process. For example, in the chart in Figure 9, C45 is the cutoff value calculated by the model (M45), and when the probability value is calculated to be 0.547 or more, it means that the person in question is predicted to belong to the age group of 45 years or older.

[0105] The age prediction probability correction process is a process of applying the cutoff (Cm) value obtained through the age prediction probability calculation process to the probability (Pm) predicted to be in the over-age group (OAGm) (Pm-Cm) to correct it to the excess probability (Dm) predicted to be in the over-age group (OAGm).

[0106] FIG. 10 is a graph showing the overage group (OAGm) and the predicted exceedance probability (Dm) obtained through the age prediction probability correction process.

[0107] In the chart of FIG. 10, D26 to D75 are the probability values ​​(P26 to P75) calculated for each individual, minus the cutoffs (C26 to C75) calculated through the ROC curve (Dm=Pm-Cm).

[0108] For example, the chronological age of a person with ID=1 is 35 years old, and D45, the probability that this person is predicted to be 45 years old or older, is "D45=0.108(P45-C45;0.655-0.547)".

[0109] Here, a (-) value can be considered to be below the age. The excess aging calculation process is a process of calculating an individual's excess aging to calculate biological age by calculating the weighted average (Δi) of the over-age group (OAGm) and the predicted excess probability (Dm) obtained through the above process.

[0110] The following Equation 4 shows a process of calculating the weighted average (Δi) for the overage group (OAGm) and the predicted exceedance probability (Dm).

[0111]

number

[0112] where N is the sample number i=1, 2, …, N Δi: weighted mean of (Pim-Cm) Cm: Cutoff (Cm) value obtained via the cutoff extraction means 150 (cutoff of Pm to predict individual's aging status from ROC curve analysis) In other words, the average of the values ​​obtained by multiplying Dm (m = 26, ..., 75) calculated for each individual by the relevant age (= m) and adding them all up is defined as each individual's "excess age."

[0113] Here, the individual excess age is calculated as the weighted average for the over-age group (OAGm) and the predicted excess probability (Dm). If there is an additional weight (Wm) to be applied, this can be applied to calculate the weighted average.

[0114] The following equation 5 shows a process of calculating the weighted average (Δi) for the overage group (OAGm) and the predicted exceedance probability (Dm).

[0115]

number

[0116] where N is the sample number i=1, 2, …, N Δi: weighted mean of (Pim-Cm) Cm: Cutoff (Cm) value obtained via the cutoff extraction means 150 (cutoff of Pm to predict individual's aging status from ROC curve analysis) Wm: Weight applied for the model to predict CA≧m

[0117] The biological age calculation step is a step of calculating biological age by adding the excess age calculated in the excess age calculation step to chronological age.

[0118] The present invention has a technical feature of generating a model (algorithm) for predicting biological age using health insurance medical checkup data.

[0119] In the present invention, the excess age (Δi) relative to the chronological age (CA) is calculated to enable prediction of biological age.

[0120] First, the subjects of health insurance medical checkup data are set for use as training data for determining biological age.

[0121] In the embodiment of the present invention, 26 to 75 years old is set as the training data age target (x to y), which is the age interval for obtaining the binary logistic regression model.

[0122] As described above, taking into consideration the characteristics of the health insurance medical checkup data, the age range from 26 to 75 years old is set as the age interval for obtaining a binary logistic regression model.

[0123] In addition, the method may further include a health check item information setting process for setting health check items to be used as training data as health check item information, allowing a user (administrator) to set health check items to be used as training data for biological age prediction.

[0124] 12 is a flow chart showing an embodiment of a model generation process for predicting biological age in the present invention. Next, an embodiment of the operation process will be described with reference to FIG.

[0125] First, we initialize the age used in the training data and set m=26 years old. The players will then be divided into an under-age group (UAG26) for those under 26 years old and an over-age group (OAG26) for those 26 years old or older based on their training data.

[0126] In other words, the health checkup data is divided into those under 26 years old and those over 26 years old. The sample subjects (people) of the health checkup data are checked for specific values ​​for each health checkup item, and samples (people) under 26 years old are set to the under age group (UAGm) "0," and samples (people) over 26 years old are set to the over age group (OAGm) "1," thereby generating a binary logistic regression model (M26) corresponding to the age of 26.

[0127] The binary logistic regression model is used to calculate the probability (Pm) of being considered over-age (OAGm) in two groups, and as mentioned above, it uses health check data for each health insurance checkup item, such as physical examination indicators such as body mass index, waist circumference, systolic blood pressure, and diastolic blood pressure, and blood test indicators such as three liver values ​​(AST, ALT, γ-GTP), creatinine, three cholesterol values ​​(HDL, LDL, TG), fasting blood sugar, and hemoglobin, and can be added or deleted as necessary to set as health checkup item information.

[0128] Thereafter, according to the binary logistic regression model (M26) generated as described above, the probability (P26) of being predicted as being in the over-age group (OAG26) for each individual is calculated using Equation 3 to obtain the age prediction probability.

[0129] That is, such age prediction probability indicates an individual's aging status, and indicates the probability to be predicted as OAGm.

[0130] Thereafter, as described above, in order to obtain the cutoff (Cm), which is the standard value for determining biological age, the underage group (UAGm) and overage group (OAGm) are set as dichotomous response variables, and the overage group (OAGm) and the predicted probability (Pm) are set as predictor variables. The cutoff (Cm) is extracted through ROC curve analysis, and the cutoff (C26) value for determining biological age is obtained through ROC curve analysis for the predicted probability (P26) of being 26 years old or older.

[0131] Then, the cutoff (C26) obtained as described above is applied to correct the age predicted probability.

[0132] In the age prediction probability correction process, the cutoff (C26) value obtained through the age prediction probability calculation process is calculated (P26-C26) from the probability (P26) predicted to be in the over-age group (OAG26) to obtain the excess probability (D26) predicted to be in the over-age group (OAG26).

[0133] In this way, a cutoff (C26) is applied to each individual to determine the overage group (OAG26) and the predicted exceedance probability (D26).

[0134] As described above, once the predicted exceedance probability (D26) for the over-age group (OAG26) is calculated for each individual (sample), return and set m = 27, and go through the process described above to calculate each binary logistic model (M27), the predicted probability (P27) for the over-age group (OAG27), the cutoff (C27), and the predicted exceedance probability (D27) for the over-age group (OAG27).

[0135] This process is repeated up to m=75 to obtain the overage group (OAG75) and predicted excess probability (D75) for each individual.

[0136] A total of 50 units can be set in the age range of 26 to 75 years old. For each age unit, the training data is divided into an under-age group (UAGm) and an over-age group (OAGm), and a binary logistic regression model is generated for the 50 models, as shown in FIG. 7.

[0137] As explained above using an example, a binary logistic regression model (M26) is generated by dividing training data such as physical test indicators such as body mass index, waist circumference, systolic blood pressure, and diastolic blood pressure, and blood test indicators such as three liver values ​​(AST, ALT, γ-GTP), creatinine, three cholesterol values ​​(HDL, LDL, TG), fasting blood sugar, and hemoglobin into people with values ​​under 26 years old and people with values ​​over 26 years old, and this process is used to generate binary logistic regression models (M27 to M75) for ages 27, 28, ..., 75.

[0138] According to the binary logistic regression model (M26-M75) generated as described above, the predicted probability (Pm) of being in the over-age group (OAGm) is calculated by calculating Pm (P26-P75) for all age units (m=26-75) as shown in Figure 8, which shows the individual aging status.

[0139] In the example above, this means that a person with sample ID=1 has a probability of 0.655 of belonging to the 45+ age group (P45), and a probability of 0.211 of belonging to the 75+ age group.

[0140] The cutoffs (C26 to C75) obtained as described above are values ​​extracted by ROC curve analysis, and mean that the cutoff (Cm) is extracted at the point where Youden's J statistic is maximized.

[0141] The over-age group (OAGm) and predicted excess probability (Dm) calculated through the age prediction probability calculation correction process are obtained by applying a cutoff (Cm) to the over-age group (OAGm) and predicted probability (Pm) obtained in the age prediction probability process, and D26 to D75 are calculated for each individual from 26 to 75 years old as shown in Figure 10.

[0142] This process is repeated up to m=75 to obtain the over-age group (OAG75) and predicted exceedance probability (D75) for each individual. Then, the weighted average (Δi) of the over-age group (OAGm) and predicted exceedance probability (Dm) obtained through the above process is calculated to obtain the individual's excess aging, which is used to calculate biological age.

[0143] The weighted average (Δi) of such an individual excess age can be calculated using Equation 4.

[0144] In other words, according to equation 4, the average of the values ​​obtained by multiplying Dm (m = 26, ..., 75) calculated for each individual by the relevant age (= m) and adding them all up is defined as each individual's "excess age."

[0145] The weighted average thus obtained can be converted into an individual's excess age and applied to the chronological age to obtain the biological age.

[0146] FIG. 11 is a diagram showing an example of an overage profile for each individual, in which the X-axis is set to the training data age target of 26 to 75, and the Y-axis is the overage group (OAGm) and the predicted overage probability (Dm), showing the overage group (OAGm) and the predicted overage probability (Dm) for each age target.

[0147] The present invention uses health insurance medical checkup data to obtain average information of information showing the degree of aging for each individual, and generates a model (algorithm) capable of predicting biological age based on the average information.

[0148] On the other hand, FIG. 13 shows the configuration of the above-mentioned personalized biological age model generating system of the present invention.

[0149] A health checkup data collection means 110 for collecting health checkup data provided from a health checkup system and storing and managing the collected data in a data storage means 190; a training data setting means for determining effective training data from the medical examination data collected by the medical examination data collecting means according to the set training data reference age range (x to y) and medical examination item information; a binary logistic regression model generating means 130 for generating a binary logistic regression model (Mx to My) for each age unit within the age interval (x to y) set for the training data set by the training data setting means 120; an age prediction probability calculation means 140 for calculating a probability (Pm) of being predicted as being in the over-age group (OAGm) for each individual of the training data according to the binary logistic regression model generated by the binary logistic regression model generation means 130; A cutoff extraction means 150 for extracting a cutoff (Cm) through an ROC curve analysis by setting an underage group (UAGm) and an overage group (OAGm) as dichotomous response variables and setting a probability (Pm) predicted to be in the overage group (OAGm) as a predictor variable; an age prediction probability correction means 160 for calculating an excess probability (Dm) predicted as an individual over-age group (OAGm) by applying a cutoff (Cm) (Pm-Cm) to the probability (Pm) predicted as an over-age group (OAGm) calculated through the age prediction probability calculation means 140, and correcting the probability (Pm) predicted as an over-age group (OAGm) calculated by the age prediction probability calculation means 140; an excess age calculation means 170 for calculating an individual's excess aging by calculating a weighted average (Δi) for the over-age group (OAGm) and the predicted excess probability (Dm) calculated through the age prediction probability correction means 160; a biological age calculation means 180 for calculating a biological age from a chronological age using the individual excess age obtained by the excess age calculation means 170; and data storage means 190 for storing and managing the health checkup data collected by the health checkup data collection means 110 and the training data set via the training data setting means 120.

[0150] The technical feature of the personalized biological age prediction system of the present invention is that it sets training data from health checkup data provided by a health checkup system, extracts individual overage information from the training data, and predicts biological age.

[0151] a biological age prediction model generation system for generating a personalized biological age model by receiving medical examination data from a medical examination system, In the biological age prediction model generation system, The medical examination data collecting means 110 is a means for collecting medical examination data provided from a medical examination system, and is a means for storing and managing the collected medical examination data in the data storing means 190.

[0152] The training data setting means 120 is a means for setting training data for generating a biological age prediction model, and is a means for determining effective training data for the binary logistic regression model generating means from the medical checkup data stored in the data storage means 190 according to the set training data reference age range (x to y) and medical checkup item information.

[0153] The binary logistic regression model generating means 130 is a means for generating a binary logistic regression model (Mx to My) for each age unit within the age range set for the training data set by the training data setting means 120, This is a means of dividing the training data for each age unit into two groups, an under-age group (UAGm) and an over-age group (OAGm), with each age unit being one unit, and using the under-age group (UAGm) and over-age group (OAGm) and the training data (health check data) as response variables to generate a binary logistic regression model (Mx~My) for each age unit.

[0154] The age prediction probability calculation means 140 is a means for calculating the probability (Pm) of an individual being predicted to be in the over-age group (OAGm) according to the 50 binary logistic regression models generated by the binary logistic regression model generation means 130.

[0155] The cutoff extraction means 150 is a means for extracting a cutoff (Cm) for correcting the probability (Pm) predicted to be in the over-age group (OAGm) calculated by the age prediction probability calculation means 140, and is a means for setting the under-age group (UAGm) and the over-age group (OAGm) as dichotomous response variables, setting the probability (Pm) predicted to be in the over-age group (OAGm) as a predictor variable, and extracting the cutoff (Cm) through ROC curve analysis.

[0156] The age prediction probability correction means 160 is a means for correcting the probability (Pm) predicted to be in the over-age group (OAGm) calculated via the age prediction probability calculation means 140, and is a means for correcting the probability (Pm) predicted to be in the over-age group (OAGm) calculated by applying a cutoff (Cm) to the probability (Pm) predicted to be in the over-age group (OAGm) (Pm-Cm) to calculate an excess probability (Dm) predicted to be in the individual over-age group (OAGm) and calculating the probability (Pm) predicted to be in the over-age group (OAGm) calculated by the age prediction probability calculation means 140.

[0157] The excess age calculation means 170 is a means for calculating an individual excess age in order to obtain a biological age, and is a means for calculating an individual excess age by calculating a weighted average (Δi) for the over-age group (OAGm) obtained via the age prediction probability correction means 160 and the predicted excess probability (Dm).

[0158] The biological age calculation means 180 is a means for calculating biological age from chronological age by using the individual excess age obtained via the excess age calculation means 170.

[0159] The operation of the system of the present invention having such a configuration will be described below. The medical checkup data collection means 110 collects the medical checkup data provided from the medical examination system and stores it in the data storage means 190 .

[0160] The training data setting means 120 sets training data for obtaining a binary logistic regression model from the medical examination data stored in the data storage means 190 .

[0161] The training data setting means 120 determines training data for the set age range (x to y) and health check items.

[0162] In the embodiment of the present invention, health insurance medical checkup data is used, and an age range is set from 26 years old (x) to 75 years old (y).

[0163] The training data setting means 120 may further include a user setting means for providing a process for a user (administrator) to inquire about and reset the age range and health check item information.

[0164] Also, the training data setting means 120 may further include a user setting means for providing a process for allowing a user to set condition information for determining training data.

[0165] The condition information may be configured as gender information, and a biological age prediction model according to gender may be configured by setting gender information.

[0166] Then, in the binary logistic regression model generating means 130, 50 age units are set within the age range of the training data setting means 120, and the training data for each unit is divided into two groups, an under-age group (UAGm) and an over-age group (OAGm), to generate a binary logistic regression model.

[0167] This is the process for generating a binary logistic regression model to determine the probability (Pm) of being considered overage in the two groups.

[0168] With m = 26 years of age, we set a group under 26 (UAG26) and a group over 26 (OAG26), and divide the training data into samples (people) under 26 as 0 and samples (people) over 26 as 1 to generate a binary logistic regression model (M26).

[0169] In other words, a binary logistic regression model (M26) is generated by dividing training data for health insurance checkup items, such as physical examination indicators such as body mass index, waist circumference, systolic blood pressure, and diastolic blood pressure, and blood test indicators such as three liver values ​​(AST, ALT, γ-GTP), creatinine, three cholesterol values ​​(HDL, LDL, TG), fasting blood sugar, and hemoglobin, into people under 26 years old and people over 26 years old.

[0170] In other words, a binary logistic regression model is generated by using the two groups, the under-age group (UAGm) and the over-age group (OAGm), as the response variables on the Y-axis and the training data (health checkup data by health checkup item) as the predictor variables on the X-axis. This process is repeated for ages 26 to 76 to generate a total of 50 binary logistic regression models (M26 to M75).

[0171] Once the binary logistic model is generated as described above, the probability (Pm) of being predicted as being in the over-age group (OAGm) for each individual is calculated according to the binary logistic regression models (M26 to M75) generated as described above.

[0172] The probability (Pm) of being predicted to be in the overage group (OAGm) is information for determining an individual's overage age in order to predict biological age, and can be calculated by Equation 3.

[0173] As shown in FIG. 8, an individual probability value (Pm) can be obtained according to a binary logistic regression model.

[0174] For example, this means that a person with sample ID=1 has a probability of belonging to the 45 years old or older group (P45) of 0.655, and a probability of belonging to the 75 years old or older group of 0.211.

[0175] Meanwhile, the cutoff extraction means 150 extracts a cutoff (Cm) for the probability (Pm) of being predicted to be in the over-age group (OAGm) for each individual by ROC curve analysis.

[0176] The cutoff (Cm) is a standard value for determining biological age. The underage group (UAGm) and overage group (OAGm) are set as dichotomous response variables, and the probability (Pm) of being predicted as being in the overage group (OAGm) is set as a predictor variable, and a cutoff (Cm) value as shown in FIG. 9 can be obtained by performing ROC curve analysis.

[0177] Thereafter, the age predicted probability correcting means 160 uses the cutoff (Cm) value determined by the cutoff extraction means 150 to correct the probability (Pm) of being predicted as being in the over-age group determined by the age predicted probability calculating means 140 .

[0178] Such age prediction probability correction is performed by applying (Pm-Cm) the cutoff (Cm) value calculated by the age prediction probability calculation means 140 to the probability (Pm) predicted to be in the over-age group (OAGm) to calculate the excess probability (Dm) predicted to be in the over-age group (OAGm). As shown in FIG. 10, the excess probability (Dm) predicted to be in the over-age group (OAGm) corrected for each individual can be obtained.

[0179] According to FIG. 10, the chronological age of the person with ID=1 is 35 years old, and when calculated using the D45 model, D45, which is the predicted probability that this person belongs to the group of 45 years old or older, is as follows: D45=0.108 (P45-C45; 0.655~0.547). Here, a (-) value can be considered to be below the age.

[0180] The excess age calculation means 170 calculates the weighted average (Δi) by equation 4 for the over-age group (OAGm) and the predicted excess probability (Dm) to obtain the excess age for each individual.

[0181] In this case, the overage age for each individual is calculated as the weighted average for the overage group (OAGm) and the predicted overage probability (Dm). If there is an additional weight (Wm) to be applied, this can be applied to calculate the weighted average as shown in Equation 5.

[0182] The biological age calculation means calculates the biological age (BA=CA+Δi) from the chronological age using the excess age calculated by the excess age calculation means.

[0183] According to the present invention, by calculating the excess age relative to the chronological age from the health insurance medical checkup data and making it possible to predict the biological age from this, it is possible to provide a more reliable biological age. [Industrial Applicability]

[0184] The present invention develops a biological age prediction model by utilizing high-quality, large-scale health checkup data accumulated by the National Health Insurance Service, and is a technology that can be widely used in the medical and statistical analysis industries to realize its practical and economic value.

Claims

1. A personalized biological age prediction model generation system for generating a biological age prediction model from health checkup data collected from a health checkup system, An age interval setting process of a training data setting means (120) for setting an age interval (x to y) used as training data to generate a binary logistic regression model; a binary logistic regression model generating step of a binary logistic regression model generating means (130) for dividing training data into two groups, an under-age group (UAGm) and an over-age group (OAGm), for each age group set in the age group setting step, and generating a binary logistic regression model (Mx to My) for each age group; an age prediction probability calculation step of an age prediction probability calculation means (140) for calculating a probability (Pm) of being predicted as an over-age group (OAGm) for each individual sample subject according to a binary logistic regression model; A cutoff extraction process of a cutoff extraction means (150) for extracting a cutoff (Cm) by ROC curve analysis by setting an under-age group (UAGm) and an over-age group (OAGm) as dichotomous response variables and setting a probability (Pm) predicted to be in the over-age group (OAGm) as a predictor variable; an age prediction probability correction step of an age prediction probability correction means (160) for calculating an excess probability (Dm) predicted to be in the over-age group (OAGm) by applying a cutoff (Cm) to the probability (Pm) predicted to be in the over-age group (OAGm) (Pm-Cm); an excess age calculation step of an excess age calculation means (170) for calculating an individual's excess age by calculating a weighted average (Δi) of the over-age group (OAGm) and the predicted excess probability (Dm) obtained by the age prediction probability correction step; A method for generating a personalized biological age prediction model, comprising: a biological age calculation process in a biological age calculation means (180) for calculating biological age by adding the individual excess age calculated by the excess age calculation process to chronological age.

2. The training data in the binary logistic regression model generation process is configured according to medical examination item information; The health check item information is The method for generating a personalized biological age prediction model as described in claim 1, characterized in that the health insurance checkup item data includes physical examination indices including body mass index, waist circumference, systolic blood pressure, and diastolic blood pressure, and blood test indices including three liver values ​​(AST, ALT, γ-GTP), creatinine, three cholesterol values ​​(HDL, LDL, TG), fasting blood glucose, and hemoglobin.

3. The training data in the binary logistic regression model generation process is configured according to medical examination item information; The method for generating a personalized biological age prediction model according to claim 1 or 2, further comprising a health check item information setting process for querying, adding, deleting, and setting health check item information used as training data.

4. The method of claim 1 , further comprising a condition information setting step for setting condition information for training data in the binary logistic regression model generation step.

5. The method for generating a personalized biological age prediction model according to claim 4 , wherein the condition information in the condition information setting step is gender information.

6. In the binary logistic regression model generation process, The binary logistic regression model (Mx to My) is The method for generating a personalized biological age prediction model as described in claim 1, characterized in that each age unit in the set age range is treated as one unit, the training data is divided into two groups, an under age group (UAGm) and an over age group (OAGm), the two groups, the under age group (UAGm) and the over age group (OAGm), are used as response variables, and the training data is used as a predictive variable, and a personalized biological age prediction model is generated for each age unit.

7. In the age prediction probability calculation process, the probability (Pm) of a sample subject being predicted to be in the over-age group (OAGm) according to a binary logistic regression model is calculated using the following Equation 1: [0010] Where: Y: Individual's aging status P(Y=OAGm): Probability to be predicted as OAGm Yi: ith individual's aging status i = 1, 2, ...: sample number m = 26(x), 27, ..., 75(y); ages used in training data (chronological age observed in the training data) CA: Chronological age Xk: kth independent variable βk: regression coefficient of kth independent variable p: number of independent variables, The method for generating a personalized biological age prediction model according to claim 1, characterized in that the method is performed by:

8. In the process of calculating the overage, the overage for each individual is as follows: The following formula 2 shows the average of the sum of the Dm (m = 26, ..., 75) calculated for each individual multiplied by the relevant age (= m): [0025] Here, N: sample number i=1, 2, ..., N Δi: weighted mean of (Pim-Cm) Cm: Cutoff (Cm) value obtained through the age prediction probability calculation process (cutoff of Pm to predict individual's aging status from ROC curve analysis), The method for generating a personalized biological age prediction model according to claim 1 , wherein the calculation is performed as follows:

9. In the process of calculating the overage, the overage for each individual is as follows: The weighted average is calculated for the overage group (OAGm) and the predicted exceedance probability (Dm). By applying the additional weight (Wm), the weighted average is calculated as follows: [0030] Here, N: sample number i=1, 2, ..., N Δi: weighted mean of (Pim-Cm) Cm: Cutoff (Cm) value obtained through the age prediction probability calculation process (cutoff of Pm to predict individual's aging status from ROC curve analysis) Wm: weight applied for the model to predict chronological age >= m; The method for generating a personalized biological age prediction model according to claim 1 , wherein the calculation is performed by:

10. A health checkup data collection means (110) for collecting health checkup data provided from a health checkup system and storing and managing the data in a data storage means; a training data setting means (120) for determining effective training data from medical examination data provided by a medical examination data collecting means (110) according to a set training data reference age range (x to y) and medical examination item information; a binary logistic regression model generating means (130) for generating a binary logistic regression model (Mx to My) for each age unit within the age range (x to y) set for the training data set by the training data setting means (120); an age prediction probability calculation means (140) for calculating a probability (Pm) of being predicted as being in the over-age group (OAGm) for each individual of the training data according to the binary logistic regression model generated by the binary logistic regression model generation means (130); A cutoff extraction means (150) for extracting a cutoff (Cm) through an ROC curve analysis by setting an under-age group (UAGm) and an over-age group (OAGm) as a dichotomous response variable and setting a probability (Pm) predicted to be in the over-age group (OAGm) as a predictor variable; an age prediction probability correction means (160) for calculating an excess probability (Dm) predicted to be an individual over-age group (OAGm) by applying a cutoff (Cm) (Pm-Cm) to the probability (Pm) predicted to be an over-age group (OAGm) calculated through the age prediction probability calculation means (140) and correcting the probability (Pm) predicted to be an over-age group (OAGm) calculated by the age prediction probability calculation means (140); an excess age calculation means (170) for calculating an individual's excess age by calculating a weighted average (Δi) for the over-age group (OAGm) and the predicted excess probability (Dm) calculated through the age prediction probability correction means (160); a biological age calculation means (180) for calculating a biological age from a chronological age using the individual excess age obtained through the excess age calculation means (170); A personalized biological age prediction model generation system comprising: a data storage means (190) for storing and managing health checkup data collected from a health checkup data collection means (110) and training data set by a training data setting means.

11. The personalized biological age prediction model generation system according to claim 10, further comprising a user setting means for providing a process for a user to inquire about and set the age range and health check item information of the training data setting means (120).

12. The personalized biological age prediction model generation system according to claim 10 or 11, further comprising a user setting means for providing a process for a user to set condition information for determining training data in the training data setting means (120).

13. The personalized biological age prediction model generating system according to claim 12 , wherein the condition information of the user setting means is gender information.

14. The binary logistic regression models (Mx to My) in the binary logistic regression model generation means (130) are The personalized biological age prediction model generation system of claim 10, characterized in that each age unit is set as one unit in the set age range, the training data is divided into two groups, an under age group (UAGm) and an over age group (OAGm), the two groups, the under age group (UAGm) and the over age group (OAGm), are used as response variables, and the training data is used as a predictive variable, and a personalized biological age prediction model generation system is generated for each age unit.

15. The medical examination item information of the training data setting means (120) is The personalized biological age prediction model generation system according to claim 10 or 11, characterized in that the health insurance checkup item data includes physical examination indices including body mass index, waist circumference, systolic blood pressure, and diastolic blood pressure, and blood test indices including three liver values ​​(AST, ALT, γ-GTP), creatinine, three cholesterol values ​​(HDL, LDL, TG), fasting blood glucose, and hemoglobin.

16. The age prediction probability calculation means (140) calculates the probability (Pm) of a sample subject being predicted to be in the over-age group (OAGm) according to a binary logistic regression model using the following Equation 4: [0045] Where: Y: Individual's aging status P(Y=OAGm): Probability to be predicted as OAGm Yi: ith individual's aging status i = 1, 2, ...: sample number m = 26(x), 27, ..., 75(y); ages used in training data (chronological age observed in the training data) CA: Chronological age Xk: kth independent variable βk: regression coefficient of kth independent variable p: number of independent variables, The personalized biological age prediction model generation system according to claim 10, wherein the system is performed by

17. The overage calculation means (170) calculates the following equation 5 for the probability (Dm) of being predicted as being in the overage group (OAGm): [0050] Here, N: sample number i=1, 2, ..., N Δi: weighted mean of (Pim-Cm) Cm: Cutoff (Cm) value obtained via cutoff extraction means (150) (cutoff of Pm to predict individual's aging status from ROC curve analysis), The personalized biological age prediction model generating system according to claim 10, wherein the weighted average (Δi) is calculated through the above to obtain the individual excess age.

18. The excess age calculation means (170) calculates the following equation 6 for the excess probability (Dm) predicted for the overage group (OAGm): [006] Here, N: sample number i=1, 2, ..., N Δi: weighted mean of (Pim-Cm) Cm: Cutoff (Cm) value obtained via cutoff extraction means (150) (cutoff of Pm to predict individual's aging status from ROC curve analysis) Wm: weight applied for the model to predict chronological age >= m; The personalized biological age prediction model generating system according to claim 10, wherein the weighted average (Δi) is calculated through the above to obtain the individual excess age.

Citation Information

Patent Citations

  • Human biological age measuring-calculating device and system

    CN108847284A

  • Integrated circuit module and operating method thereof

    JP1987065356A

  • Device for determining health condition

    JP2010026855A

  • Method for predicting health age

    JP2019145057A

  • Health condition diagnosis system

    JP2020017153A