Disease risk assessment method, disease risk assessment system, and health information processing device.

A data-driven method using unsupervised and supervised learning on health checkup data separates high-risk groups and quantifies disease risk, addressing the challenge of pre-symptomatic disease detection and management.

JP7837023B2Active Publication Date: 2026-03-30THE INSTITUTE OF PHYSICAL & CHEMICAL RESEARCH +2
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Current methods fail to accurately determine the risk of developing specific diseases at the health stage before symptoms appear, as they rely on genetic information influenced by environmental factors and lack temporal inference, making it difficult to achieve ultra-early health management.

Method used

A data-driven approach using unsupervised learning for clustering health checkup data to separate high-risk and low-risk groups, followed by supervised learning to quantify disease risk, excluding disease progression markers, enabling continuous risk assessment from the healthy stage to symptom onset.

Benefits of technology

Enables the identification of high disease risk before symptoms appear and quantifies future risk, allowing for ultra-early health management and intervention guidelines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007837023000001
    Figure 0007837023000001
  • Figure 0007837023000002
    Figure 0007837023000002
  • Figure 0007837023000003
    Figure 0007837023000003
Patent Text Reader

Abstract

To provide a disease risk evaluation method, a disease risk evaluation system and a health information processing device capable of detecting a potential tendency to develop a disease in a healthy stage in advance and quantizing a future disease risk to a target disease.SOLUTION: According to the present invention, a disease risk evaluation method, a disease risk evaluation system and a health information processing device include a step for classifying whether or not to be easily affected by a specific disease regardless of a degree of progress of a disease from a healthy stage to a stage after the incidence of a disease into groups, and a plurality of steps for further classifying the degrees of affection in a group determined to have a high affection risk, and achieves disease risk evaluation by changing types of data to be used for an analysis mainly based on data in each step.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a disease risk assessment method, a disease risk assessment system, and a health information processing device for determining whether a person has a high risk of suffering from a specific disease at the health stage.

Background Art

[0002] With the development of advanced medical care in recent years, the treatment and medical technology after the onset of diseases have advanced. Affected by this, while life expectancy has increased, the medical expenses of the entire nation have increased, and the financial burden has become a major social problem. In addition to physical diseases, there is also a need to control and improve mental health problems such as depression due to stress and unhealthy lifestyles. In order to reduce medical expenses and enable the nation to work in good health, that is, to extend the healthy life expectancy, it is necessary to manage the risk of contracting diseases at the health stage when the disease has not yet manifested, and to realize ultra-early health management so as not to approach the onset of the disease. For this purpose, it is necessary to know what diseases a person has a high risk of contracting before symptoms appear at the health stage. Indices (measurement value standards) used in health check-ups such as general medical examinations and disease diagnoses indicate the appearance of symptoms of the onset of the disease, and do not provide risk criteria at the health stage. Therefore, new indices effective at the health stage are required. Currently, genes and their mutations are used as indices for what diseases a person is likely to contract, but it is known that the expression of genes varies depending on environmental factors. Also, in many cases, there are not one but many gene mutations said to be related to a specific disease. Therefore, it is difficult to determine from only gene information which diseases a person has a high risk of contracting and whether they are approaching the onset stage based on the current health condition. By analyzing big data, it may be possible to determine which diseases an individual is susceptible to without using genetic information, thus identifying the risk of developing certain diseases during the healthy stage. Furthermore, if this is determined by measurements that change depending on health status, it may be possible to analyze what measurement values ​​reduce the risk of developing a disease. There is a need for methods that determine whether an individual is at high risk of developing a specific disease, regardless of whether they are in the healthy stage or have already developed the disease, based on their environment and measurement values, without relying on genetic information. In order to manage the risk of developing a disease and prevent it from progressing to full-blown illness during the healthy stage before the disease becomes apparent, we believe that the following requirements must be met to achieve ultra-early health management. One purpose is to provide a measure of health that is used in disease diagnosis and health checkups, allowing us to determine whether there is a high risk of contracting a particular disease before any signs of the disease appear. Another purpose is to provide information about the degree of the disease from the time the aforementioned signs appear until the disease develops. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2013-191020 [Disclosure of the Invention] [Problems that the invention aims to solve]

[0004] Patent Document 1 states that it estimates the state using self-organizing map technology, but it only describes very general unsupervised learning and does not explain the effectiveness of its application. Furthermore, since it only classifies health check data, it does not identify the risk of disease at different health stages, which is our goal. Methods that do not include temporal inference cannot be said to have achieved ultra-early health management technology that manages the risk of illness and prevents it from progressing. Furthermore, temporal verification is necessary. We are developing a technology that can determine the degree of risk of developing a specific disease in healthy individuals, using various environmental data including health checkup data, without relying on genetic information. Furthermore, we will quantify whether individuals, even in a healthy state, are approaching the risk of developing the disease. We devised a method and a system that achieve the two requirements mentioned above. (1) A measurement used in disease diagnosis or health checkups that can determine whether there is a high risk of contracting a particular disease before the onset of symptoms appears. This function can be realized through so-called data-driven analysis, and Patent Document 1 takes such an approach. Supervised learning methods cannot be used in the healthy stage because there are no biomarkers to indicate diagnosis or signs of illness. However, the unsupervised learning self-organizing map used in Patent Document 1 maps based on the presence or absence of signs, that is, whether the risk of a particular disease has manifested, and cannot be used to determine the magnitude of the risk of contracting a particular disease before signs appear, or whether one is approaching the onset of the disease before signs appear. In other words, Patent Document 1 merely states that it performed mapping using a self-organizing map already known to the public, and therefore does not fulfill the function of a method to achieve our objectives. Furthermore, it does not present an invention that would achieve our objectives. (2) The extent of the illness from before the appearance of the symptoms mentioned in the preceding paragraph until the onset of the disease is known. Furthermore, not only in Patent Document 1, but no method has been devised to determine the degree of risk of contracting a specific disease from a healthy stage based on health checkup data. In addition, recent studies have shown that genetic testing is heavily influenced by acquired changes in gene expression and environmental factors, making it important to determine the degree of disease risk from health checkup data.

[0005] Therefore, the present invention aims to provide a disease risk assessment method, a disease risk assessment system, and a health information processing device that can detect potential disease onset tendencies in advance during the health stage and quantify future disease risk for a target disease. [Means for solving the problem]

[0006] We devised a method to achieve our objective by taking multiple steps: first, dividing individuals into high-risk and low-risk groups, including those in the healthy stage before the onset of the disease, and then further dividing the high-risk group based on the degree of disease. We also considered and verified that each step could be achieved by changing the type of data used in the data-driven analysis. Specifically, we will perform unsupervised learning clustering to separate individuals into high-risk and low-risk groups by removing data that changes depending on the disease progression (biomarkers used to diagnose the disease). By removing data that changes depending on the disease progression, both those with advanced disease and those in a healthy stage can be classified into a single group based on their risk of developing the disease. Next, data whose values ​​change depending on the stage of disease progression is returned, and the progression (high or low risk of developing the disease or the appearance of symptoms) is categorized or quantified. This can be achieved using conventional supervised learning techniques. In conventional methods for determining disease progression, existing biomarkers were not applicable to the progression of all patients. However, by applying existing biomarkers to high-risk groups, it becomes possible to more accurately understand and manage the progression of the disease. Because this is a data-driven analysis, we cannot show the detailed mechanisms between individual data points and results. However, we verified the effectiveness of our method using publicly available data. This verification demonstrates that our devised method is an effective way to solve the problem. [Effects of the Invention]

[0007] As can be seen from the verification results, according to the present invention, it is possible to determine whether there is a high risk of contracting a specific disease even at a healthy stage where no symptoms are present. Furthermore, it is possible to obtain a continuous disease risk score from the healthy stage to the onset of the disease. [Brief explanation of the drawing]

[0008] [Figure 1] Figure showing the evaluation step of the disease risk assessment method according to an embodiment of the present invention [Figure 2] Figure showing an example of score display [Figure 3] Graph showing verification results [Figure 4] Conceptual diagram of the disease risk assessment method according to this embodiment [Figure 5] Figure showing the filter processing flow of data used when the target disease is cardiovascular disease [Figure 6] Parameter list shown in FIG. 5 [Figure 7] Figure showing the filter processing flow of data used when the target disease is diabetes [Figure 8] Parameter list shown in FIG. 7 [Figure 9] Figure showing the filter processing flow of data used when the target disease is depression [Figure 10] Parameter list shown in FIG. 9 [Figure 11] Process of generating cardiovascular disease subtypes [Figure 12] Parameter list used for generating cardiovascular disease subtypes [Figure 13] Cardiovascular disease subcategory analysis [Figure 14] Overview of glaucoma subcategory classification [Figure 15] Graph of diabetes progression rate analysis [Figure 16] Conceptual diagram showing the whole of the disease risk assessment system according to this embodiment [Figure 17] Figure showing the configuration related to the clustering process of the disease risk assessment system according to this embodiment [Figure 18] Figure showing the configuration related to the mapping process of the disease risk assessment system according to this embodiment [Figure 19] Figure showing the configuration related to the verification process of the disease risk assessment system according to this embodiment [Modes for carrying out the invention]

[0009] The following describes embodiments of the disease risk assessment method of the present invention. Figure 1 shows the evaluation steps according to this embodiment. In S1, data that may be related to health status is collected and compiled into a database. As with gene mutation analysis, it is best to collect as much data as possible. For example, when studying diabetes, all or some of the following data are used: white blood cell count, lymphocyte percentage, red blood cell count, platelet count, HDL cholesterol, creatinine, albumin, height, systolic blood pressure, and medical history. When studying cardiovascular disease, all or some of the following data are used: white blood cell count, lymphocyte percentage, red blood cell count, platelet count, HbA1c, creatinine, albumin, height, systolic blood pressure, and sleep duration. When studying depression, all or some of the following data are used: white blood cell count, lymphocyte percentage, HbA1c, total / HDL cholesterol, creatinine, albumin, height, systolic blood pressure, sleep duration, and medical history. As illustrated, a favorable evaluation can be obtained by using at least 10 data points for each disease. Some data items may be replaced with those exemplified above, and if some of the exemplified items are unavailable, other available data items can be used.

[0010] In step S2, data (features) that correlate with the level of disease are excluded. For example, if the target is diabetes, at least HbA1c, which is used to diagnose diabetes, is excluded. In this way, in step S2, data (features) that correlate with the level of disease are excluded. The step in S2 creates a database that does not contain data (features) that correlate with the level of disease.

[0011] In S3, a database that does not contain data (features) correlated with disease level is used to divide data into high-risk and low-risk groups. Semi-supervised clustering is suitable for separation in S3. Unsupervised clustering may also be used. By clustering the data to include cases where the disease has developed, it is possible to extract groups with a high risk of developing the disease.

[0012] In S4, data (features) correlated with the level of the disease are returned to the high-risk group. In other words, data that was excluded in the process of dividing patients into high-risk and low-risk groups (S3), such as HbA1c which was excluded in the case of diabetes, is returned.

[0013] Then, in S5, for groups at high risk of developing the disease, the disease level from the healthy stage to the post-symptomatic stage is quantified, including data (features) that correlate with the disease level. Supervised learning is suitable for S5. Supervised learning allows us to quantify data even in areas where actual data is unavailable. After processing up to S5 for a specific disease, processing from S1 to S5 is performed for the next target disease. In this way, all target diseases are classified into groups based on whether the risk of developing the disease is high or low, and the disease level is quantified. Once processing for all target diseases is complete, S6 creates individual subject scores. For groups with a high risk of disease, the scores are displayed, showing a quantified disease incidence level indicating the disease with a high risk. The scores can be displayed numerically, for example, on a scale from 0 to 100, and graphical methods for displaying the scores include bar charts or radar charts.

[0014] S6 allows individuals undergoing this test to learn about the diseases they should be aware of and the degree of their risk of developing them.

[0015] Figure 2 shows an example of how the score is displayed. This is an example of the display that should be output when this evaluation method and evaluation system are implemented in a health information processing device. It shows the requirements for realizing the aforementioned ultra-early health management: "diseases judged to have a high risk of contracting" and "the degree from before the appearance of symptoms to the onset of the disease."

[0016] Verification method Figure 3 shows the verification results; Figure 3(a) shows the verification results for group classification by S3, and Figure 3(b) shows the verification results for disease severity in the group classified as high risk. The model was trained using publicly available data (CDC (Centers for Disease Control and Prevention) NHANES 2013-2014) and validated using another publicly available dataset (CDC NHANES 2011-2012). Thus, validation was performed using a different dataset than the one used for training. In Figure 3(a), validation is indicated by dots.

[0017] The solid red line in Figure 3(a) shows the average incidence rate by age for individuals classified as high-risk after clustering data trained on Cardio-vasculature. It can be seen that the incidence rate is divided into high-risk and low-risk groups consecutively, from older age to younger age with a low incidence rate. This means that even at a young age when individuals are still healthy, they can be divided into high-risk and low-risk groups. This means that it is possible to predict the likelihood of developing cardiovascular disease in the future based on current environmental values ​​(measurements). The solid line overlaps with the data, indicating that the trained knowledge is valid for other data as well.

[0018] Figure 3(b) shows the disease severity (Disease Score) from the healthy stage to the post-symptomatic stage for the group classified as high-risk after separation, by returning the data excluded in the separation step and performing Supervised Learning. Previously, data-driven analyses primarily used biomarkers indicating symptoms used in diagnosis as the main explanatory variables, and therefore could not analyze the severity before symptoms appeared. However, in this invention, semi-supervised clustering is used in the pre-symptomatic stage, and supervised learning and data whose values ​​change depending on the disease severity are used in the stage when symptoms appear, so it is possible to show the disease severity (Disease Score) continuously from the healthy stage to the post-symptomatic stage.

[0019] Figure 4 is a more detailed conceptual diagram of the group classification (S1 to S3) in the disease risk assessment method of this embodiment. In the disease risk assessment method according to this embodiment, disease risk is estimated for estimated subjects in a healthy stage by clustering at least two categories of data from blood test data, physical measurement data, demographic data, medical interview data, and urine test data into at least two groups, and determining which group they belong to or are closest to. The data used excludes disease parameters that are used to diagnose the target disease or to determine the progression of the target disease.

[0020] As shown in Figure 4, in the disease risk assessment method according to this embodiment, the computer includes a learning data acquisition step S10 in which it acquires at least two categorical data, a filtering step S20 in which it removes specific parameters from the data, a learning step S30 in which it performs machine learning using the data from which the specific parameters have been removed, a mapping step S40 in which it displays the clustering results, and a display step S50 in which it displays the clustered groups and judgment results obtained in the learning step S30.

[0021] The filtering step S20 comprises a first filtering step S21 and a second filtering step S22. In the first filtering step S21, disease parameters used for diagnosing or determining the progression of a pre-defined target disease are excluded from the data. In the second filtering step S22, display parameters used for displaying the clustering results, one of the parameters that have a strong correlation with each other, and parameters that degrade the clustering performance are excluded. In learning step S30, the system learns parameters in a heuristic manner so that clustering based on disease risk separates the group into, for example, a low-risk group and a high-risk group. In the mapping processing step S40, for example, mapping is performed on two axes: disease risk rate and age distribution. In display step S50, for example, the X-axis represents age distribution and the Y-axis represents disease risk, and the low-risk group and the high-risk group are displayed two-dimensionally as line graphs.

[0022] The computer has a verification step S60 which verifies the groups clustered by the learning step S30. In the learning step S30, data from a past first predetermined period is used as learning data, and in the validation step S60, data from a second predetermined period prior to the first predetermined period is used as validation data. For example, CDC (Centers for Disease Control and Prevention) 2013-2014 data is used as learning data, and CDC 2011-2012 data is used as validation data. The validation data used in validation step S60 is filtered by the first filtering step S21 to remove disease parameters, and by the second filtering step S22 to remove display parameters used for displaying clustering results, or one of the parameters that have a strong correlation with each other, as well as parameters that degrade clustering performance. In display step S50, the agreement with the line graph is shown by plotting the low-risk group and the high-risk group.

[0023] The computer has a determination step S70 that determines which group the estimated subject belongs to or is closest to. The subject data of the estimated subjects used in the determination step S70 is filtered in the first filtering step S21 to remove disease parameters, and in the second filtering step S22 to remove display parameters used for displaying clustering results, one of the parameters that have a strong correlation with each other, and parameters that degrade the performance of clustering. In display step S50, the assessment results for the estimated target individuals are plotted, allowing for comparison with line graphs representing low-risk and high-risk groups, enabling users to determine their risk position and which group they are closer to. Furthermore, the distribution by age group allows for risk assessment over time.

[0024] Furthermore, it is preferable to normalize the parameters used in the learning step S30, verification step S60, and judgment step S70, particularly gender, age group, and questionnaire data, for example, by using the standard deviation (SD) value. In disease risk clustering, it is preferable to extract high-risk groups for the target disease, encompassing stages from the healthy stage to the onset and progression stages, and then grade these extracted groups according to their progression. Furthermore, for disease risk clustering, the Kernel k-means method or a custom kernel function can be used. For example, initialization (setting the center point) is performed on 40% of the training data with disease labels, and clustering of high-risk and low-risk groups is performed for each age group at the center point (each pre-disease category).

[0025] In the validation step S60, validation can be performed using the training data used in the learning step S30. The validation data is input to the constructed clustering model, and the results from the training data are compared with the error in the prevalence of disease risk for verification. Validation can also be performed using the past history of individuals who developed the disease.

[0026] In this way, by performing machine learning using data that excludes disease parameters used in diagnosing or judging the progression of the target disease, it is possible to detect potential disease tendencies in the healthy stage and quantify future disease risk for the target disease. Furthermore, by analyzing the lifestyles of high-risk and low-risk disease groups, it becomes possible to create applications that enable health promotion management and provide intervention guidelines to reduce disease risk.

[0027] Figure 5 shows the data filtering process used when the target disease is cardiovascular disease, and Figure 6 is a list of the parameters shown in Figure 5. As shown in Figure 5, when the target disease is cardiovascular disease, six parameters are removed in the first filtering step S21, and another six parameters are removed in the second filtering step S22.

[0028] In the first filtering step S21, among the parameters shown in Figure 6, total cholesterol and direct HDL-cholesterol, which are blood test data, are excluded as disease parameters. Also in the first filtering step S21, among the parameters shown in Figure 6, interview parameters regarding the current or past illnesses of the person estimated to have had a heart attack, coronary heart disease, angina pectoris, or congestive heart failure, which are interview data, are excluded as disease parameters.

[0029] In the second filtering step S22, among the parameters shown in Figure 6, the blood test data, segmented neutrophil percentage and epi-25-hydroxyvitamin D3, and the anthropometric data, BMI, are excluded. Segmented neutrophil percentage is excluded to improve clustering performance, epi-25-hydroxyvitamin D3 is excluded because it has a strong correlation with 25-hydroxyvitamin D3, and BMI is excluded because it has a strong correlation with mean sagittal abdominal diameter. Furthermore, in the second filtering step S22, among the parameters shown in Figure 6, the demographic data parameters of age and sex are excluded, and the questionnaire parameter "Have you eaten?" which is questionnaire data is excluded. Gender was included to improve clustering performance, and the question parameter "Did you not eat?" was included because it strongly correlates with the question parameter "Did you not have time to eat a balanced meal?".

[0030] Figure 7 shows the data filtering process used when diabetes is the target disease, and Figure 8 is a list of the parameters shown in Figure 7. As shown in Figure 7, when the target disease is diabetes, two parameters are removed in the first filtering step S21, and seven more parameters are removed in the second filtering step S22.

[0031] In the first filtering step S21, among the parameters shown in Figure 8, HbA1c, which is blood test data, is excluded as a disease parameter. Also in the first filtering step S21, among the parameters shown in Figure 8, the questionnaire parameters regarding the current or past illnesses of the person estimated to have diabetes, which is questionnaire data, are excluded as disease parameters.

[0032] In the second filtering step S22, among the parameters shown in Figure 8, erythrocyte folate, which is blood test data, and BMI, which is anthropometric data, are excluded. Erythrocyte folate is excluded to improve clustering performance, and BMI is excluded because it has a strong correlation with the mean abdominal sagittal diameter. Furthermore, in the second filtering step S22, demographic data parameters such as age and sex are excluded from the parameters shown in Figure 8, and questionnaire data parameters such as "I didn't have time to eat a balanced diet," "Did you eat anything?", and "I'm worried about food shortages" are excluded. Gender and these questionnaire parameters are included to improve clustering performance.

[0033] Figure 9 shows the data filtering process used when the target disease is depression, and Figure 10 is a list of the parameters shown in Figure 9. As shown in Figure 9, when the target disease is depression, there are no parameters to exclude in the first filtering step S21, and 13 parameters are removed in the second filtering step S22.

[0034] In the second filtering step S22, among the parameters shown in Figure 10, the following blood test data are excluded: red blood cell distribution width, red blood cell count, platelet count, monocyte percentage, mean platelet volume, mean corpuscular volume, hemoglobin, basophil percentage, and eosinophil percentage. The anthropometric data, mean abdominal sagittal diameter, is also excluded. Red blood cell distribution width, red blood cell count, platelet count, monocyte percentage, mean platelet volume, mean corpuscular volume, hemoglobin, basophil percentage, and eosinophil percentage are excluded to improve clustering performance, while mean abdominal sagittal diameter is excluded because it has a strong correlation with BMI. Furthermore, in the second filtering step S22, among the parameters shown in Figure 10, the demographic data parameters of age and gender are excluded, and the questionnaire parameter "Have you been told by a doctor that you have diabetes?", which is questionnaire data, is excluded. Gender and these questionnaire parameters are included to improve clustering performance.

[0035] In this embodiment, when the target disease is cardiovascular disease, blood test data, physical measurement data, interview data, and urine test data are used as category data, and 35 parameters are selected from this category data. When the target disease is diabetes, blood test data, physical measurement data, interview data, and urine test data are used as category data, and 38 parameters are selected from this category data. When the target disease is depression, blood test data, physical measurement data, interview data, and urine test data are used as category data, and 34 parameters are selected from this category data. However, only one of the category data may be used, and it is preferable to use at least two category data. In particular, by not using category data from blood test data, it is possible to estimate the risk of developing the disease at the healthy stage without performing highly invasive tests that cause mental distress.

[0036] Furthermore, the number of parameters can be any number. For example, if the target disease is a cardiovascular disease, if the blood test data includes total cholesterol and direct HDL cholesterol, these will be excluded from the judgment data as disease parameters. However, if the blood test data includes 25-hydroxyvitamin D2, white blood cell count, vitamin B12, segmented neutrophil percentage, red blood cell distribution width, red blood cell folate, red blood cell count, platelet count, monocyte percentage, mean platelet volume, mean corpuscular volume, lymphocyte percentage, hemoglobin, HbA1c, epi-25-hydroxyvitamin D3, 25-hydroxyvitamin D3, basophil percentage, or eosinophil percentage as blood test parameters, at least one blood test parameter can be used as judgment data. Furthermore, if the target disease is a cardiovascular disease, and the physical measurement data includes systolic blood pressure, diastolic blood pressure, arm circumference, mean sagittal abdominal diameter, BMI, or height, then at least one physical measurement parameter can be used as judgment data. Furthermore, if the target disease is a cardiovascular disease, if the interview data includes an interview parameter indicating that the estimated subject has had a heart attack, coronary heart disease, angina pectoris, or congestive heart failure, either currently or in the past, it will be excluded from the judgment data as a disease parameter. However, if the interview data includes an interview parameter related to kidney stones, diabetes, asthma, kidney disease, hepatitis, or sleep, at least one of these interview parameters can be used as judgment data. Furthermore, if the target disease is a cardiovascular disease, and the urine test data includes creatinine or albumin as urine test parameters, at least one urine test parameter can be used as judgment data.

[0037] Furthermore, if the target disease is diabetes, HbA1c is excluded from the judgment data as a disease parameter in the blood test data. However, if the blood test data includes 25-hydroxyvitamin D2, white blood cell count, vitamin B12, total cholesterol, segmented neutrophil percentage, red blood cell distribution width, red blood cell folate, red blood cell count, platelet count, monocyte percentage, mean platelet volume, mean corpuscular volume, lymphocyte percentage, hemoglobin, epi-25-hydroxyvitamin D3, 25-hydroxyvitamin D3, basophil percentage, eosinophil percentage, or direct HDL cholesterol as blood test parameters, at least one blood test parameter can be used as judgment data. Furthermore, if the target disease is diabetes, and the physical measurement data includes systolic blood pressure, diastolic blood pressure, arm circumference, mean sagittal abdominal diameter, BMI, or height, then at least one physical measurement parameter can be used as judgment data. Furthermore, if the target disease is diabetes, the questionnaire data indicating that the estimated subject currently or has had diabetes in the past will be excluded from the judgment data as a disease parameter. However, if the questionnaire data includes questionnaires regarding kidney stones, asthma, kidneys, hepatitis, heart attack, coronary heart disease, angina pectoris, congestive heart failure, or sleep, at least one of these questionnaire parameters can be used as judgment data. Furthermore, if the target disease is diabetes, and the urine test data includes creatinine or albumin as urine test parameters, at least one urine test parameter can be used as judgment data. In this way, using at least one categorical data and judgment data with any number of parameters, it is possible to determine which group an estimated subject belongs to or is closest to, and to map the group and the judgment result using at least two axes: risk rate and age.

[0038] Furthermore, if the target disease is depression, and the blood test data includes 25-hydroxyvitamin D2, white blood cell count, vitamin B12, total cholesterol, segmented neutrophil percentage, red blood cell distribution width, red blood cell folate, red blood cell count, platelet count, monocyte percentage, mean platelet volume, mean corpuscular volume, lymphocyte percentage, hemoglobin, HbA1c, epi-25-hydroxyvitamin D3, 25-hydroxyvitamin D3, basophil percentage, eosinophil percentage, or direct HDL cholesterol as blood test parameters, then at least one blood test parameter can be used as judgment data. Furthermore, if the target disease is depression, and the physical measurement data includes systolic blood pressure, diastolic blood pressure, arm circumference, mean sagittal abdominal diameter, BMI, or height, then at least one physical measurement parameter can be used as judgment data. Furthermore, if the target disease is depression, and the questionnaire data includes questions about diabetes, kidney stones, asthma, kidney disease, hepatitis, heart attack, coronary heart disease, angina pectoris, congestive heart failure, or sleep, then at least one of these questionnaire parameters can be used as judgment data. Furthermore, if the target disease is diabetes, and the urine test data includes creatinine or albumin as urine test parameters, at least one urine test parameter can be used as judgment data. In this way, using at least one categorical data and judgment data with any number of parameters, it is possible to determine which group an estimated subject belongs to or is closest to, and to map the group and the judgment result using at least two axes: risk rate and age.

[0039] The relative importance shown in Figures 6, 8, and 10 was calculated by normalizing the importance values ​​of all parameters to a range between 0 and 1. The relative importance (X) of parameter X is given by the following formula: Relative importance (X) = (Importance X - Minimum importance of all parameters) / (Maximum importance of all parameters - Minimum importance of all parameters) Here, importance X = separating power of all parameters - separating power excluding X The importance of a parameter is calculated by removing one parameter and measuring how much this removal affects the separation force.

[0040] Figure 11 shows the process of generating cardiovascular disease subtypes. Figure 11 shows the data filtering process used when the target disease is cardiovascular disease, and Figure 12 is a list of parameters used in the cardiovascular disease subtype generation process shown in Figure 11. As shown in Figure 11, when the target disease is cardiovascular disease, four parameters are removed in the first filtering step S21 and six parameters are removed in the second filtering step S22.

[0041] In the first filtering step S21, the following questionnaire parameters are excluded: "Have you ever said you've had a heart attack?", "Have you ever said you have coronary heart disease?", "Have you ever said you have angina pectoris / are experiencing angina pectoris?", and "Have you ever said you've had congestive heart failure?".

[0042] In the second filtering step S22, the following parameters are excluded from the parameters shown in Figure 12: the percentage of segmented neutrophils and epi-25-hydroxyvitamin D3, which are blood test data; the BMI, which is anthropometric data; the age and sex parameters, which are demographic data; and the questionnaire parameter, "Have you experienced loss of appetite?", which is questionnaire data. The percentage of segmented neutrophils, epi-25-hydroxyvitamin D3, age, sex, and questionnaire parameter are excluded to improve clustering performance.

[0043] Figure 13 shows the cardiovascular disease subcategory analysis. The confusion matrix in Figure 13 illustrates the separation of various cardiovascular disease subtypes. For example, in the example in Figure 13, if a person has previously had a heart attack, the algorithm has a 60% chance of identifying that person as a heart attack subtype, a 26% chance as a heart failure subtype, and a 14% chance as a stroke subtype. In subcategory analysis, the input is measured values ​​excluding diseases that show biomarkers, and the output is the disease subtype the patient has. Depending on the progression of the disease, from the healthy stage to the post-symptomatic stage, specific diseases are further subdivided, and the degree of prevalence in each subcategory is displayed.

[0044] The matrix in Figure 13 shows the validation results for the classification of cardiovascular disease subtypes. Here, it shows the degree of agreement between the data of subjects who have actually been diagnosed with the disease and the categories subclassified by AI without using this data. In the validation of cardiovascular disease subtype classification in Figure 13, the input is the results of cardiovascular risk analysis, and the output is an output or display of which cardiovascular disease subtype the patient has: heart attack, heart failure, or stroke. The clustering algorithm used here is almost the same as the algorithm used in the risk analysis, but the processing of cardiovascular disease subcategory analysis differs from the processing of risk analysis in the following respects. First, the outputs of the two are different. In risk analysis, there are only two outputs: low risk or high risk. In contrast, in subtype classification, the number of outputs is the same as the number of subtype classes. In this experiment, three subtypes are considered: heart attack, heart failure, and stroke. Second, the training data (ground truth data) is different. In risk analysis, two types of labeled data are required: healthy subjects and diseased subjects. In contrast, subtype classification requires labeled data for each subtype of the disease. In this experiment, three types of labeled data were used: subjects who had a heart attack, subjects who had heart failure, and subjects who had a stroke.

[0045] Figure 14 shows an overview of the glaucoma subcategory classification. Figure 14 shows the difference between the glaucoma subcategory classification method according to the present invention and a conventional method using unsupervised clustering. First, as for the clustering method, the conventional method uses unsupervised clustering, while the method of the present invention uses semi-supervised clustering. Semi-supervised clustering may preferably be multi-stage semi-supervised clustering. The disadvantages of conventional unsupervised clustering are that the clustering results are unpredictable and there is no guarantee that the resulting clusters correspond to the target subtype. In contrast, the advantage of using semi-supervised clustering in the method of the present invention is that the cluster type of the predetermined cluster group is determined in advance by a small amount of labeled data.

[0046] Furthermore, while conventional unsupervised clustering methods use biomarkers whose values ​​are proportional to disease progression as input data, the method of the present invention excludes biomarkers whose values ​​are proportional to disease progression. The disadvantages of using biomarkers whose values ​​are proportional to disease progression in conventional methods are that predictions are limited to the current state of the subject and future progression cannot be predicted. In contrast, the advantages of excluding biomarkers whose values ​​are proportional to disease progression in the method of the present invention are that predictions are not limited to the current state of the subject and the current level of disease progression can be predicted.

[0047] Furthermore, while conventional unsupervised clustering methods output the disease subtype as a single output result, the method of the present invention provides two stages of output: the disease subtype is output as the first stage output, and the current level of disease progression is output as the second stage output.

[0048] Figure 15 shows a graph of the diabetes progression rate analysis. In other embodiments of the present invention, the steps may further include predicting or indicating the expected rate of progression of a particular disease, depending on the degree of risk, from a healthy stage to the onset of the disease. The inputs and outputs in diabetes progression rate analysis are the same as those in risk analysis, but the method of visualizing the results differs. In age-related risk analysis, the x-axis is age and the y-axis is prevalence. This type of graph shows the proportion of people with the disease or at risk in different risk groups at different ages. In progression rate analysis, the x-axis is age and the y-axis is the mean value of the biomarker indicating the disease. For example, in the case of diabetes, the y-axis is the mean value of HbA1c for subjects of the same age and risk group. Since HbA1c is proportional to the progression of diabetes, a faster change in HbA1c indicates a faster progression of diabetes. Therefore, the slope of the progression rate analysis shows the rate of diabetes progression for subjects of different risk groups at different ages.

[0049] Figure 16 is a conceptual diagram showing the overall disease risk assessment system 1 according to this embodiment. The disease risk assessment system 1 can be implemented as part of a cloud AI platform. The cloud AI platform has a customer data management center that manages customer data from medical institutions, etc., and a health map API that provides a health map to the user terminal 50 based on data input from the user terminal 50. The disease risk assessment system 1 of the present invention is a system for realizing the health map API and is a system that performs specific processing for generating a health map. The customer data management center, the user terminal, and the health map API including the disease risk assessment system 1 are connected via a network and exchange data.

[0050] The disease risk assessment system 1 comprises a data processing unit 10 and a database 20. The data processing unit 20 includes a first filtering unit 11, a first clustering unit 12, a second filtering unit 13, a second clustering unit 14, and a clustering model storage unit 15 for clustering processing. The data processing unit 20 may also further include a mapping unit 16 for mapping processing. Furthermore, the data processing unit 20 may further include a verification unit 17 for verifying machine learning in the clustering process.

[0051] Database 20 includes a training data database 21 and an AI parameter database 22 for storing data related to the clustering process. Database 20 may also include a validation data database 24 for storing data related to the validation of machine learning in the clustering process.

[0052] Figure 17 shows the configuration of the clustering process of the disease risk assessment system 1 according to this embodiment. The disease assessment system 1 according to this embodiment, which evaluates the risk of contracting a specific disease, includes a diagnostic data database 21 that stores health-related diagnostic data, a first filtering unit 11 that reads diagnostic data from the diagnostic database 21 and excludes diagnostic data that changes depending on the disease level, a first clustering unit 12 that clusters the diagnostic data not excluded by the first filtering unit 11 and divides it into a high-risk group and a low-risk group, a second filtering unit 13 that extracts only the diagnostic data clustered into the high-risk group by the first clustering unit 12 from the diagnostic data database, and a second clustering unit 14 that clusters the diagnostic data extracted by the second filtering unit 13 and divides it into multiple disease levels. The system includes a clustering result storage unit 15 that stores the clustering results from the first clustering unit 12 and the second clustering unit 14.

[0053] The diagnostic data database 21 stores health-related diagnostic data received from the user terminal 50 or data input terminal 30, such as a terminal of an external system. Here, health-related diagnostic data refers to the results of any diagnosis, examination, or test related to health, such as diagnostic results obtained from health checkups or medical examinations, or diagnostic results obtained during consultations or tests at medical institutions. The diagnostic data includes measurement items as shown in the tables in Figures 6, 8, 10, and 12.

[0054] The first filtering unit 11 reads diagnostic data from the diagnostic database 21 and excludes diagnostic data that changes depending on the disease level. In other words, the first filtering unit 11 excludes data (features) that correlate with the disease level.

[0055] The first clustering unit 12 clusters the diagnostic data that was not excluded by the first filtering unit 11, dividing it into a group with a high risk of disease and a group with a low risk of disease.

[0056] The second filtering unit 13 extracts only the diagnostic data that has been clustered into groups with a high risk of disease by the first clustering unit 12 from the diagnostic data database.

[0057] The second clustering unit 14 performs clustering on the diagnostic data extracted by the second filtering unit 13, dividing it into multiple disease levels.

[0058] The clustering result storage unit 15 stores the clustering results from the first clustering unit 12 and the second clustering unit 14.

[0059] The AI ​​parameter database 22 stores the parameters that have been trained and optimized for the AI ​​engine. For example, if the AI ​​engine is built using a neural network, the AI ​​parameter database 22 stores the weights of the nodes in each layer.

[0060] Figure 18 shows the configuration of the mapping process for the disease risk assessment system according to this embodiment. As shown in Figure 18, the disease assessment system 1 according to this embodiment, which evaluates the risk of contracting a specific disease, may further include a mapping processing unit 16 that performs mapping processing for displaying the clustering results stored in the clustering result storage unit 15 in a graph.

[0061] The diagnostic data database 21 stores customer data such as customer ID and name, along with the customer's health diagnosis results. The customer data stored in the diagnostic data database 21 may be used to display the disease risk assessment results for each customer in graphs or other formats.

[0062] The mapping processing unit 16 performs mapping processing to display the clustering results stored in the clustering result storage unit 15 as a graph.

[0063] Figure 19 shows the configuration of the verification process for the disease risk assessment system according to this embodiment. The disease assessment system 1 according to this embodiment, which evaluates the risk of contracting a specific disease, may further include a verification unit 17 that compares verification data stored in a verification database 24 with AI prediction data which is the result of clustering in a data processing unit 10.

[0064] The verification data database 24 stores verification data. Preferably, the verification data is stored chronologically, covering several years.

[0065] The verification unit 17 compares the verification data stored in the verification database 24 with the AI-predicted data, which is the result of clustering in the data processing unit 10. The AI-predicted data to be compared is, for example, the result of clustering in the first clustering unit 12 or the result of clustering in the second clustering unit 14. For example, when verifying clustering in the first clustering unit 12 to divide patients into groups with a high risk of disease and groups with a low risk of disease, if there is accumulated data for four years, the AI ​​is trained using the accumulated data from the first two years, and the verification is performed by comparing the disease labels of the predicted data for the latter two years with the actual data for the latter two years using that AI engine.

[0066] Furthermore, the verification unit 17 may, for example, verify whether the distribution is appropriate by comparing the degree of disease risk for each age group with the actual disease distribution. [Industrial applicability]

[0067] According to the present invention, it is possible to propose reducing the risk of developing the disease through interventions such as improving lifestyle habits. As described above, the present invention makes it possible to determine the degree of risk of contracting a specific disease from the health stage using health checkup data. This is achieved through a two-stage method: a first stage in which individuals, including those who are healthy and whose risk of contracting the disease has not yet manifested (health stage), and those who have already developed the disease, are divided into high-risk and low-risk groups; and a second stage in which the degree of disease is further divided within the group judged to be at high risk. In the first stage in which individuals are divided into high-risk and low-risk groups, data (features) that correlate with the level of the disease, such as HbA1c in diabetes, are excluded. This avoids the influence of data (features) that correlate with the level of the disease on clustering, making it possible to divide individuals into high-risk and low-risk groups for a specific disease regardless of the stage of disease progression and even before the onset of the disease. As a result, it is possible to estimate the risk of contracting a specific disease from a healthy state before the onset of the disease, regardless of its progression, and to perform preventive health management for that specific disease from the health stage. For example, to assess the risk of developing diabetes from a healthy stage, data that changes proportionally with the progression of the disease, such as HbA1c, can be excluded. However, data on parameters that do not change with the progression of the disease, such as being overweight (obesity), but which accumulate damage and can increase the risk of developing diabetes in the future, can be retained and then clustered. Although the above description concerns examples, it will be apparent to those skilled in the art that the present invention is not limited thereto, and various changes and modifications can be made within the scope of the principles of the present invention and the appended claims. [Explanation of symbols]

[0068] S10 Step to acquire training data S20 Filtering step S21 First filtering step S22 Second filtering step S30 Learning Steps S40 Mapping processing step S50 Display Step S60 Verification Steps S70 Judgment Step 1. Disease Risk Assessment System 10 Data Processing Unit 11. First Filtering Section 12. First Clustering Unit 13. Second Filtering Section 14. Second Clustering Unit 15 Mapping section 16. Comparison Section 20 Databases 21. Training Data Database 22 AI Parameter Database 24. Verification Data Database 30 Data entry terminals 40 Clustering Model Memory Unit 50 User Terminals

Claims

1. A disease risk assessment method that evaluates the risk of contracting a specific disease, regardless of the stage of disease progression, from the healthy stage to after the onset of illness. Computers A learning data acquisition step in which at least one data is obtained from blood test data, physical measurement data, demographic data, medical interview data, and urine test data, A filtering step is performed to remove data from the acquired data that changes depending on the disease level of the disease for which the risk is to be assessed, A learning step in which machine learning is performed using data from which data that changes depending on the level of the disease has been removed, The steps include inputting the subject's data into the machine learning model generated by the learning step, and classifying the subject into a group with a high risk of contracting a specific disease and a group with a low risk of contracting a specific disease. A method that includes performing [something].

2. The aforementioned computer, The machine learning model generated by the learning step is input with data from subjects in a group at high risk of contracting the specific disease, and the step of determining the degree of risk of contracting the specific disease for subjects in the group at high risk of contracting the specific disease is determined. The method according to claim 1, which includes performing the following:

3. The method according to claim 2, characterized in that the items of data of the subject entered in the classification step are different from the items of data of the subject entered in the step of determining the degree of risk of contracting the specific disease.

4. The method according to claim 2 or 3, characterized in that the subject data entered in the classification step is obtained by excluding data whose value changes depending on the degree of risk of contracting the specific disease from the subject data entered in the step of determining the degree of risk of contracting the specific disease.

5. The method according to any one of claims 2 to 4, wherein, in the step of determining the degree of risk of contracting the particular disease, supervised learning and data whose value changes depending on the degree of risk of contracting the particular disease are used to determine the degree of risk of contracting the particular disease.

6. The method according to claim 4, wherein the data of the subject to be input does not include genetic information.

7. The aforementioned computer, A step to display the results of the classification in the classification step and the results of determining the degree of risk of contracting the specific disease in the step to determine the degree of risk of contracting the specific disease. This further includes performing, The method according to any one of claims 2 to 6, characterized in that the display step involves normalizing the results of the classification and the results of determining the degree of risk of contracting the specific disease, and displaying them as a radar chart.

8. The method according to any one of claims 1 to 7, wherein the classification step is performed by semi-supervised clustering or unsupervised clustering.

9. A disease risk assessment system that evaluates the risk of contracting a specific disease, regardless of the stage of disease progression, from the healthy stage to after the onset of illness. Equipped with a computer, The aforementioned computer, Obtain at least one data set from among blood test data, physical measurement data, demographic data, medical interview data, and urine test data. From the acquired data, remove the data that changes depending on the disease level of the disease being assessed for risk. Machine learning is performed using data from which data that changes depending on the level of the disease has been removed. A disease risk assessment system characterized by inputting subject data into a generated machine learning model and classifying the subjects into a group with a high risk of contracting a specific disease and a group with a low risk of contracting a specific disease.

10. The disease risk assessment system according to claim 9, wherein the computer classifies the specific disease into a plurality of subtypes and generates labeled data for each subtype.

11. The disease risk assessment system according to claim 9, wherein the computer predicts the predicted rate of progression of a specific disease according to the degree of risk of contracting the specific disease.

12. The disease risk assessment system according to claim 9, wherein the computer determines the degree of risk of contracting the specific disease for a group of people at high risk of contracting the specific disease.

13. The disease risk assessment system according to claim 12, characterized in that the items of data of the subject entered in the classification process are different from the items of data of the subject entered in the determination of the degree of risk of contracting the specific disease.

14. The disease risk assessment system according to claim 12 or 13, characterized in that the subject data entered in the classification is obtained by excluding data whose value changes depending on the degree of risk of contracting the specific disease from the subject data entered in determining the degree of risk of contracting the specific disease.

15. A disease risk assessment system according to any one of claims 10 to 14, wherein, in determining the degree of risk of contracting the aforementioned specific disease, supervised learning and data whose value changes depending on the degree of risk of contracting the aforementioned specific disease are used in determining the degree of risk of contracting the aforementioned specific disease.

16. The aforementioned computer, The method further includes displaying the results of the classification in the aforementioned classification and the results of determining the degree of risk of contracting the aforementioned particular disease in determining the degree of risk of contracting the aforementioned particular disease, The disease risk assessment system according to any one of claims 12 to 14, characterized in that the display involves normalizing the results of the classification and the results of determining the degree of risk of contracting the specific disease, and displaying them as scores.

17. The disease risk assessment system according to claim 9, 13, or 14, wherein the classification is performed by semi-supervised clustering or unsupervised clustering.

18. A health information processing device that evaluates the risk of contracting a specific disease, regardless of the stage of disease progression, from the healthy stage to after the onset of illness, Obtain at least one data set from among blood test data, physical measurement data, demographic data, medical interview data, and urine test data. From the acquired data, remove the data that changes depending on the disease level of the disease being assessed for risk. Machine learning is performed using data from which data that changes depending on the level of the disease has been removed. A health information processing device that inputs subject data into a generated machine learning model, classifies subjects into groups with a high risk of contracting a specific disease and groups with a low risk of contracting a specific disease, and generates disease risk assessment information.

19. The health information processing device according to claim 18, wherein data of subjects in a group with a high risk of contracting the specific disease are input into the generated machine learning model, and the degree of risk of contracting the specific disease is determined for subjects in the group with a high risk of contracting the specific disease.

20. The health information processing device according to claim 19, characterized in that the items of data of the subject that are input when classifying the subject are different from the items of data of the subject that are input when determining the degree of risk of contracting the specific disease.

21. The health information processing device according to claim 19 or 20, characterized in that the subject data entered in the classification is obtained by excluding data whose value changes depending on the degree of risk of contracting the specific disease from the subject data entered in determining the degree of risk of contracting the specific disease.

22. A health information processing device according to any one of claims 19 to 21, wherein, in determining the degree of risk of contracting the aforementioned specific disease, supervised learning and data whose value changes depending on the degree are used in determining the degree of risk of contracting the aforementioned specific disease.

23. The method further includes displaying the results of the classification in the aforementioned classification and the results of determining the degree of risk of contracting the aforementioned particular disease in determining the degree of risk of contracting the aforementioned particular disease, The health information processing device according to any one of claims 19 to 22, characterized in that the display involves normalizing the results of the classification and the results of determining the degree of risk of contracting the specific disease, and displaying them as scores.

24. The health information processing device according to any one of claims 18 to 21, wherein the classification is performed by semi-supervised clustering or unsupervised clustering.

25. Computers A learning data acquisition step in which at least one data is obtained from blood test data, physical measurement data, demographic data, medical interview data, and urine test data, A filtering step is performed to remove data from the acquired data that changes depending on the disease level of the disease for which the risk is to be assessed, A learning step in which machine learning is performed using data from which data that changes depending on the level of the disease has been removed. This includes performing the following: A method wherein the machine learning model generated by the learning step is a machine learning model that quantifies the level of disease of the disease to be assessed by inputting subject data into the machine learning model.

Citation Information

Patent Citations

  • Diagnostic prediction device of lifestyle related disease, diagnostic prediction method of lifestyle related disease, and program

    JP2012064087A

  • Physical examination information providing device and physical examination information providing method

    JP2013191020A

  • Information processor, radiation imaging system, and method for support

    JP2020102037A

  • Risk response analysis system, risk response analysis method and risk response analysis program

    JP2020173525A

  • Machine learning program, machine learning method, and machine learning apparatus

    JP2020190935A