Prediction device, prediction method, and program
The prediction device addresses the challenge of predicting health status over long periods by using distribution feature information to estimate future health data without requiring time-series data from the same individual, ensuring accurate predictions based on age group distributions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-04-07
AI Technical Summary
Existing methods for predicting health status over long periods, such as 10 or 20 years, require long-term time-series data on the same individual, which is often difficult to obtain.
A prediction device that calculates distribution feature information based on the relationship between health data distributions of different age groups and generates a predicted distribution using this information to estimate future health data without relying on time-series data from the same individual.
Enables accurate prediction of health status after a predetermined period even when time-series data is unavailable, considering the characteristics of health data distributions across age groups.
Smart Images

Figure 2026059250000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to a prediction device, a prediction method, and a program. [Background technology]
[0002] Technologies exist for future analysis of personal information in fields such as healthcare.
[0003] Patent Document 1 discloses a technique for predicting the state of a subject at a predetermined future time based on the subject's medical records. Specifically, Patent Document 1 discloses a method for obtaining a subject's past medical records, generating multiple input vectors based on the medical records, and inputting each input vector into a learning algorithm to obtain an output vector that indicates a prediction of the subject's state at a predetermined future time. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2018-194904 [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] In Patent Document 1, the input vector is generated from medical claim data, health checkup data, or questionnaire data, etc., from multiple past points in time for the same subject. Then, learning is performed using this input vector.
[0006] In other words, Patent Document 1 uses time-series data obtained by observing the same person over a certain period of time in order to predict the state of the subject. However, when making predictions for longer periods, such as 10 or 20 years from now, in order to perform a similar prediction method, time-series data on the same person over a long period of time is required. In this case, long-term observation of the same person is necessary. Therefore, it is highly likely that obtaining such time-series data on the same person over a long period of time will be difficult.
[0007] One of the purposes of this disclosure is to address the above-mentioned issues and to provide a predictive device, etc., that can predict the health status after a predetermined period, even when there is no time-series data for a predetermined period. [Means for solving the problem]
[0008] A prediction device according to one aspect of the present disclosure includes: acquisition means for acquiring health data, which is data relating to the health of multiple persons, in a first period and a second period predetermined prior to the first period; calculation means for calculating distribution feature information that indicates the characteristics of the distribution of health data of the first age group in the first period, based on the relationship between the distribution of health data in the first period, which is the distribution of health data of the first age group corresponding to the age of the target person, and the distribution of health data in the second period, which is the distribution of health data of the first age group or the distribution of health data of the age group corresponding to the first age group prior to the predetermined period; generation means for generating a predicted destination distribution by integrating the distribution in the first period, which is the distribution of health data of a predicted destination age group corresponding to the age of the target person after the predetermined period has elapsed, and the distribution feature information; and prediction means for predicting health data in the predicted destination distribution, which is the destination of the data corresponding to the health data of the target person, from the distribution of health data of the first age group in the first period.
[0009] A prediction method according to one aspect of the present disclosure acquires health data, which is data relating to the health of multiple individuals, in a first period and a second period predetermined prior to the first period. Based on the relationship between the distribution of health data in the first period, which is the distribution of health data for a first age group corresponding to the age of the target individual, and the distribution of health data in the second period, which is the distribution of health data for the first age group or the distribution of health data for an age group corresponding to the first age group prior to the predetermined period, distribution feature information indicating the characteristics of the distribution of health data for the first age group in the first period is calculated. A predicted target distribution is generated by integrating the distribution in the first period, which is the distribution of health data for a predicted target age group corresponding to the age of the target individual after the predetermined period has elapsed, with the distribution feature information. The method predicts the health data in the predicted target distribution that will be the transition destination of the data corresponding to the health data of the target individual, from the distribution of health data for the first age group in the first period.
[0010] A program according to one aspect of this disclosure causes a computer to perform the following steps: a process of acquiring health data, which is data relating to the health of multiple persons, in a first period and a second period a predetermined period prior to the first period; a process of calculating distribution feature information that shows the characteristics of the distribution of health data of the first age group in the first period, based on the relationship between the distribution of health data in the first period, which is the distribution of health data of the first age group corresponding to the age of the target person, and the distribution of health data in the second period, which is the distribution of health data of the first age group or the distribution of health data of an age group corresponding to the first age group prior to the predetermined period; a process of generating a predicted target distribution by integrating the distribution in the first period, which is the distribution of health data of a predicted target age group corresponding to the age of the target person after the predetermined period has elapsed, and the distribution feature information; and a process of predicting the health data in the predicted target distribution, which is the transition destination of the data corresponding to the health data of the target person, from the distribution of health data of the first age group in the first period. [Effects of the Invention]
[0011] According to the present disclosure, even when there is no time-series data regarding the same person, it is possible to estimate a state transition considering the past state.
Brief Description of Drawings
[0012] [Figure 1] FIG. 1 is a first block diagram showing an example of the functional configuration of the prediction device of the present disclosure. [Figure 2] FIG. 2 is a first flowchart for explaining an example of the operation of the prediction device of the present disclosure. [Figure 3] FIG. 3 is a second block diagram showing an example of the functional configuration of the prediction device of the present disclosure. [Figure 4] FIG. 4 is a diagram showing an example of the probability density distribution of the present disclosure. [Figure 5] FIG. 5 is a diagram for explaining an image when solving the transition from one probability density distribution to another probability density distribution of the present disclosure as an optimal transport problem. [Figure 6] FIG. 6 is a second flowchart for explaining an example of the operation of the prediction device of the present disclosure. [Figure 7] FIG. 7 is a third block diagram showing an example of the functional configuration of the prediction device of the present disclosure. [Figure 8] FIG. 8 is a third flowchart for explaining an example of the operation of the prediction device of the present disclosure. [Figure 9] FIG. 9 is a fourth flowchart for explaining an example of the operation of the prediction device of the present disclosure. [Figure 10] FIG. 10 is a block diagram showing an example of the hardware configuration of a computer device for realizing the prediction device of the present disclosure.
Modes for Carrying Out the Invention
[0013] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. <0000The prediction device described in this disclosure makes future predictions about data based on state transitions between data, which are estimated using accumulated data. In this disclosure, an example of the data covered is health data, which is data related to a person's health. That is, the prediction device described in this disclosure can predict how a person's health data will change in the future based on state transitions from health data in one stage to health data in another stage.
[0016] Health data may include, for example, the values of test items in a health checkup, or information about a person's exercise habits. Health data is not limited to these examples. Furthermore, health data may be pre-stored, for example. In this case, the stored health data may be data for multiple people at a predetermined point in time. For example, the results of health checkups for 10,000 people in a predetermined year may be stored as health data. In this disclosure, we mainly describe an example in which a prediction device estimates data transitions based on health data, but the data to be covered is not limited to this example.
[0017] Figure 1 is a first block diagram showing an example of the functional configuration of the prediction device 100. As shown in Figure 1, the prediction device 100 comprises an acquisition unit 110, a calculation unit 120, a generation unit 130, and a prediction unit 140.
[0018] The acquisition unit 110 acquires health data. More specifically, the acquisition unit 110 acquires health data for a first period. For example, the acquisition unit 110 acquires the results of health examinations conducted on multiple individuals in year t as health data. The acquisition unit 110 also acquires health data for a second period. The second period is a predetermined period prior to the first period. For example, the acquisition unit 110 acquires the results of health examinations conducted on multiple individuals in year t-1 as health data.
[0019] Health data may be datasets classified by stage. A stage may be information indicating a tier when the dataset is stratified. In other words, a stage can be said to be information indicating the conditions under which the population data is classified into subsets based on predetermined conditions. For example, a dataset for each stage may be health data for each age group. More specifically, if the health data includes data showing the blood glucose levels of multiple individuals, the health data for each stage may include data showing the blood glucose levels of individuals in their teens, individuals in their twenties, ..., and individuals in their eighties. In this example, the dataset includes data showing blood glucose levels for each age group of 10 years. Thus, the stages may have a fixed order. For example, the stage following the stage for teenagers is the stage for individuals in their twenties. Note that the age groups may be any age groups. For example, the health data may be classified into data showing blood glucose levels for each age group of one year.
[0020] Furthermore, health data may be stored in a storage device (not shown). In this case, the storage device may be a device owned by the prediction device 100, or it may be an external device that is communicatively connected to the prediction device 100.
[0021] In this way, the acquisition unit 110 acquires health data, which is data on the health of multiple individuals, during a first period and a second period that precedes the first period by a predetermined period. The acquisition unit 110 is an example of an acquisition means.
[0022] The calculation unit 120 calculates distribution feature information. Distribution feature information indicates the characteristics of the distribution of health data for a specific age group. For example, the distribution feature information indicates the characteristics of the distribution of health data for the age group corresponding to the age of the subject. Here, the age group corresponding to the age of the subject is referred to as the first age group.
[0023] The calculation unit 120 may calculate the difference between the distribution of health data for the first age group in the first period and the distribution of health data for the first age group in the second period as distribution feature information. For example, the calculation unit 120 calculates the difference between the distribution of health data for 50-year-olds in year t and the distribution of health data for 50-year-olds in year t-1. The second age group is defined as an age group a predetermined period prior to the first age group. In this case, the calculation unit 120 may calculate information based on the transition from the distribution of health data for the second age group in the second period to the distribution of health data for the first age group in the first period as distribution feature information. For example, the calculation unit 120 may use the transition from the distribution of health data for 49-year-olds in year t-1 to the distribution of health data for 50-year-olds in year t. That is, the calculation unit 120 calculates distribution feature information based on the relationship between the distribution of the first age group in the first period and the distribution of the first age group in the second period, or the distribution of the second age group in the second period. Note that the examples for calculating distributional feature information are not limited to those described above.
[0024] Thus, the calculation unit 120 calculates distribution feature information that shows the characteristics of the distribution of health data for the first age group in the first period, based on the relationship between the distribution of health data for the first age group in the first period, which corresponds to the age of the target person, and the distribution of health data for the second period, which corresponds to the distribution of health data for the first age group or the distribution of health data for an age group corresponding to a predetermined period prior to the first age group. The calculation unit 120 is an example of a calculation means.
[0025] The generation unit 130 generates a predicted distribution using distribution feature information. Suppose we want to predict the health data of a target person after a predetermined period of time has elapsed. Here, the age group corresponding to the target person's age after the predetermined period of time has elapsed is called the predicted age group. The generation unit 130 generates the predicted distribution by integrating the distribution of health data for the predicted age group in the first period with the distribution feature information. For example, suppose we want to predict the health data of a 50-year-old target person 10 years from now. In this case, the generation unit 130 may integrate the distribution of health data for a 60-year-old in the first period with the distribution feature information. One example of integration is a linear sum of a matrix representing the distribution of health data for the predicted age group and a matrix representing the distribution feature information. Note that the method for generating the predicted distribution is not limited to this example.
[0026] Thus, the generation unit 130 generates a predicted distribution by integrating the distribution of health data for a predicted age group corresponding to the age of the target person after a predetermined period, which is the distribution during the first period, with distribution feature information. The generation unit 130 is an example of a generation means.
[0027] The prediction unit 140 uses the predicted distribution to predict the future health data of the target individual. Specifically, the prediction unit 140 estimates the state transition to the predicted distribution from the distribution of health data for the first age group in the first period.
[0028] Assume that the health data consists of data showing the blood glucose levels of multiple individuals. Also, assume that the first age group is 50 years old. Furthermore, assume that the predicted distribution is a distribution generated by integrating the distribution of health data for 60-year-olds during the first period with distribution feature information. In this case, the state transition of the data from the distribution of blood glucose levels of 50-year-olds to the predicted distribution is estimated. This state transition may be estimated using an optimal transport problem algorithm. That is, the most likely transition from the probability distribution showing the probability of a 50-year-old individual having each blood glucose level to the probability distribution showing the probability of a 60-year-old individual having each blood glucose level may be estimated. In this estimation, a mapping from the probability distribution of 50-year-olds to the probability distribution of 60-year-olds is estimated. Estimating this mapping is equivalent to estimating the state transition probability regarding the transition from the distribution of blood glucose levels of 50-year-olds to the distribution of blood glucose levels of 60-year-olds.
[0029] The prediction unit 140 uses, for example, such state transition probabilities to calculate which of the predicted distributions the data corresponding to the target person's health data will transition to within the distribution of health data for the first age group. Based on this, the prediction unit 140 may predict the target person's future health data.
[0030] Thus, the prediction unit 140 predicts the health data in the predicted destination distribution, which is the data transition destination corresponding to the health data of the target person, from the distribution of health data for the first age group during the first period. The prediction unit 140 is an example of a prediction means.
[0031] Next, an example of the operation of the prediction device 100 will be explained using Figure 2. In this disclosure, each step of the flowchart will be represented by a number assigned to each step, such as "S1".
[0032] Figure 2 is a flowchart illustrating an example of the operation of the prediction device 100.
[0033] The acquisition unit 110 acquires health data, which is data on the health of multiple individuals, during a first period and a second period that precedes the first period by a predetermined period (S1).
[0034] The calculation unit 120 calculates distribution feature information that shows the characteristics of the distribution of health data for the first age group in the first period, based on the relationship between the distribution of health data for the first age group in the first period, which corresponds to the age of the target person, and the distribution of health data for the second period, which corresponds to the distribution of health data for the first age group or the distribution of health data for an age group corresponding to a predetermined period prior to the first age group (S2).
[0035] The generation unit 130 generates a predicted distribution by integrating the distribution of health data for a predicted age group corresponding to the age of the target person after a predetermined period has elapsed, which is the distribution during the first period, with distribution feature information (S3).
[0036] The prediction unit 140 predicts the health data in the predicted destination distribution, which is the destination of the data corresponding to the health data of the target person, from the distribution of health data for the first age group in the first period (S4).
[0037] As described above, the prediction device 100 of the first embodiment acquires health data, which is data relating to the health of multiple individuals, during a first period and a second period that precedes the first period by a predetermined period. The prediction device 100 also calculates distribution feature information that shows the characteristics of the distribution of health data for the first age group during the first period, based on the relationship between the distribution of health data for the first age group during the first period, which corresponds to the age of the target individual, and the distribution of health data for the second period, which corresponds to the distribution of health data for the first age group or the age group that precedes the first age group by a predetermined period. The prediction device 100 also generates a predicted target distribution by integrating the distribution of health data for the predicted target age group during the first period, which corresponds to the age of the target individual after a predetermined period has elapsed, and the distribution feature information. The prediction device 100 then predicts the health data in the predicted target distribution, which is the transition destination of the data corresponding to the health data of the target individual, from the distribution of health data for the first age group during the first period.
[0038] In other words, the prediction device 100 makes predictions based on the transition from the distribution of health data for the first age group in the first period to the distribution of health data for the target age group. At this time, the prediction device 100 does not use a method that requires time-series data of the same person observed over a certain period in order to make predictions about data concerning the subject over a certain period. That is, the prediction device 100 can make predictions about the health status after a predetermined period has elapsed, even if there is no time-series data for a predetermined period.
[0039] Furthermore, distribution feature information is incorporated into the predicted destination distribution. Distribution feature information is information that indicates the characteristics of the distribution of health data for the first age group in the first period, which is the source of the transition. In other words, the prediction device 100 can make predictions that take into account the individuality of the distribution of health data for the first age group in the first period.
[0040] <Second Embodiment> Next, a prediction device according to a second embodiment will be described. In the second embodiment, further examples relating to the prediction device described in the first embodiment will be described. In the second embodiment, the main description will be of an example in which the prediction device makes predictions using the transition of health data, but the data to be used is not limited to the examples described below. Note that some explanations that overlap with the first embodiment will be omitted.
[0041] [Details of prediction device 100] Figure 3 is a block diagram showing an example of the functional configuration of the prediction device 100. The prediction device 100 comprises an acquisition unit 110, a calculation unit 120, a generation unit 130, and a prediction unit 140. The prediction device 100 may also include a classification unit 150 and an estimation unit 160. Furthermore, the prediction device 100 may include a state transition model generation unit 170. In addition, the prediction device 100 may include a storage device 190. The storage device 190 may be a device owned by the prediction device 100, or it may be an external device that is communicatively connected to the prediction device 100.
[0042] The prediction device 100 is, for example, a device installed in a terminal device such as a personal computer. The terminal device is a device operated by a user. Although not limited to this example, the prediction device 100 may be a device implemented in a server device that is communicably connected to the terminal device via a wired or wireless network. The prediction device 100 may perform various processes in response to instructions from the terminal device.
[0043] Furthermore, the prediction device 100 may be connected to other devices via a wired or wireless network for communication. For example, the prediction device 100 may be able to communicate with an external server device that holds health data. The external server device may be, for example, a device managed by a hospital, local government, or company.
[0044] The acquisition unit 110 acquires a dataset related to health data. At this time, the dataset is stored in the storage device 190. For example, the acquisition unit 110 may acquire a dataset related to health data by reading the dataset stored in the storage device 190 in accordance with instructions from the terminal device.
[0045] At this time, the storage device 190 may store data previously acquired by the acquisition unit 110. Specifically, the acquisition unit 110 acquires health data for a first period and health data for a second period from an external server device that manages health data. For example, suppose the external server device manages the results of health checkups for 10,000 people for year t and year t-1. The acquisition unit 110 acquires the results of health checkups for 10,000 people for year t and year t-1 as health data from the external server device. The health data may be information corresponding to the examination items of the health checkup. Also, the health data may be the results of a health checkup received by each of the 10,000 people at a single point in time within a single year. For example, the health data for year t may include the results of health checkups received by 10,000 people at a single point in time within year t. Also, for example, the health data for year t-1 may include the results of health checkups received by 10,000 people at a single point in time within year t-1. Thus, a dataset for a single period does not have to be data showing the change over time of the same object, but may be data measured at a specific point in time for each of multiple objects. Furthermore, the individuals in the health data for year t may be different from those in the health data for year t-1. Also, the number of individuals in the health data for year t may be different from the number of individuals in the health data for year t-1. The acquisition unit 110 then stores the acquired health data in the storage device 190.
[0046] The predetermined period, which is the interval between the first period and the second period, does not have to be one year. The predetermined period may be two years or more, or less than one year. Furthermore, the method of acquiring health data is not limited to this example. For example, there may be a recording medium that stores the health data. In this case, the terminal device reads the health data from the recording medium. The acquisition unit 110 may then acquire the health data read by the terminal device.
[0047] Health data is processed, for example, by a classification unit 150. The processed health data may then be stored in a storage device 190. The classification unit 150 processes the data acquired by the acquisition unit 110 into datasets according to certain conditions. For example, the classification unit 150 classifies the health data by age group. For example, the classification unit 150 classifies the health data in 1-year increments. However, this is not the only example, and the classification unit 150 may classify the health data at any age interval. For example, the classification unit 150 may classify the health data in 10-year increments.
[0048] Furthermore, the classification unit 150 may extract specific data from the health data and classify the extracted data by age group. For example, suppose the health data includes information such as height, weight, blood pressure, blood glucose level, HbA1c, and BMI (Body Math Index). In this case, the classification unit 150 may classify the data showing blood glucose level and BMI from the health data by age group.
[0049] The classification criteria and data to be extracted may be information provided in response to instructions from the terminal device. That is, the user operating the terminal device inputs information indicating the classification criteria and data to be extracted into the terminal device. The terminal device transmits the input information to the prediction device 100. The classification unit 150 processes the health data using the information indicating the classification criteria and data to be extracted transmitted from the terminal device.
[0050] Furthermore, the classification unit 150 generates a distribution for the acquired data. Specifically, the classification unit 150 generates a probability density distribution for each condition based on the acquired data. For example, the classification unit 150 generates a distribution plotting data showing blood glucose levels and BMI for each age group. The distribution generated at this time is a two-dimensional distribution relating blood glucose levels and BMI. The classification unit 150 then generates a probability density distribution showing the probability of existence for each value of blood glucose level and BMI. i Let the data values of other dimensions be x j In this case, the probability density distribution is p([x i ,x jThis can be expressed as ]). In this way, the classification unit 150 generates a probability density distribution for each age group regarding the acquired health data.
[0051] In this case, the classification unit 150 may classify the data in each distribution into data groups. In this case, the probability density distribution will be a distribution showing the probability of existence for each data group. Figure 4 is a diagram showing an example of a probability density distribution. The probability density distribution shown in Figure 4 has blood glucose level and BMI as axes. Also, in the example in Figure 4, 64 cells are shown. These cells are classified data groups showing each person's blood glucose level and BMI. The probability of existence for each data group is shown. In this way, the classification unit 150 may classify the data in each distribution of health data for each age group in each period into data groups. The classification unit 150 is an example of a classification means.
[0052] The classification unit 150 may generate distributions of two or more dimensions. For example, the classification unit 150 may generate probability density distributions with blood glucose level, BMI, and average daily step count as axes. The probability density distribution divided into cells as shown in Figure 4 corresponds to the marginal distribution based on each value of the health data. The probability density distributions discussed hereafter may be probability distributions like those shown in Figure 4, or they may be probability distributions that do not take the form of marginal distributions.
[0053] The classification unit 150 stores the processed health data again in the storage device 190. For example, the classification unit 150 may store the probability density distribution of health data for each age group for each period in the storage device 190.
[0054] The acquisition unit 110 acquires health data classified according to conditions as stage-specific datasets. For example, the acquisition unit 110 may acquire the probability density distribution of health data for each age group as described above as stage-specific datasets. That is, the acquisition unit 110 may acquire the probability density distribution of health data for each age group in a first period and the probability density distribution of health data for each age group in a second period.
[0055] Furthermore, the acquisition unit 110 acquires the health data of the subject person. The health data of the subject person is also referred to as the target data. For example, if a probability density distribution for blood glucose level and BMI has been acquired, the acquisition unit 110 acquires the target data showing the blood glucose level and BMI of the subject person. In this case, the target data includes information indicating the age of the subject person.
[0056] The acquisition unit 110 may acquire further information. For example, the acquisition unit 110 may acquire health data for other periods.
[0057] The calculation unit 120 calculates distribution feature information that shows the characteristics of the distribution of health data for the first age group during the first period. Then, the generation unit 130 generates a predicted target distribution using the calculated distribution feature information. At this time, the predicted target distribution may be a probability density distribution that incorporates the distribution feature information into the health data for the predicted target age group. Details of the method for calculating the distribution feature information and the method for generating the predicted target distribution will be described later.
[0058] The prediction unit 140 predicts the changes in the target person's health data. In other words, the prediction unit 140 predicts the future health data values of the target person based on the target data. Specifically, the prediction unit 140 predicts the health data values when the target person reaches the age corresponding to the predicted age group. In this case, the first age group is the age group corresponding to the target person's age.
[0059] For example, suppose the subject is 51 years old. In this case, the subject's age corresponds to the first age group. The prediction unit 140 identifies which data or data group the subject data belongs to within the probability density distribution of health data for 51-year-olds in the first period. Then, the prediction unit 140 predicts, based on state transition probabilities, which data or data group the identified data or data group will transition to within the predicted distribution.
[0060] The state transition probability is estimated by the estimation unit 160. The estimation unit 160 estimates the state transition probability between distributions. Specifically, the estimation unit 160 estimates the state transition probability based on the transition to the predicted target distribution from the distribution of health data for the first age group in the first period. Here, the predicted target age group corresponds to the age group after a predetermined period has elapsed from the first age group. For example, suppose the first period is t years. Also, suppose the second period, which is a predetermined period earlier than the first period, is t-1 years. That is, suppose the difference between the first period and the second period is 1 year. In this case, if the first age group is 51 years old, the predicted target age group is 52 years old and older. Hereafter, the distribution of health data for the first age group in the first period will also be simply referred to as the first distribution.
[0061] The estimation unit 160 estimates the transition from the first distribution to the predicted distribution using an algorithm for the optimal transport problem (hereinafter referred to as the optimal transport algorithm). The optimal transport algorithm is an algorithm that finds a transport method that optimizes the cost required to transition a given probability distribution to another probability distribution.
[0062] Specifically, for distributions μ and ν in a probability space X, the direct product X 2 A distribution π in a given system is said to be a coupling if the following equations 1 and 2 hold true.
[0063]
number
[0064]
number
[0065] Let Π(μ,ν) be the set of all couplings. Let c(x,y) be the cost function for transporting an element x from distribution μ to an element y from distribution ν. In this case, for example, in equation 3 below, the coupling that minimizes the cost is called the optimal transport.
[0066]
Number
[0067] Alternatively, if the mapping from distribution μ to ν is T(x), the direct transition may be obtained by finding T that minimizes the following number 4. In this case, T is a one-to-one mapping. Also, for a subset U of μ, assume that the volume of its mapping T(U) is the same as that of U.
[0068]
Number
[0069] When performing optimal transport for discrete data, it can also be formulated as follows. Specifically, C ij is used as the cost matrix, and the distributions are μ i , ν j . At this time, under the conditions shown in number 6, find P ij that minimizes the number 5 representing the total cost.
[0070]
Number
[0071]
Number
[0072] In this way, the optimal transport algorithm can calculate a pair of pre-transport data and destination data that optimizes the cost of transporting from the first distribution to the predicted destination distribution.
[0073] The estimation unit 160 estimates the state transition probability based on the transition from the first distribution (i.e., the distribution of the health data of the first age group in the first period) to the predicted destination distribution. At this time, the estimation unit 160 solves the transition from the first distribution to the predicted destination distribution as an optimal transport problem.
[0074] Figure 5 illustrates the concept of solving the transition from one probability density distribution to another as an optimal transport problem. Figure 5 shows the probability density distribution for health data of the first age group in the first period, and the target distribution. The target distribution can also be called the probability density distribution of the target age group. The estimation unit 160 solving the transition from the first distribution to the target distribution as an optimal transport problem is equivalent to estimating which of the cells in the target age group's probability density distribution each cell in the first age group's probability density distribution has a higher probability of transitioning to. In other words, the estimation unit 160 estimates the state transition probability based on the transition from each data group in the first distribution to each data group in the target distribution.
[0075] For example, let μ be the probability density distribution of the first age group and ν be the probability density distribution of the target age group (target distribution). In this case, the estimation unit 160 estimates the mapping T using, for example, equation 4. For example, the estimation unit 160 models the mapping T as a function using a neural network such as a multi-layer fully connected layer. The estimation unit 160 also obtains the mapping T by optimizing using machine learning so that equation 3 becomes small. Furthermore, the estimation unit 160 generates multiple y values that transition from a given x using this mapping T. Then, the estimation unit 160 calculates the state transition probability from the generated y values. In this way, the estimation unit 160 estimates the state transition probability. Note that the prediction unit 140 may have the same functionality as the estimation unit 160.
[0076] The prediction unit 140 predicts the data set of health data in the predicted destination distribution, which is the data set to which the health-classified data set of the target person will transition, from the distribution of health data for the first age group in the first period. For example, the prediction unit 140 uses the state transition probabilities estimated in this way to predict the transition of the target person's health data.
[0077] Next, we will explain a specific example of generating a predicted target distribution. In the following example, the first period is year t, and the second period, which is a predetermined period before the first period, is year t-1. The first age group is set to 51 years old, and the second age group is set to 50 years old. The second age group is set to an age group that is a predetermined period before the first age group. Furthermore, the predicted target age group is set to 70 years old. In other words, the following example explains how to generate a predicted target distribution when estimating the transition from the distribution of health data for 51-year-olds in year t to the distribution of health data for 70-year-olds. It is also assumed that the memory device 190 stores the probability density distribution of health data for each year in year t-1 and the probability density distribution of health data for each year in year t. Hereafter, the distribution of health data will also be simply referred to as "distribution".
[0078] [First example of generating a predicted distribution] In the first example, we will explain how to generate a predicted target distribution using distribution feature information based on the relationship between the first distribution (the distribution of the first age group in the first period) and the distribution of the first age group in the second period. Specifically, the prediction device 100 generates a predicted target distribution by applying a state transition model, which is generated based on the distribution of the first age group in the second period, to the first distribution.
[0079] In this example, the prediction device 100 may include a state transition model generation unit 170. The state transition model generation unit 170 generates a state transition model that predicts the transition from the distribution of health data for one age group in a second period to the distribution of health data for another age group in a first period. The other age group is an age group that is a predetermined period after the first age group.
[0080] For example, the state transition model generation unit 170 identifies the distribution of health data for 51-year-olds in year t-1. The state transition model generation unit 170 also identifies the distribution of health data for 52-year-olds in year t. Then, from the distribution of health data for 51-year-olds in year t-1 and the distribution of health data for 52-year-olds in year t, the state transition model generation unit 170 generates a state transition model that shows the state transition probability of a 51-year-old person when they turn 52. This state transition model is called P(X 52 |X 51 This is expressed as ). In this example, the predetermined period is 1 year. Therefore, the state transition model when a person in one age group i (where i is a natural number) moves to another age group i+1 can be expressed as shown in the following equation 7.
[0081]
number
[0082] The state transition model generation unit 170 generates a state transition model corresponding to each age group. In this case, the state transition model generation unit 170 may generate a state transition model in the range where i is from the first age group to an age group that is a predetermined period before the predicted target age group. In this example, the state transition model generation unit 170 may generate a state transition model in the range (51 ≤ i < 70).
[0083] Thus, the state transition model generation unit 170 generates a state transition model that predicts the transition of the distribution when moving from one age group to another, based on the relationship between the distribution of health data for one age group in a second period and the distribution of health data for other age groups that are age groups a predetermined period after the one age group in the first period. The state transition model generation unit 170 is an example of a state transition model generation means. Note that the calculation unit 120 may also have the functions of the state transition model generation unit 170.
[0084] The calculation unit 120 uses the generated state transition model to calculate distribution feature information. For example, the calculation unit 120 calculates P(X 52 |X 51The formula ) is applied to the first distribution. This predicts the distribution of 52-year-olds based on the first distribution. Furthermore, the calculation unit 120 applies P(X) to the predicted distribution of 52-year-olds. 53 |X 52 The calculation unit 120 applies the following process. This predicts the distribution of 53-year-olds based on the predicted distribution of 52-year-olds. The calculation unit 120 performs the same process until the distribution of 70-year-olds, which is the target age group, is predicted. The calculation unit 120 calculates the predicted distribution of 70-year-olds as distribution feature information.
[0085] In this way, the calculation unit 120 uses a state transition model that predicts the transition of the distribution when moving from one age group to an age group after a predetermined period, and calculates the distribution of the predicted target age group estimated from the first distribution. The distribution of the predicted target age group at this time corresponds to the post-transition distribution that shows the destination of the first distribution, using a state transition model based on the accumulated health data for two periods.
[0086] In other words, the calculation unit 120 uses a state transition model to calculate a post-transition distribution as distribution feature information, which indicates the transition destination of the distribution of health data for the first age group during the first period when the group moves from the first age group to the predicted target age group.
[0087] The generation unit 130 integrates the calculated distribution feature information with the distribution of health data for the predicted age group during the first period. For example, suppose each distribution is represented by matrix M. In this case, the distribution of health data for the predicted age group during the first period is M b Let's assume that the distribution feature information in this example is M f Let's assume that the predicted distribution is M. p Let's assume that in this case, the predicted distribution M p This can be expressed as shown in the following number 8.
[0088]
number
[0089] α is a coefficient. Thus, the generation unit 130 may generate the predicted distribution as a linear combination of matrices. This is not limited to this example, and each distribution may be represented as a list. Also, α may be arbitrarily determined. For example, it may be set based on the true values of the predicted distribution and the distribution of the predicted age group that were generated in the past. Specifically, suppose that in the past, the predicted distribution for 70-year-olds was generated using the distribution for 51-year-olds in year t'. Also, suppose that the prediction device 100 has health data for year t'+19. In this case, the prediction device 100 adjusts the coefficient α so that the formula that generated the predicted distribution in the past calculates a distribution similar to the distribution of the predicted age group in year t'+19. The prediction device 100 then generates the predicted distribution M described above. p The adjusted α may be used when generating the result.
[0090] [Second example of generating a predicted distribution] The second example describes another example of generating a predicted target distribution using distribution feature information based on the relationship between the first distribution (the distribution of the first age group in the first period) and the distribution of the first age group in the second period. Specifically, the prediction device 100 generates a predicted target distribution by utilizing the difference between the first distribution and the distribution of the first age group in the second period.
[0091] For example, the calculation unit 120 identifies the distribution of 51-year-olds in year t and the distribution of 51-year-olds in year t-1. The calculation unit 120 then calculates the difference between the distribution of 51-year-olds in year t and the distribution of 51-year-olds in year t-1. This difference is referred to as the first difference.
[0092] The distribution of 51-year-olds in year t is shown in matrix M. t_51 Let's assume that the distribution of 51-year-olds in year t-1 is represented by matrix M. t-1_51 Let's assume that the first difference is D1. Then the first difference is calculated as shown in equation 9 below.
[0093]
number
[0094] The calculation unit 120 may calculate the first difference as distribution feature information. That is, the calculation unit 120 may calculate the first difference, which is the difference between the distribution of health data for the first age group in the first period and the distribution of health data for the first age group in the second period, as distribution feature information.
[0095] The generation unit 130 integrates the calculated first difference with the distribution of health data for the predicted age group during the first period. For example, the predicted distribution M p It can be expressed as shown in the following number 10.
[0096]
number
[0097] Similar to the first example, the generation unit 130 may generate the predicted distribution as a linear combination of matrices. This example is not limited to this one; each distribution may be represented as a list. Furthermore, α may be arbitrarily determined.
[0098] Note that the method for generating a predicted distribution using differences is not limited to this example. Specifically, the calculation unit 120 may calculate distribution feature information by combining the first difference and the second difference.
[0099] The second difference is the difference between the distribution of the first age group in the second period and the distribution of the previously predicted age group in the second period. For example, the calculation unit 120 identifies the distribution of 51-year-olds in year t-1 and the distribution of 70-year-olds in year t-1. Then, the calculation unit 120 calculates the difference between the distribution of 51-year-olds in year t-1 and the distribution of 70-year-olds in year t-1.
[0100] The distribution of 51-year-olds in year t-1 is shown in matrix M. t-1_51 Let's assume that the distribution of 70-year-olds in year t-1 is represented by matrix M. t-1_70 Let's assume that the first difference is D2. Then the second difference is calculated as shown in equation 11 below.
[0101]
number
[0102] The calculation unit 120 calculates information by combining the first difference and the second difference as distribution feature information. Then, the generation unit 130 integrates the calculated distribution feature information with the distribution of health data for the target age group during the first period. For example, the target distribution M p It can be expressed as shown in the following number 12.
[0103]
number
[0104] [Third example of generating a predicted distribution] The third example describes how to generate a predicted target distribution using distribution feature information based on the relationship between the first distribution (the distribution of the first age group in the first period) and the distribution of health data for age groups corresponding to a predetermined period prior to the first age group in the second period. Specifically, the prediction device 100 generates a predicted target distribution by utilizing the difference between the predicted distribution of the first age group, which is predicted based on the distribution of the second age group in the second period, and the first distribution.
[0105] In this example, the acquisition unit 110 may pre-acquire the predicted distribution of the first age group, which is predicted based on the distribution of the second age group in the second period. For example, the acquisition unit 110 acquires the predicted distribution of 51-year-olds, which is predicted based on the distribution of 50-year-olds in year t-1, as the predicted distribution. The predicted distribution may be a distribution generated by any method. For example, the predicted distribution may be a distribution generated using a machine learning model that is based on past health data. As an example, the predicted distribution may be the distribution of 51-year-olds predicted from the distribution of 50-year-olds in year t-1 using a linear model. The method for generating the predicted distribution is not limited to this example. The predicted distribution may be a distribution generated by a machine learning model based on health data from the second period or earlier, or by a machine learning model based on health data from other datasets.
[0106] The calculation unit 120 calculates the difference between the predicted distribution of the first age group, which is predicted based on the distribution of the second age group in the second period, and the first distribution, as distribution feature information. Assume that the distribution of 51-year-olds, predicted based on the distribution of 50-year-olds in year t-1, is obtained as the predicted distribution. This predicted distribution corresponds to the distribution of 51-year-olds in year t, which is predicted from the health data in year t-1. On the other hand, the distribution of 51-year-olds in year t can be considered the true value for this predicted distribution. In other words, the difference between this predicted distribution and the distribution of 51-year-olds in year t corresponds to the prediction error.
[0107] The distribution of 51-year-olds in year t is shown in matrix M. t_51 Let's assume that the distribution of 51-year-olds predicted from the health data of 50-year-olds in year t-1 is represented by matrix M. p_51 Let's assume the prediction error is D3. Then the prediction error is calculated as shown in equation 13 below.
[0108]
number
[0109] Thus, the calculation unit 120 may calculate the difference between the predicted distribution of the first age group, which is predicted based on the distribution of age groups prior to the first age group by a predetermined period in the second period, and the distribution of the first age group in the first period, as distribution feature information.
[0110] The generation unit 130 integrates the calculated prediction error with the distribution of health data for the target age group during the first period. For example, the target distribution M p It can be expressed as shown in the following number 14.
[0111]
number
[0112] Similar to the first example, the generation unit 130 may generate the predicted distribution as a linear combination of matrices. This example is not limited to this one; each distribution may be represented as a list. Furthermore, α may be arbitrarily determined.
[0113] [Example of operation of prediction device 100] Next, an example of the operation of the prediction device 100 will be explained using Figure 6.
[0114] Figure 6 is a second flowchart illustrating an example of the operation of the prediction device 100. Specifically, Figure 6 is a flowchart illustrating an example of how the prediction device 100 uses the generated prediction distribution to predict future health data of a target person.
[0115] The acquisition unit 110 acquires health data (S101). For example, the acquisition unit 110 acquires health data for a first period and a second period from an external server device. The acquisition unit 110 then stores the health data in the storage device 190. The classification unit 150 processes the health data (S102). For example, the classification unit 150 classifies the health data for each period by age group. The classification unit 150 then generates a probability density distribution for each age group of the health data.
[0116] The acquisition unit 110 acquires a dataset related to health data (S103). Specifically, the acquisition unit 110 acquires the probability density distribution for each age group of health data for the first and second periods from the storage device 190. The acquisition unit 110 also acquires target data, which is the health data of the target person (S104). The target data may be stored in the storage device 190 in advance. Alternatively, the acquisition unit 110 may acquire the target data from an external server or from a terminal device.
[0117] The calculation unit 120 calculates distribution feature information (S105). For example, the calculation unit 120 generates distribution feature information by applying a state transition model generated based on the distribution of the first age group in the second period to the first distribution. In this case, the state transition model may be generated by the state transition model generation unit 170. Alternatively, the calculation unit 120 may calculate the difference between the first distribution and the distribution of the first age group in the second period as distribution feature information. Alternatively, the calculation unit 120 may calculate the difference between the predicted distribution of the first age group predicted based on the distribution of the second age group in the second period and the first distribution as distribution feature information.
[0118] The generation unit 130 generates a predicted target distribution (S106). For example, the generation unit 130 generates a predicted target distribution by integrating the calculated distribution feature information with the distribution of health data for the predicted target age group in the first period.
[0119] The estimation unit 160 estimates the state transition probability (S107). Specifically, the estimation unit 160 estimates the state transition probability based on the transition from the distribution of health data for the first age group in the first period (first distribution) to the predicted distribution. In this case, an optimal transport algorithm may be used to estimate the state transition probability.
[0120] The prediction unit 140 then predicts the health status of the subject (S108). For example, the prediction unit 140 identifies data or data sets in the first distribution corresponding to the target data. Then, based on the state transition probability, the prediction unit 140 predicts the health data for when the subject reaches the predicted age group. For example, the prediction unit 140 predicts the data sets in the predicted distribution, which are the transition destinations based on the state transition probability of the identified data sets, as the health data for when the subject reaches the predicted age group.
[0121] It should be noted that this example of operation is merely one example. In other words, the operation of the prediction device 100 of this disclosure is not limited to this example.
[0122] As described above, the prediction device 100 of the second embodiment acquires health data, which is data relating to the health of multiple individuals, during a first period and a second period that precedes the first period by a predetermined period. The prediction device 100 also calculates distribution feature information that indicates the characteristics of the distribution of health data for the first age group during the first period, based on the relationship between the distribution of health data for the first age group during the first period, which corresponds to the age of the target individual, and the distribution of health data for the second period, which corresponds to the distribution of health data for the first age group or the age group that precedes the first age group by a predetermined period. The prediction device 100 also generates a predicted target distribution by integrating the distribution of health data for the predicted target age group during the first period, which corresponds to the age of the target individual after a predetermined period has elapsed, and the distribution feature information. The prediction device 100 then predicts the health data in the predicted target distribution, which is the transition destination of the data corresponding to the health data of the target individual, from the distribution of health data for the first age group during the first period.
[0123] In other words, the prediction device 100 makes predictions based on the transition from the distribution of health data for the first age group in the first period to the distribution of health data for the target age group. At this time, the prediction device 100 does not use a method that requires time-series data of the same person observed over a certain period in order to make predictions about data concerning the subject over a certain period. That is, the prediction device 100 can make predictions about the health status after a predetermined period has elapsed, even if there is no time-series data for a predetermined period.
[0124] Furthermore, distribution feature information is incorporated into the predicted destination distribution. Distribution feature information is information that indicates the characteristics of the distribution of health data for the first age group in the first period, which is the source of the transition. In other words, the prediction device 100 can make predictions that take into account the individuality of the distribution of health data for the first age group in the first period.
[0125] For example, the prediction device 100 generates a state transition model that predicts the transition of the distribution when moving from one age group to another, based on the relationship between the distribution of health data for one age group in the second period and the distribution of health data for other age groups that are predetermined periods after the first age group in the first period. Then, using the state transition model, the prediction device 100 calculates a post-transition distribution as distribution feature information, which indicates the destination of the distribution of health data for the first age group in the first period when moving from the first age group to the predicted target age group.
[0126] Furthermore, for example, the prediction device 100 calculates a first difference as distribution feature information, which is the difference between the distribution of health data for the first age group in the first period and the distribution of health data for the first age group in the second period.
[0127] Furthermore, the prediction device 100 may calculate distribution feature information by combining the first difference and the second difference. In this case, the second difference is the difference between the distribution of the first age group in the second period and the distribution of the target age group in the second period.
[0128] Furthermore, for example, the prediction device 100 calculates the difference between the predicted distribution of the first age group, which is predicted in the second period based on the distribution of age groups prior to the first age group by a predetermined period, and the distribution of the first age group in the first period, as distribution feature information.
[0129] The predicted distribution, which incorporates such distributional feature information, is a corrected distribution of the distribution of health data for the predicted age group in the first period. In other words, the distributional feature information acts as correction information. Therefore, the prediction device 100 can make predictions with higher accuracy compared to simply using the health data of the predicted age group in the first period as the destination for the health data of the first age group in the first period.
[0130] <Third Embodiment> Next, we will describe the prediction device of the third embodiment. In the third embodiment, we will mainly describe an example in which the prediction device makes predictions using the transition of health data, but the data to be used is not limited to the example described below. Note that some explanations will be omitted as they overlap with the first and second embodiments.
[0131] [Details of prediction device 101] The prediction device 101 is a device that adds further functional units to the prediction device 100. Figure 7 is a block diagram showing an example of the functional configuration of the prediction device 101. The prediction device 101 includes an acquisition unit 110, a calculation unit 120, a generation unit 130, a prediction unit 140, a classification unit 150, and an estimation unit 160. The prediction device 101 may also include a state transition model generation unit 170. Furthermore, the prediction device 101 may include a learning model generation unit 180. In addition, the prediction device 101 may include a storage device 190.
[0132] The prediction device 101, like the prediction device 100, may be a device installed in a terminal device, or it may be a device implemented in a server device that is communicably connected to the terminal device via a wired or wireless network.
[0133] The prediction device 101 calculates the state transition probabilities between each stage in advance. Then, the prediction device 101 generates a learning model using the calculated state transition probabilities. Furthermore, the prediction device 101 uses the generated learning model to predict future health data for the target person. In this embodiment, the stage of generating the learning model is referred to as the generation phase, and the stage of making predictions is referred to as the prediction phase.
[0134] (Generation Phase) Assume that the probability density distribution of health data for each age group during the first and second periods is stored in the memory device 190.
[0135] The acquisition unit 110 acquires the probability density distribution of health data for each age group during the first period and the probability density distribution of health data for each age group during the second period.
[0136] The calculation unit 120 calculates distribution feature information corresponding to each age group. Specifically, it considers the transition from the distribution of one age group in the first period to the distributions of each age group above that age group. That is, the calculation unit 120 calculates the distribution feature information corresponding to each age group above the first age group when they are designated as the target age groups. For example, let's say the source distribution is the distribution for 50-year-olds. In this case, the calculation unit 120 calculates the distribution feature information corresponding to each transition from the distribution for 50-year-olds to the distributions of each age group 51 years and older.
[0137] The calculation unit 120 similarly calculates distribution feature information even when the source distribution is an age group different from 50 years old. In other words, the calculation unit 120 calculates distribution feature information for each combination of the source age group and the predicted target age group from the acquired health data. Here, the source age group corresponds to the role of the first age group.
[0138] The generation unit 130 generates a predicted distribution for each combination of the source age group and the predicted target age group.
[0139] The estimation unit 160 estimates the state transition probability for each combination of the source age group and the target age group, based on the generated predicted distribution. More specifically, the estimation unit 160 estimates the state transition probability for each combination of the source age group and the target age group in the health data for the first period, based on the transition of the distribution of health data.
[0140] The learning model generation unit 180 generates a machine learning model. Specifically, the learning model generation unit 180 generates a predictive model that outputs data in other stages, which are the transition destinations for data in one stage, based on the estimated state transition probabilities. This predictive model corresponds to a machine learning model that has learned the relationship between the data distribution in one stage and the data distribution in other stages. Alternatively, the learning model generation unit 180 may generate a predictive model that outputs a data set of data in other stages, which are the transition destinations for the data set of data distribution in one stage, based on the estimated state transition probabilities. The generated predictive model is a machine learning model that takes age and health data as input and outputs health data, or a data set thereof, to which the input health data will transition after a predetermined period of time.
[0141] In this way, the learning model generation unit 180 generates a machine learning model that learns the relationship between the data in the source stage and the data in the destination stage based on state transition probabilities. More specifically, the learning model generation unit 180 generates a machine learning model that learns the relationship between the health data in the source age group and the health data in the destination age group based on state transition probabilities. The learning model generation unit 180 is an example of a learning model generation means.
[0142] (Prediction Phase) The acquisition unit 110 acquires target data, which is the health data of the target person.
[0143] The prediction unit 140 uses a machine learning model to estimate the future health status of the target person. Specifically, the prediction unit 140 predicts the transition of the target data. For example, the prediction unit 140 inputs the target data and the target person's age into a machine learning model. The machine learning model outputs health data for age groups above the target person's age. That is, it outputs health data for when the target person reaches an age after a predetermined period of time has passed. The prediction unit 140 outputs this health data as the target person's future health data.
[0144] For example, suppose the subject is 51 years old. And suppose we want to predict the health data of the subject when they turn 70. In this case, the prediction unit 140 inputs information indicating that the subject is 51 years old and the subject data into the machine learning model. At this point, the machine learning model outputs health data for 70 years old, which is the age group to which the input health data will transition. The prediction unit 140 outputs the outputted health data as the health data for the subject when they turn 70 years old.
[0145] Furthermore, the machine learning model may be a model that outputs health data after a specific period of time has elapsed for the target data. For example, the machine learning model may output health data for the target person when they reach the age group 20 years from now. Alternatively, the machine learning model may be a model that outputs the changes in health data up to a specific period of time. For example, the prediction unit 140 may output the changes in health data for the target person up to the age group 20 years from now.
[0146] In this way, the prediction unit 140 uses a machine learning model to predict the data in the predicted destination distribution, which is the data transition destination corresponding to the health data of the target person.
[0147] [Example of operation of prediction device 101] Next, an example of the operation of the prediction device 101 will be explained using Figures 8 and 9.
[0148] Figure 8 is a third flowchart illustrating an example of the operation of the prediction device 101. Specifically, Figure 8 is a flowchart illustrating an example of the operation of the prediction device 101 in the generation phase. In the example operation shown in Figure 8, it is assumed that the probability density distribution of health data for each age group is pre-stored in the storage device 190.
[0149] The acquisition unit 110 acquires the probability density distribution of health data for each age group during the first and second periods (S201). For example, the acquisition unit 110 acquires the probability density distribution of health data for each year during the first and second periods.
[0150] The calculation unit 120 generates distribution feature information for each combination of the source age group and the predicted age group (S202). For example, suppose the distribution of health data for each year from 20 to 80 years old has been obtained. If 20 years old is the source age group, the calculation unit 120 calculates the distribution feature information for each age group from 21 to 80 years old as the predicted age group. Furthermore, the calculation unit 120 similarly calculates the distribution feature information when the source age group is changed to 21 to 79 years old.
[0151] The generation unit 130 generates a predicted distribution for each combination of the source age group and the predicted target age group (S203).
[0152] The estimation unit 160 estimates the state transition probability for each combination of the source age group and the predicted target age group (S204). Specifically, the estimation unit 160 uses the predicted target distribution for each combination of the source age group and the predicted target age group to estimate the state transition probability for each combination of the source distribution and the predicted target distribution.
[0153] The learning model generation unit 180 generates a machine learning model based on the estimated state transition probabilities (S205). Specifically, the learning model generation unit 180 generates a machine learning model that learns the relationship between health data in the source age group and health data in the destination age group based on the state transition probabilities.
[0154] Figure 9 is a fourth flowchart illustrating an example of the operation of the prediction device 101. Specifically, Figure 9 is a flowchart illustrating an example of the operation of the prediction device 101 during the prediction phase.
[0155] The acquisition unit 110 acquires target data, which is the health data of the target person (S301). For example, the acquisition unit 110 acquires target data from a terminal device.
[0156] The prediction unit 140 uses a machine learning model to predict the future health status of the subject (S302). Specifically, the prediction unit 140 inputs information indicating the subject's age and target data into the machine learning model. The machine learning model then outputs health data. The prediction unit 140 outputs this output health data as health data for when the subject reaches an age after a predetermined period of time has elapsed.
[0157] It should be noted that this example of operation is merely one example. In other words, the operation of the prediction device 101 of this disclosure is not limited to this example.
[0158] In this way, the prediction device 101 estimates the state transition probability based on the transition of the distribution of health data for each combination of the source age group and the predicted age group in the health data for the first period. The prediction device 101 also generates a machine learning model that learns the relationship between the health data in the source age group and the health data in the transitioned age group based on the state transition probability. Then, the prediction device 101 uses the machine learning model to predict the data in the predicted distribution that corresponds to the data transition destination for the health data of the target person.
[0159] [Example 1] This disclosure primarily describes an example in which a predictive device estimates data transitions based on health data. In other words, it primarily describes examples in which the predictive device is used in the healthcare or medical field. However, the applications of the predictive device are not limited to these examples. For example, the predictive device may also be applied to estimate state transitions of various machines.
[0160] For example, if measurement data is obtained regarding the operating status of a machine, the prediction device may accept the measurement data for each state based on the aging process, from a state where the machine is operating normally to a state where the machine is malfunctioning, as a dataset for each stage. The prediction device may also estimate the state transition probability based on the distribution of measurement data between each state. The prediction device may then output the state transition probability for each transition between states.
[0161] <Example of hardware configuration for a prediction device> The hardware constituting the prediction devices of the first, second, and third embodiments described above will now be explained. Figure 10 is a block diagram showing an example of the hardware configuration of the computer device constituting the prediction device in each embodiment. The computer device 90 implements the prediction device and estimation method described in each embodiment and each modified example. For example, the prediction devices described in each embodiment and each modified example may have the hardware configuration shown in Figure 10.
[0162] As shown in Figure 10, the computer device 90 includes a processor 91, RAM (Random Access Memory) 92, ROM (Read Only Memory) 93, storage device 94, input / output interface 95, bus 96, and drive device 97. Note that the prediction device and the like may be implemented by multiple electrical circuits.
[0163] The storage device 94 stores the program (computer program) 98. The processor 91 executes the program 98 of the prediction device using the RAM 92. Specifically, for example, the program 98 includes a program that causes the computer to execute the processes shown in Figures 2, 6, and 8. The functions of each component of the prediction device are realized in accordance with the execution of the program 98 by the processor 91. The program 98 may also be stored in the ROM 93. Alternatively, the program 98 may be recorded on the recording medium 80 and read using the drive device 97, or it may be transmitted to the computer device 90 from an external device (not shown) via a network (not shown).
[0164] The input / output interface 95 exchanges data with peripheral devices (keyboard, mouse, display device, etc.) 99. The input / output interface 95 functions as a means of acquiring or outputting data. The bus 96 connects each component.
[0165] Furthermore, there are various variations in how the prediction device can be implemented. For example, each component included in the prediction device can be implemented as a dedicated device. Also, each prediction device can be implemented based on a combination of multiple devices.
[0166] The processing method for recording a program to realize each configuration in the function of each embodiment on a recording medium, reading the program recorded on the recording medium as code, and executing it on a computer is also included in the scope of each embodiment. In other words, a computer-readable recording medium is also included in the scope of each embodiment. Furthermore, the recording medium on which the above-mentioned program is recorded, and the program itself, are also included in each embodiment.
[0167] The recording medium in question is, but is not limited to, a floppy disk, hard disk, optical disk, magneto-optical disk, CD (Compact Disc)-ROM, magnetic tape, non-volatile memory card, or ROM. Furthermore, the programs recorded on the recording medium are not limited to programs that perform processing independently, but also include programs that operate on an OS (Operating System) in cooperation with other software and the functions of expansion boards to perform processing, and these are also included in the scope of each embodiment.
[0168] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments. Various modifications to the structure and details of the present invention can be made within the scope of the present invention as can be understood by those skilled in the art.
[0169] Furthermore, the above embodiments and modifications can be combined as appropriate.
[0170] Some or all of the above embodiments may also be described as follows, but are not limited to the following:
[0171] <Note> [Note 1] A means for acquiring health data, which is data relating to the health of multiple individuals, during a first period and a second period that precedes the first period by a predetermined period; A calculation means for calculating distribution feature information that shows the characteristics of the distribution of health data for the first age group in the first period, based on the relationship between the distribution of health data for the first age group in the first period, which corresponds to the age of the subject person, and the distribution of health data for the second period, which corresponds to the distribution of health data for the first age group or the distribution of health data for an age group corresponding to the predetermined period prior to the first age group. A generation means for generating a predicted distribution by integrating the distribution of health data for a predicted target age group corresponding to the age of the target person after the predetermined period has elapsed, and the distribution feature information, The system comprises a prediction means for predicting the health data in the predicted destination distribution, which is the destination of the data corresponding to the health data of the target person, from the distribution of health data of the first age group during the first period, Prediction device.
[0172] [Note 2] The system includes a state transition model generation means that generates a state transition model that predicts the transition of the distribution when moving from one age group to the other age group, based on the relationship between the distribution of health data for one age group during the second period and the distribution of health data for another age group, which is the age group of the one age group after a predetermined period during the first period. The calculation means uses the state transition model to calculate, as distribution feature information, a post-transition distribution that indicates the transition destination of the distribution of health data for the first age group during the first period when the age group moves from the first age group to the predicted target age group. The prediction device described in Appendix 1.
[0173] [Note 3] The calculation means calculates a first difference, which is the difference between the distribution of health data for the first age group during the first period and the distribution of health data for the first age group during the second period, as the distribution feature information. The prediction device described in Appendix 1.
[0174] [Note 4] The calculation means calculates the distribution feature information by combining the first difference and the second difference. The second difference is the difference between the distribution of the first age group during the second period and the distribution of the predicted age group during the second period. The prediction device described in Appendix 3.
[0175] [Note 5] The calculation means calculates the difference between the predicted distribution of the first age group, which is predicted in the second period based on the distribution of age groups prior to the first age group by a predetermined period, and the distribution of the first age group in the first period, as distribution feature information. The prediction device described in Appendix 1.
[0176] [Note 6] The system includes a classification means for classifying the data in the distribution of health data for each age group within each period into data groups. The prediction means predicts the data set of health data in the predicted destination distribution, which is the data set to which the health-classified data set of the target person will transition, from the distribution of health data of the first age group in the first period. The prediction device described in Appendix 1.
[0177] [Note 7] The system includes estimation means for estimating the state transition probability based on the transition from the distribution of the first age group during the first period to the predicted destination distribution, The estimation means estimates the state transition probability by using an optimal transport algorithm that calculates a pair of data before transport and data at the destination, optimizing the cost of transporting from the distribution of the first age group to the predicted destination distribution. The prediction device described in Appendix 1.
[0178] [Note 8] The system further includes a means for generating machine learning models, The estimation means estimates the state transition probability based on the transition of the distribution of health data for each combination of the source age group and the predicted target age group in the health data for the first period. The aforementioned learning model generation means generates a machine learning model that learns the relationship between health data in the source age group and health data in the destination age group based on the state transition probability. The prediction means uses the machine learning model to predict the data in the predicted destination distribution that corresponds to the data transition destination for the health data of the target person. The prediction device described in Appendix 7.
[0179] [Note 9] Health data, which is data concerning the health of multiple individuals, is obtained for a first period and a second period that precedes the first period by a predetermined period. Based on the relationship between the distribution of health data in the first period, which is the distribution of health data for the first age group corresponding to the age of the subject, and the distribution of health data in the second period, which is the distribution of health data for the first age group or the distribution of health data for the age group corresponding to the predetermined period prior to the first age group, distribution feature information indicating the characteristics of the distribution of health data for the first age group in the first period is calculated. A predicted distribution is generated by integrating the distribution in the first period, which is the distribution of health data for the predicted target age group corresponding to the age of the target person after the predetermined period has elapsed, with the distribution feature information. Among the distribution of health data for the first age group during the first period, the health data in the predicted destination distribution that corresponds to the health data of the target person is predicted. Prediction method.
[0180] [Note 10] A process for acquiring health data, which is data relating to the health of multiple individuals, during a first period and a second period that precedes the first period by a predetermined period. A process for calculating distribution feature information that shows the characteristics of the distribution of health data for the first age group in the first period, based on the relationship between the distribution of health data for the first age group in the first period, which corresponds to the age of the subject, and the distribution of health data for the second period, which corresponds to the distribution of health data for the first age group or the distribution of health data for the age group corresponding to the predetermined period prior to the first age group. A process for generating a predicted distribution by integrating the distribution of health data for a predicted target age group corresponding to the age of the target person after the predetermined period has elapsed, and the distribution feature information, during the first period described above. The computer is made to perform the following: a process of predicting the health data in the predicted destination distribution, which is the destination of the data corresponding to the health data of the target person, from the distribution of health data of the first age group during the first period. program.
[0181] Furthermore, some or all of the configurations described in Appendices 2 through 8, which are dependent on Appendice 1 above, may also be dependent on Appendices 9 and 10 in the same way as those described in Appendices 2 through 8. Moreover, within the scope that does not depart from each of the embodiments described above, some or all of the configurations described as appendices may also be dependent on various hardware, software, various recording means for recording software, or systems. [Explanation of Symbols]
[0182] 100, 101 Prediction device 110 Acquisition Department 120 Calculation Unit 130 Generation part 140 Prediction Section 150 Classification Department 160 Estimation Department 170 State Transition Model Generation Unit 180 Learning Model Generation Unit 190 Memory Device
Claims
1. A means for acquiring health data, which is data relating to the health of multiple individuals, during a first period and a second period that precedes the first period by a predetermined period; A calculation means for calculating distribution feature information that shows the characteristics of the distribution of health data for the first age group during the first period, based on the relationship between the distribution of health data for the first age group during the first period, which corresponds to the age of the subject person, and the distribution of health data for the second period, which corresponds to the distribution of health data for the first age group or the distribution of health data for the age group corresponding to the predetermined period prior to the first age group. A generation means for generating a predicted distribution by integrating the distribution of health data for a predicted target age group corresponding to the age of the target person after the predetermined period has elapsed, and the distribution feature information, The system comprises a prediction means for predicting the health data in the predicted destination distribution, which is the destination of the data corresponding to the health data of the target person, from the distribution of health data of the first age group during the first period, Prediction device.
2. The system includes a state transition model generation means that generates a state transition model that predicts the transition of the distribution when moving from one age group to the other age group, based on the relationship between the distribution of health data for one age group during the second period and the distribution of health data for another age group, which is the age group of the one age group after the predetermined period during the first period. The calculation means uses the state transition model to calculate, as distribution feature information, a post-transition distribution that indicates the transition destination of the distribution of health data for the first age group during the first period when the age group moves from the first age group to the predicted target age group. The prediction device according to claim 1.
3. The calculation means calculates a first difference, which is the difference between the distribution of health data for the first age group during the first period and the distribution of health data for the first age group during the second period, as the distribution feature information. The prediction device according to claim 1.
4. The calculation means calculates the distribution feature information by combining the first difference and the second difference. The second difference is the difference between the distribution of the first age group during the second period and the distribution of the predicted age group during the second period. The prediction device according to claim 3.
5. The calculation means calculates the difference between the predicted distribution of the first age group, which is predicted in the second period based on the distribution of age groups prior to the first age group by a predetermined period, and the distribution of the first age group in the first period, as distribution feature information. The prediction device according to claim 1.
6. The system includes a classification means for classifying the data in the distribution of health data for each age group during each period into data groups. The prediction means predicts the data set of health data in the predicted destination distribution, which is the data set to which the health-classified data set of the target person will transition, from the distribution of health data of the first age group during the first period. The prediction device according to claim 1.
7. The system includes estimation means for estimating the state transition probability based on the transition from the distribution of the first age group during the first period to the predicted destination distribution, The estimation means estimates the state transition probability by using an optimal transport algorithm that calculates a pair of data before transport and data at the destination, optimizing the cost of transporting from the distribution of the first age group to the predicted destination distribution. The prediction device according to claim 1.
8. The system further includes a means for generating machine learning models, The estimation means estimates the state transition probability based on the transition of the distribution of health data for each combination of the source age group and the predicted target age group in the health data for the first period. The aforementioned learning model generation means generates a machine learning model that learns the relationship between health data in the source age group and health data in the destination age group based on the state transition probability. The prediction means uses the machine learning model to predict the data in the predicted destination distribution that corresponds to the data transition destination for the health data of the target person. The prediction device according to claim 7.
9. Health data, which is data concerning the health of multiple individuals, is obtained for a first period and a second period that is a predetermined period prior to the first period. Based on the relationship between the distribution of health data in the first period, which is the distribution of health data for the first age group corresponding to the age of the subject, and the distribution of health data in the second period, which is the distribution of health data for the first age group or the distribution of health data for the age group corresponding to the predetermined period prior to the first age group, distribution feature information indicating the characteristics of the distribution of health data for the first age group in the first period is calculated. A predicted distribution is generated by integrating the distribution in the first period, which is the distribution of health data for the predicted target age group corresponding to the age of the target person after the predetermined period has elapsed, with the distribution feature information. Among the distribution of health data for the first age group during the first period, the health data in the predicted destination distribution that corresponds to the health data of the target person is predicted. Prediction method.
10. A process for acquiring health data, which is data relating to the health of multiple individuals, during a first period and a second period that precedes the first period by a predetermined period. A process for calculating distribution feature information that shows the characteristics of the distribution of health data for the first age group in the first period, based on the relationship between the distribution of health data for the first age group in the first period, which corresponds to the age of the subject person, and the distribution of health data for the second period, which corresponds to the distribution of health data for the first age group or the distribution of health data for the age group corresponding to the predetermined period prior to the first age group. A process for generating a predicted distribution by integrating the distribution of health data for a predicted target age group corresponding to the age of the target person after the predetermined period has elapsed, and the distribution feature information, during the first period described above. The computer is made to perform the following: a process of predicting the health data in the predicted destination distribution, which is the destination of the data corresponding to the health data of the target person, from the distribution of health data of the first age group during the first period. program.
Citation Information
Patent Citations
Prediction system, prediction method, and prediction program
JP2018194904A