Disease prediction device, learning model generation device, disease prediction method, learning model generation method, and program

By generating and interpolating graphs from biological data to fill in missing values, the method addresses the variability in training data, enhancing the accuracy of disease prediction systems.

JP7896829B2Active Publication Date: 2026-07-29NEC SOLUTION INNOVATORS LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC SOLUTION INNOVATORS LTD
Filing Date
2024-02-20
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing disease prediction systems face challenges in achieving high accuracy due to variability in training data, as the types of values included in human health and vital data differ among individuals, leading to missing data that hinder effective prediction models.

Method used

A graph-based approach is employed to generate and interpolate missing data points from a dataset containing biological information, constructing a predictive model using machine learning to improve accuracy by filling in missing values and enhancing the training data quality.

Benefits of technology

This method enhances the accuracy of disease prediction by compensating for missing data through graph interpolation, resulting in improved predictive models even when training data is incomplete.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007896829000001
    Figure 0007896829000001
  • Figure 0007896829000002
    Figure 0007896829000002
  • Figure 0007896829000003
    Figure 0007896829000003
Patent Text Reader

Abstract

A learning model generation device 10 comprises: a graph generation unit 11 which generates, from a data group including biometric information of persons and information indicating the presence or absence of occurrence of diseases in the persons, a graph composed of nodes representing data points and edges representing relationships between the nodes; a graph supplementation unit 12 which supplements the generated graph for a deficiency therein; and a model generation unit 13 which generates, from the supplemented graph, a data group in which the deficiency is supplemented, performs machine learning using the generated data group as training data, and generates a prediction model for predicting the occurrence of diseases in a person.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , , ,

[0005] ,

[0003] , , , ,

[0001] The present disclosure relates to a disease prediction apparatus and a disease prediction method for assisting a doctor in diagnosing a disease, and further to a program for realizing these. Mu The present disclosure also relates to a learning model generation apparatus and a learning model generation method for generating a learning model used in the above-described disease prediction apparatus, and further to a program for realizing these. Mu The present disclosure relates to a program for realizing these.

Background Art

[0002] In order to maintain a healthy life of a person, it is important to predict the occurrence of serious diseases that endanger life, such as heart failure, kidney failure, and cerebral infarction. For this reason, for example, Patent Document 1 discloses a system for predicting a person's future diseases using a prediction model.

[0003] In the system disclosed in Patent Document 1, the prediction model is a machine learning model. The prediction model is constructed by machine learning using a person's health data or vital data (explanatory variables) and the presence or absence of the onset of a disease as teacher data (objective variables). Note that examples of health data include data obtained by examinations such as blood tests, for example, weight, height, CCR (creatinine clearance), cholesterol level, and the like. Examples of vital data include data obtained from a living body via a sensor, such as blood pressure, heart rate, respiratory rate, body temperature, wake-up time, bedtime, and sleep time.

[0004] Therefore, the system disclosed in Patent Document 1 acquires the health data or vital data of a patient to be predicted, and inputs the acquired data into a machine learning model. Then, the system disclosed in Patent Document 1 displays the output of the machine learning model as a prediction result. A doctor can make a final diagnosis while checking the prediction result.

Prior Art Documents

Patent Documents

[0005] [Patent Document 1] Japanese Patent Publication No. 2022-169193 [Overview of the project] [Problems that the invention aims to solve]

[0006] By the way, in the system disclosed in Patent Document 1, in order to improve prediction accuracy, it is necessary to prepare as much training data as possible and perform machine learning. However, the types of values ​​included in the human health data and vital data used as training data differ depending on the person from whom the data is collected, and are not uniform. In other words, in the training data obtained from person A, item a is missing, but in the training data obtained from person B, item a is not missing, but item b is missing. For this reason, it is difficult to use a prediction model with high prediction accuracy in the system disclosed in Patent Document 1, and it is difficult to improve prediction accuracy.

[0007] One example of the purpose of this disclosure is to improve the accuracy of disease prediction, even when there are missing data in the training data of a disease prediction model. [Means for solving the problem]

[0008] To achieve the above objective, the learning model generation device in one aspect of this disclosure is: A graph generation unit generates a graph consisting of nodes representing data points and edges representing the relationships between nodes from a data set including a person's biological information and information indicating whether or not the person has a disease. A graph interpolation unit that fills in the missing parts of the generated graph, A model generation unit generates a set of data with missing values ​​filled in from the completed graph, performs machine learning using the generated set of data as training data, and generates a predictive model to predict the occurrence of the disease in the person. It is characterized by having the following features.

[0009] To achieve the above objective, the disease prediction device in one aspect of this disclosure is An information acquisition unit that acquires biometric information of the person to be predicted, A disease prediction unit inputs the acquired biological information into a prediction model and predicts whether or not the person to be predicted will develop a disease based on the output results of the prediction model. Equipped with, The aforementioned prediction model, From a data set containing a person's biological information and information indicating whether or not the person has a disease, a graph is generated consisting of nodes representing data points and edges representing the relationships between nodes. The missing parts of the generated graph are filled in, From the completed graph, a set of data with missing values ​​is generated, The generated data set is used as training data to perform machine learning. It is generated by It is characterized by the following:

[0010] Furthermore, in order to achieve the above objective, the learning model generation method in one aspect of this disclosure is: A graph generation step involves generating a graph from a data set containing a person's biological information and information indicating whether or not the person has a disease, the graph being composed of nodes representing data points and edges representing the relationships between nodes. A graph interpolation step to fill in the missing parts of the generated graph, A model generation step involves generating a set of data with missing values ​​filled in from the completed graph, performing machine learning using the generated set of data as training data, and generating a predictive model to predict the occurrence of the disease in the person. It is characterized by having the following:

[0011] Furthermore, in order to achieve the above objectives, the disease prediction method in one aspect of this disclosure is: The information acquisition step involves obtaining biometric information of the person to be predicted, and Input the obtained biological information into a prediction model, and predict the presence or absence of a disease in the person to be predicted from the output result of the prediction model. A disease prediction step; having The prediction model generates a graph composed of nodes representing data points and edges representing relationships between nodes from a data group including biological information of a person and information indicating the presence or absence of a disease in the person; completes the deficiency of the generated graph; generates a data group with deficiencies complemented from the complemented graph; performs machine learning using the generated data group as training data; and is generated by characterized by this.

[0012] Furthermore, in order to achieve the above object, a first program in one aspect of the present disclosure causes a computer to generate a graph composed of nodes representing data points and edges representing relationships between nodes from a data group including biological information of a person and information indicating the presence or absence of a disease in the person, a graph generation step; completes the deficiency of the generated graph, a graph complementation step; generates a data group with deficiencies complemented from the complemented graph, performs machine learning using the generated data group as training data, and generates a prediction model for predicting the occurrence of a disease in the person, a model generation step; and causes it to execute 、 characterized by this.

[0013] Furthermore, in order to achieve the above object, a second program in one aspect of the present disclosure causes a computer to acquire biological information of a person to be predicted, an information acquisition step; input the obtained biological information into a prediction model, and predict the presence or absence of a disease in the person to be predicted from the output result of the prediction model. A disease prediction step; Execute height, The prediction model generates a graph composed of nodes representing data points and edges representing the relationships between nodes from a data group including human biological information and information indicating the presence or absence of the occurrence of a disease in the person, completes the deficiency of the generated graph, generates a data group with the deficiency completed from the completed graph, and executes machine learning using the generated data group as training data, and is characterized by being generated by this.

Advantages of the Invention

[0014] As described above, according to the present disclosure, even when there are deficiencies in the training data of the prediction model for predicting a disease, the prediction accuracy of the disease can be improved.

Brief Description of the Drawings

[0015] [Figure 1] FIG. 1 is a configuration diagram showing a schematic configuration of an example of a learning model generation device. [Figure 2] FIG. 2 is a configuration diagram showing the configuration of an example of a learning model generation device more specifically. [Figure 3] FIG. 3 is a diagram showing an example of a data group used for generating a graph. [Figure 4] FIG. 4 is a diagram showing an example of a generated graph. [Figure 5] FIG. 5 is a diagram showing an example of a data group with the deficiency completed. [Figure 6] FIG. 6 is a flowchart showing an example of the operation of a learning model generation device. [Figure 7] FIG. 7 is a configuration diagram showing an example of the configuration of a disease prediction device. [Figure 8] FIG. 8 is a flowchart showing an example of the operation of a disease prediction device. <000013Figure 9 is a block diagram showing an example of a computer that implements a learning model generation device and a disease prediction device. [Modes for carrying out the invention]

[0016] (Embodiment 1) In the following description of Embodiment 1, the learning model generation device, learning model generation method, and program will be explained with reference to Figures 1 to 6.

[0017] [Device configuration] First, we will explain the schematic configuration of an example of a learning model generation device using Figure 1. Figure 1 is a configuration diagram showing the schematic configuration of an example of a learning model generation device in Embodiment 1.

[0018] The learning model generation device 10 in Embodiment 1, shown in Figure 1, is a device for generating a learning model for disease prediction. As shown in Figure 1, the learning model generation device 10 comprises a graph generation unit 11, a graph interpolation unit 12, and a model generation unit 13.

[0019] The graph generation unit 11 generates a graph from a data set that includes human biological information and information indicating the presence or absence of disease in a person. The graph consists of nodes representing data points and edges representing the relationships between nodes. The graph completion unit 12 completes the missing data in the graph generated by the graph generation unit 11.

[0020] The model generation unit 13 first generates a data set with missing data points filled in from the graph that has been filled in by the graph completion unit 12. Then, the model generation unit 13 uses the generated data set as training data to perform machine learning and generate a predictive model 20 that predicts the occurrence of diseases in people.

[0021] In this way, the learning model generator 10 compensates for missing data in the training data for the learning model by graph interpolation. The missing data is then filled in with high accuracy through graph interpolation. Therefore, the learning model generator 10 improves the accuracy of disease prediction even when there is variability in the training data for the disease prediction model.

[0022] Next, we will explain the configuration and function of the learning model generation device in detail using Figures 2 to 5. Figure 2 is a configuration diagram showing a more specific example of the configuration of a learning model generation device. Figure 3 is a diagram showing an example of a data set used to generate a graph. Figure 4 is a diagram showing an example of a generated graph. Figure 5 is a diagram showing an example of a data set with missing data imputed.

[0023] First, as shown in Figure 2, the learning model generation device 10 includes, in addition to the graph generation unit 11, graph interpolation unit 12, and model generation unit 13 described above, a prediction model 20, a graph generation model 30, and a graph interpolation model 40.

[0024] The prediction model 20, the graph generation model 30, and the graph completion model 40 are machine learning models. Examples of machine learning models include neural networks. Furthermore, the prediction model 20, the graph generation model 30, and the graph completion model 40 are actually implemented by machine learning programs executed on a computer. These machine learning models may also be built outside the learning model generation device 10. Details of these machine learning models will be described later.

[0025] In Embodiment 1, the data set includes, as described above, human biological information and information indicating the presence or absence of disease in the person. In Embodiment 1, examples of biological information include sex, age, height, weight, blood pressure, creatinine clearance (CCr), Ivmass, cholesterol level, etc. Furthermore, the biological information may also be time-series information acquired over time. Examples of diseases include heart failure, renal failure, cerebral infarction, diabetes, etc.

[0026] In the example shown in Figure 3, the data set includes biometric information for each patient (each ID), such as sex, age, height, weight, creatinine clearance (CCr), and Ivmass. The data set also includes information indicating whether or not heart failure occurred in each patient. The information indicating the presence or absence of disease may also include the presence or absence of disease over time.

[0027] In Embodiment 1, the graph generation unit 11 first acquires the data set shown in Figure 3 from an external server device or the like. As shown in Figure 3, the acquired data set contains missing data. Therefore, if a prediction model is generated using this data set as is, it will be difficult to increase the accuracy of the prediction model.

[0028] Next, the graph generation unit 11 generates a graph 31 (see Figure 4) from the data set in which missing data exists. Figure 4 conceptually shows the generated graph. As shown in Figure 4, the graph 31 consists of a set of numerous nodes 32 and a set of edges 33 connecting the nodes. The nodes 32 represent entities.

[0029] As shown in Figure 4, in graph 31, edges 33 connect nodes 32 whose entities are similar. For example, in the example in Figure 4, since patients A and B have similar attributes, the node representing patient A and the node representing patient B are connected by an edge 33. Each node 32 contains the biometric information of the corresponding patient.

[0030] Furthermore, in Embodiment 1, the graph generation unit 11 generates a graph by inputting the acquired data set into the graph generation model 30. The graph generation model 30 uses machine learning to determine the relationships between nodes that should be connected by edges.

[0031] The graph generation model 30 uses combinations of entities and labels indicating whether or not those combinations should be concatenated as training data. The following literature is referenced for details on the construction of the graph generation model 30. Reference: NEC Technical Report / Vol. 72 No. 1 / Special Feature on AI Creating New Social Value 103, October 2019

[0032] In Embodiment 1, the graph interpolation unit 12 interpolates the missing nodes in the graph 31 generated by the graph generation unit 11 using the graph interpolation model 40. The graph interpolation model 40 is a machine learning model constructed by machine learning the relationship between a graph with some missing nodes and a graph with the missing nodes interpolated.

[0033] To construct the graph completion model 40, features representing the graph with missing nodes and features representing the graph with the missing nodes completed are obtained, and these two sets of features are used as training data. The above-mentioned literature is also referenced for the graph completion model 40.

[0034] When the graph interpolation unit 12 interpolates the missing data in the graph, the model generation unit 13 converts the interpolated graph back into the original data set. Specifically, the model generation unit 13 identifies the nodes corresponding to patients and identifies the non-patient nodes connected to the identified nodes. Then, the model generation unit 13 combines the entities of each identified node into a single record.

[0035] As shown in Figure 5, the conversion process described above fills in any missing data points in the dataset. Figure 5 is an example of filling in the missing data points in the dataset shown in Figure 2. In Figure 5, the gray areas represent the filled-in portions.

[0036] Next, the model generation unit 13 inputs the interpolated data set shown in Figure 5 into the prediction model 20 record by record, and updates the parameters of the prediction model 20 so that the output result matches the presence or absence of disease occurrence in the data set. In Embodiment 1, the prediction model 20 is generated by updating the parameters of the prediction model 20. In other words, in Embodiment 1, the generation of the prediction model includes updating the parameters.

[0037] [Device operation] Next, an example of the operation of the learning model generation device 10 will be explained using Figure 6. Figure 6 is a flowchart showing an example of the operation of the learning model generation device. In the following explanation, Figures 1 to 5 will be referred to as appropriate. In Embodiment 1, the learning model generation method is carried out by operating the learning model generation device 10. Therefore, the explanation of the learning model generation method in Embodiment 1 will be replaced by the following explanation of the operation of the learning model generation device 10.

[0038] As shown in Figure 6, first, the graph generation unit 11 acquires a data set (see Figure 3) from an external server device or the like, which includes human biological information and information indicating whether or not a person has a disease (step A1). The acquired data set contains missing data.

[0039] Next, the graph generation unit 11 inputs the data set containing missing data into the graph generation model 30 to generate graph 31 (see Figure 4) (step A2). Because there are missing data in the data set, there are also missing data in graph 31.

[0040] Next, the graph interpolation unit 12 uses the graph interpolation model 40 to interpolate the missing data in the graph 31 generated by the graph generation unit 11 in step A2 (step A3). Specifically, the graph interpolation unit 12 inputs the graph 31 into the graph interpolation model 40 to interpolate the missing data in the graph 31.

[0041] Next, the model generation unit 13, after the missing data in the graph has been filled in step A3, converts the filled-in graph back into the original data set (step A4). In the data set obtained in step A4, the missing data has been filled in (see Figure 5).

[0042] Next, the model generation unit 13 inputs the data set obtained in step A4 into the prediction model 20 record by record, and updates the parameters of the prediction model 20 so that the output results match the presence or absence of disease occurrence in the data set (step A5). The prediction model 20 is finally generated by updating the parameters in step A5.

[0043] As described above, in Embodiment 1, missing data in the training data set is imputed by graph interpolation. A predictive model is then constructed using the data set with the missing data imputed. Furthermore, since the imputation is performed by graph interpolation, the accuracy of the imputation is considered high. Therefore, according to Embodiment 1, even if there is variability in the training data of the predictive model that predicts diseases, it is possible to improve the accuracy of disease prediction.

[0044] [program] The program in Embodiment 1 can be any program that causes a computer to execute steps A1 to A5 shown in Figure 6. By installing and running this program on a computer, the learning model generation device 10 and the learning model generation method in Embodiment 1 can be realized. In this case, the computer's processor functions as a graph generation unit 11, a graph interpolation unit 12, and a model generation unit 13, and performs the processing. Examples of computers include general-purpose PCs, smartphones, and tablet devices.

[0045] Furthermore, the program in Embodiment 1 may be executed by a computer system constructed by multiple computers. In this case, for example, each computer may function as one of the graph generation unit 11, the graph interpolation unit 12, and the model generation unit 13, respectively.

[0046] (Embodiment 2) Next, in Embodiment 2, the disease prediction device, disease prediction method, and program will be described with reference to Figures 7 and 8.

[0047] [Device configuration] First, we will explain the configuration of an example of a disease prediction device using Figure 7. Figure 7 is a configuration diagram showing the configuration of an example of a disease prediction device.

[0048] The disease prediction device 50 in Embodiment 2, shown in Figure 7, is a device for predicting human diseases. As shown in Figure 7, the disease prediction device 50 includes an information acquisition unit 51 and a disease prediction unit 52.

[0049] The information acquisition unit 51 acquires biometric information of the person to be predicted from an external server device or the like. As with Embodiment 1, the biometric information may include gender, age, height, weight, blood pressure, creatinine clearance (CCr), IVmass, cholesterol level, etc.

[0050] The disease prediction unit 52 inputs the biological information acquired by the information acquisition unit 51 into the prediction model 20, and predicts whether or not a disease will occur in the person to be predicted based on the output of the prediction model 20.

[0051] The prediction model 20 is generated by the following steps (a) to (d). (a) A step of generating a graph from a set of data including human biological information and information indicating whether or not a person has a disease, the graph consisting of nodes representing data points and edges representing the relationships between nodes. (b) A step to fill in any missing information in the generated graph. (c) A step of generating a data set with missing values ​​filled in from the interpolated graph. (d) A step in which machine learning is performed using the generated dataset as training data.

[0052] In other words, the prediction model 20 is the same as the prediction model 20 described in Embodiment 1. When the disease prediction unit 52 receives the biological information of the person to be predicted, it inputs this biological information into the prediction model 20 described in Embodiment 1 to predict the person's disease.

[0053] [Device operation] Next, an example of the operation of the disease prediction device 50 will be explained using Figure 8. Figure 8 is a flowchart showing an example of the operation of the disease prediction device. In the following explanation, Figure 7 will be referred to as appropriate. In Embodiment 2, the disease prediction method is performed by operating the disease prediction device 50. Therefore, the explanation of the disease prediction method in Embodiment 2 will be replaced by the following explanation of the operation of the disease prediction device 50.

[0054] As shown in Figure 8, first, the information acquisition unit 51 acquires biometric information of the person to be predicted from an external server device or the like (step B1).

[0055] Next, the disease prediction unit 52 inputs the biological information acquired by the information acquisition unit 51 in step B1 into the prediction model 20, and predicts whether or not the disease will occur in the person to be predicted based on the output of the prediction model 20 (step B2).

[0056] Specifically, in step B2, the disease prediction unit 52 transmits the output results of the prediction model 20 to the terminal device of the doctor or other person who requested the prediction. This allows the doctor or other person to confirm the predicted disease for the person (patient) being predicted.

[0057] As described above, in Embodiment 2, the predictive model constructed in Embodiment 1 is used, so highly accurate disease prediction results can be obtained.

[0058] [program] The program in Embodiment 2 can be any program that causes a computer to execute steps B1 to B2 shown in Figure 8. By installing and running this program on a computer, the disease prediction device 50 and disease prediction method in Embodiment 2 can be realized. In this case, the computer's processor functions as an information acquisition unit 51 and a disease prediction unit 52, and performs processing. Examples of computers include general-purpose PCs, smartphones, and tablet devices.

[0059] Furthermore, the program in Embodiment 2 may be executed by a computer system constructed by multiple computers. In this case, for example, each computer may function as either an information acquisition unit 51 or a disease prediction unit 52.

[0060] [Physical configuration] Here, we will explain, using Figure 9, a computer that implements a learning model generation device 10 and a disease prediction device 50 by executing a program. Figure 9 is a block diagram showing an example of a computer that implements a learning model generation device and a disease prediction device.

[0061] As shown in Figure 9, the computer 110 comprises a CPU (Central Processing Unit) 111, main memory 112, storage device 113, input interface 114, display controller 115, data reader / writer 116, and communication interface 117. Each of these components is connected to the others via a bus 121, enabling data communication.

[0062] Furthermore, the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) in addition to, or instead of, the CPU 111. In this embodiment, the GPU or FPGA can execute the program in the embodiment.

[0063] The CPU 111 loads the program in the embodiment, which consists of a set of codes stored in the storage device 113, into the main memory 112, and performs various calculations by executing each code in a predetermined order. The main memory 112 is typically a volatile storage device such as DRAM (Dynamic Random Access Memory).

[0064] Furthermore, the program in this embodiment is provided stored on a computer-readable recording medium 120. The program in this embodiment may also be distributed over the internet via a communication interface 117.

[0065] Specific examples of the storage device 113 include hard disk drives and semiconductor storage devices such as flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and mouse. The display controller 115 is connected to the display device 119 and controls the display on the display device 119.

[0066] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, reads programs from the recording medium 120, and writes processing results from the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.

[0067] Furthermore, specific examples of the recording medium 120 include general-purpose semiconductor memory devices such as CF (Compact Flash®) and SD (Secure Digital), magnetic recording media such as Flexible Disks, or optical recording media such as CD-ROMs (Compact Disk Read Only Memory).

[0068] Furthermore, the learning model generation device 10 and the disease prediction device 50 can be implemented not by a computer with a program installed, but by using hardware corresponding to each part, such as electronic circuits. Moreover, the learning model generation device 10 and the disease prediction device 50 may be partially implemented by a program and the remaining parts by hardware. The computer is not limited to the one shown in Figure 9.

[0069] Some or all of the embodiments described above can be expressed by (Appendix 1) to (Appendix 18) described below, but are not limited to the following descriptions.

[0070] (Note 1) A graph generation unit generates a graph consisting of nodes representing data points and edges representing the relationships between nodes from a data set including a person's biological information and information indicating whether or not the person has a disease. A graph interpolation unit that fills in the missing parts of the generated graph, A model generation unit generates a set of data with missing values ​​filled in from the completed graph, performs machine learning using the generated set of data as training data, and generates a predictive model to predict the occurrence of the disease in the person. It is equipped with A learning model generation device characterized by the following features.

[0071] (Note 2) The graph interpolation unit interpolates the missing nodes in the generated graph using a machine learning model constructed by learning the relationship between a graph with some missing nodes and a graph with the missing nodes interpolated. The learning model generation device described in Appendix 1.

[0072] (Note 3) The aforementioned biometric information includes time-series information obtained from multiple individuals in a time-series manner. The learning model generation device described in Appendix 1.

[0073] (Note 4) Information indicating whether or not the person has developed the disease includes whether or not the disease has developed at each time interval, The learning model generation device described in Appendix 3.

[0074] (Note 5) The aforementioned disease is heart failure, renal failure, cerebral infarction, or diabetes. The learning model generation device described in Appendix 1.

[0075] (Note 6) An information acquisition unit that acquires biometric information of the person to be predicted, A disease prediction unit inputs the acquired biological information into a prediction model and predicts whether or not the person to be predicted will develop a disease based on the output results of the prediction model. Equipped with, The aforementioned prediction model, From a data set containing a person's biological information and information indicating whether or not the person has a disease, a graph is generated consisting of nodes representing data points and edges representing the relationships between nodes. The missing parts of the generated graph are filled in, From the completed graph, a set of data with missing values ​​is generated, The generated data set is used as training data to perform machine learning. It is generated by A disease prediction device characterized by the following features.

[0076] (Note 7) A graph generation step involves generating a graph from a data set containing a person's biological information and information indicating whether or not the person has a disease, the graph being composed of nodes representing data points and edges representing the relationships between nodes. A graph interpolation step to fill in the missing parts of the generated graph, A model generation step involves generating a set of data with missing values ​​filled in from the completed graph, performing machine learning using the generated set of data as training data, and generating a predictive model to predict the occurrence of the disease in the person. Having, A method for generating a learning model characterized by the following features.

[0077] (Note 8) In the graph interpolation step, the missing nodes in the generated graph are interpolated using a machine learning model constructed by learning the relationship between a graph with some missing nodes and a graph with the missing nodes interpolated. The learning model generation method described in Appendix 7.

[0078] (Note 9) The aforementioned biometric information includes time-series information obtained from multiple individuals in a time-series manner. The learning model generation method described in Appendix 7.

[0079] (Note 10) Information indicating whether or not the person has developed the disease includes whether or not the disease has developed at each time interval, The learning model generation method described in Appendix 9.

[0080] (Note 11) The aforementioned disease is heart failure, renal failure, cerebral infarction, or diabetes. The learning model generation method described in Appendix 7.

[0081] (Note 12) The information acquisition step involves obtaining biometric information of the person to be predicted, and A disease prediction step involves inputting the acquired biometric information into a prediction model and predicting, from the output results of the prediction model, whether or not the person to be predicted has developed a disease. It has, The aforementioned prediction model, From a data set containing a person's biological information and information indicating whether or not the person has a disease, a graph is generated consisting of nodes representing data points and edges representing the relationships between nodes. The missing parts of the generated graph are filled in, From the completed graph, a set of data with missing values ​​is generated, The generated data set is used as training data to perform machine learning. It is generated by A disease prediction method characterized by the following features.

[0082] (Note 13) On the computer, A graph generation step involves generating a graph from a data set containing a person's biological information and information indicating whether or not the person has a disease, the graph being composed of nodes representing data points and edges representing the relationships between nodes. A graph interpolation step to fill in the missing parts of the generated graph, A model generation step involves generating a set of data with missing values ​​filled in from the completed graph, performing machine learning using the generated set of data as training data, and generating a predictive model to predict the occurrence of the disease in the person. Let's execute it ru, pu Logra Hmm.

[0083] (Note 14) In the graph interpolation step, the missing nodes in the generated graph are interpolated using a machine learning model constructed by learning the relationship between a graph with some missing nodes and a graph with the missing nodes interpolated. As described in Appendix 13 program .

[0084] (Note 15) The aforementioned biometric information includes time-series information obtained from multiple individuals in a time-series manner. As described in Appendix 13 program .

[0085] (Note 16) Information indicating whether or not the person has developed the disease includes whether or not the disease has developed at each time interval, As described in Appendix 15 program .

[0086] (Note 17) The aforementioned disease is heart failure, renal failure, cerebral infarction, or diabetes. As described in Appendix 13 program .

[0087] (Note 18) On the computer, The information acquisition step involves obtaining biometric information of the person to be predicted, and A disease prediction step involves inputting the acquired biometric information into a prediction model and predicting, from the output results of the prediction model, whether or not the person to be predicted has developed a disease. Execute height, The aforementioned prediction model, From a data set containing a person's biological information and information indicating whether or not the person has a disease, a graph is generated consisting of nodes representing data points and edges representing the relationships between nodes. The missing parts of the generated graph are filled in, From the completed graph, a set of data with missing values ​​is generated, The generated data set is used as training data to perform machine learning. It is generated by program .

[0088] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments. Various modifications to the structure and details of the present invention can be made, as can be understood by those skilled in the art within the scope of the present invention.

[0089] This application claims priority based on Japanese Patent Application No. 2023-026656, filed on 22 February 2023, and incorporates all of its disclosures herein. [Industrial applicability]

[0090] As described above, this disclosure makes it possible to improve the accuracy of disease prediction even when there are missing data in the training data of a predictive model that predicts diseases. This disclosure is useful for various systems related to disease diagnosis. [Explanation of Symbols]

[0091] 10. Learning Model Generator 11 Graph Generation Unit 12 Graph interpolation section 13 Model Generation Unit 20 Predictive Models 30 Graph Generation Models 31 Graph 32 nodes 33 Edge 40 Graph Complementation Models 50 Disease prediction devices 51 Information Acquisition Department 52 Disease Prediction Department 110 Computer 111 CPU 112 Main Memory 113 Storage device 114 Input Interface 115 Display Controller 116 Data Readers / Writers 117 Communication Interface 118 Input devices 119 Display device 120 recording media 121 Bus

Claims

1. A graph generation unit generates a graph consisting of nodes representing data points and edges representing the relationships between nodes from a data set including a person's biological information and information indicating whether or not the person has a disease. A graph interpolation unit that fills in the missing parts of the generated graph, A model generation unit generates a set of data with missing values ​​filled in from the completed graph, performs machine learning using the generated set of data as training data, and generates a predictive model to predict the occurrence of the disease in the person. It is equipped with A learning model generation device characterized by the following features.

2. The graph interpolation unit interpolates the missing nodes in the generated graph using a machine learning model constructed by learning the relationship between a graph with some missing nodes and a graph with the missing nodes interpolated. A learning model generation device according to claim 1.

3. The aforementioned biometric information includes time-series information obtained from multiple individuals in a time-series manner. A learning model generation device according to claim 1.

4. Information indicating whether or not the person has developed the disease includes whether or not the disease has developed at each time interval, A learning model generation device according to claim 3.

5. The aforementioned disease is heart failure, renal failure, cerebral infarction, or diabetes. A learning model generation device according to claim 1.

6. An information acquisition unit that acquires biometric information of the person to be predicted, A disease prediction unit inputs the acquired biological information into a prediction model and predicts whether or not the person to be predicted will develop a disease based on the output results of the prediction model. Equipped with, The aforementioned prediction model, From a data set containing a person's biological information and information indicating whether or not the person has a disease, a graph is generated consisting of nodes representing data points and edges representing the relationships between nodes. The missing parts of the generated graph are filled in, From the completed graph, a set of data with missing values ​​is generated, The generated data set is used as training data to perform machine learning. It is generated by A disease prediction device characterized by the following features.

7. A graph is generated from a data set containing a person's biological information and information indicating whether or not the person has a disease, consisting of nodes representing data points and edges representing the relationships between nodes. The missing parts of the generated graph are filled in, From the completed graph, a set of data with missing values ​​is generated, and machine learning is performed using the generated set of data as training data to generate a predictive model that predicts the occurrence of the disease in the person. A method for generating a learning model characterized by the following features.

8. In the interpolation of the aforementioned graph, a machine learning model constructed by machine learning the relationship between a graph with some missing nodes and a graph with the missing nodes interpolated is used to interpolate the missing nodes in the generated graph. The method for generating a learning model according to claim 7.

9. The aforementioned biometric information includes time-series information obtained from multiple individuals in a time-series manner. The method for generating a learning model according to claim 7.

10. Information indicating whether or not the person has developed the disease includes whether or not the disease has developed at each time interval, The method for generating a learning model according to claim 9.

11. The aforementioned disease is heart failure, renal failure, cerebral infarction, or diabetes. The method for generating a learning model according to claim 7.

12. By acquiring biometric information of the person to be predicted, The acquired biometric information is input into a prediction model, and the presence or absence of disease in the person to be predicted is predicted from the output of the prediction model. The aforementioned prediction model, From a data set containing a person's biological information and information indicating whether or not the person has a disease, a graph is generated consisting of nodes representing data points and edges representing the relationships between nodes. The missing parts of the generated graph are filled in, From the completed graph, a set of data with missing values ​​is generated, The generated data set is used as training data to perform machine learning. It is generated by A disease prediction method characterized by the following features.

13. On the computer, From a data set including a person's biological information and information indicating whether or not the person has a disease, a graph is generated consisting of nodes representing data points and edges representing the relationships between nodes. The missing values ​​in the generated graph are filled in. From the completed graph, a set of data with missing values ​​is generated, and machine learning is performed using the generated set of data as training data to generate a predictive model that predicts the occurrence of the disease in the person. program.

14. In the interpolation of the aforementioned graph, a machine learning model constructed by machine learning the relationship between a graph with some missing nodes and a graph with the missing nodes interpolated is used to interpolate the missing nodes in the generated graph. The program according to claim 13.

15. The aforementioned biometric information includes time-series information obtained from multiple individuals in a time-series manner. The program according to claim 13.

16. Information indicating whether or not the person has developed the disease includes whether or not the disease has developed at each time interval, The program according to claim 15.

17. The aforementioned disease is heart failure, renal failure, cerebral infarction, or diabetes. The program according to claim 13.

18. On the computer, Obtain biometric information from the person to be predicted, The acquired biological information is input into a prediction model, and the presence or absence of disease in the person to be predicted is predicted from the output results of the prediction model. The aforementioned prediction model, From a data set containing a person's biological information and information indicating whether or not the person has a disease, a graph is generated consisting of nodes representing data points and edges representing the relationships between nodes. The missing parts of the generated graph are filled in, From the completed graph, a set of data with missing values ​​is generated, The generated data set is used as training data to perform machine learning. It is generated by program.