Domain extension learning device, domain extension learning method, and program
The domain extension learning method and device generate pseudo-health checkup data using a DNN to simulate health checkup data for unknown domains, addressing the limitations of existing data generation methods and enhancing disease risk prediction models with comprehensive data coverage.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2024-10-07
- Publication Date
- 2026-04-17
AI Technical Summary
Existing methods for generating data, such as those described in Patent Document 1, are limited in their ability to create a wide variety of data, particularly in domains not covered by actual health diagnosis data, which is costly and impractical to collect comprehensively.
A domain extension learning method and device that generates pseudo-health checkup data using a deep neural network (DNN) to simulate health checkup data for unknown domains, incorporating a pseudo-data generation unit, domain recognition unit, difference calculation unit, and parameter update unit to minimize differences between specified and predicted domains.
Enables the generation of pseudo-data for unknown domains, allowing comprehensive data coverage and improving the accuracy of disease risk prediction models by generating data that encompasses a wide variety of domains.
Smart Images

Figure 2026066492000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the technology of data generation.
Background Art
[0002] In recent years, the utilization of big data has been progressing. For example, big data is utilized for performing highly accurate predictions by AI (Artificial Intelligence). However, the collection of big data requires financial costs and time costs. In contrast, Patent Document 1 discloses a method for generating pseudo data for use in model learning and the like.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, even with the method of Patent Document 1, it is not always possible to generate a wide variety of data.It is equipped with.
[0007] From another perspective of this disclosure, the domain extension learning method is A domain extension learning method performed by a computer, The generation process is performed to generate simulated health checkup data. A prediction process is performed to predict the domain from the aforementioned pseudo-health checkup data. The system performs a calculation to determine the difference between the predicted domain and the specified domain. Based on the aforementioned difference, the parameters of the generation process are updated.
[0008] In yet another aspect of this disclosure, the program is The generation process is performed to generate simulated health checkup data. A prediction process is performed to predict the domain from the aforementioned pseudo-health checkup data. The system performs a calculation to determine the difference between the predicted domain and the specified domain. Based on the aforementioned difference, the computer is instructed to perform a process to update the parameters of the generation process. [Effects of the Invention]
[0009] This disclosure makes it possible to provide a domain extension learning device capable of generating pseudo-data for unknown domains. [Brief explanation of the drawing]
[0010] [Figure 1] This block diagram shows the hardware configuration of the learning device related to this disclosure. [Figure 2] This block diagram shows the functional configuration of the learning device related to this disclosure. [Figure 3] This shows an example of using a pre-trained pseudo-data generation unit. [Figure 4] This is a flowchart of the processing performed by the learning device related to this disclosure. [Figure 5] This block diagram shows the functional configuration of other learning devices related to this disclosure. [Figure 6]An example of using another trained pseudo-data generation unit is shown. [Figure 7] It is a flowchart of processing by another learning device according to the present disclosure. [Figure 8] It is a block diagram showing the functional configuration of another learning device according to the present disclosure. [Figure 9] It is a flowchart of processing by another learning device according to the present disclosure. [Mode for Carrying Out the Invention]
[0011] Hereinafter, preferred embodiments of the present disclosure will be described with reference to the drawings.
[0012] <First Embodiment> [Overview Explanation] In recent years, in the field of medical healthcare, prediction models for disease risks using big data such as health diagnosis data have been developed. The prediction model predicts the disease risk for patients having various attributes (hereinafter, also referred to as "domains") such as race, gender, disease, age, blood pressure value, etc. In order to perform highly accurate prediction for patients having various domains, learning data including a wide variety of domains is required. However, comprehensively collecting the above-mentioned learning data is unrealistic from the viewpoints of cost, data privacy, differences in data formats for each hospital, etc.
[0013] Therefore, in the present embodiment, a trained model that generates pseudo health diagnosis data for a domain specified by a user is generated. At this time, the user can specify a domain that is not included in the actually collected health diagnosis data. Thereby, data in a range not covered by the actually collected health diagnosis data can be generated.
[0014] In the present embodiment, the actually collected health diagnosis data is also referred to as "actual data", and the pseudo health diagnosis data generated by the learning model is also referred to as "pseudo data".
[0015] Furthermore, in this embodiment, domains included in the actual data are also referred to as "known domains," and domains not included in the actual data are also referred to as "unknown domains." For example, if there is no data for 30-year-old patients in the actual data, then "30 years old" becomes an unknown domain.
[0016] [Hardware configuration] Figure 1 is a block diagram showing the hardware configuration of a learning device 10 according to the first embodiment. The learning device 10 is an example of a domain extension learning device. As shown in the figure, the learning device 10 includes an interface (I / F) 11, a processor 12, a memory 13, a recording medium 14, and a database (DB) 15.
[0017] I / F11 performs data input and output with external devices. Specifically, I / F11 acquires training data used by the learning device 10 from external devices.
[0018] The processor 12 is a computer such as a CPU (Central Processing Unit) and controls the entire learning device 10 by executing a pre-prepared program. The processor 12 may be a GPU (Graphics Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating Point Number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof. The processor 12 performs the learning process described later.
[0019] Memory 13 consists of ROM (Read Only Memory), RAM (Random Access Memory), and other components. Memory 13 stores the deep neural network (DNN) model used by the learning device 10. Memory 13 is also used as working memory while the processor 12 is executing various processes.
[0020] The recording medium 14 is a non-volatile, non-temporary recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from the learning device 10. The recording medium 14 stores various programs that the processor 12 executes. When the learning device 10 performs various processes, the programs stored on the recording medium 14 are loaded into the memory 13 and executed by the processor 12. The DB 15 stores data input via the I / F 11.
[0021] In addition to the above, the learning device 10 may also be equipped with a display device such as a liquid crystal display, and input devices such as a keyboard and mouse. These display devices and input devices are used, for example, by the administrator of the learning device 10 to perform necessary management.
[0022] [Functional Configuration] Figure 2 is a block diagram showing the functional configuration of the learning device 10 of the first embodiment. Functionally, the learning device 10 comprises a pseudo-data generation unit 101, a domain recognition unit 102, a difference calculation unit 103, and a parameter update unit 104. The pseudo-data generation unit 101 is the target of learning and is composed of a DNN or the like.
[0023] The learning device 10 receives random noise, a specified label, and a specified value as input via the I / F 11. The random noise is input to the pseudo-data generation unit 101. The specified label and specified value are input to the difference calculation unit 103.
[0024] The specified label and specified value are domains provided by the user. The learning device 10 trains the pseudo-data generation unit 101 to generate pseudo-data for the domains given as the specified label and specified value. For example, if the user wants to generate pseudo-data for "having diabetes" and "being 25 years old," the user sets "Probability of having diabetes {1,0}" as the specified label and "Age {25}" as the specified value. Note that the specified label is not limited to one-hot labels; it may also be a soft label such as "Probability of having diabetes {0.8,0.2}". Furthermore, the specified value may be a known domain or an unknown domain.
[0025] The pseudo-data generation unit 101 generates pseudo-data from random noise. The pseudo-data is fictitious health checkup data and has the same items as real data. These items include, for example, systolic blood pressure, diastolic blood pressure, fasting blood glucose, and γ-GTP. The pseudo-data generation unit 101 outputs the pseudo-data to the domain recognition unit 102.
[0026] The domain recognition unit 102 comprises a recognition unit 102a and a regression unit 102b. The recognition unit 102a is composed of a classifier or an anomaly detection model pre-trained on real data. The regression unit 102b is composed of a regressor pre-trained on real data.
[0027] The recognition unit 102a makes predictions about race, gender, disease, etc., based on the input pseudo-data, and outputs the predicted labels to the difference calculation unit 103. For example, the recognition unit 102a classifies the pseudo-data and outputs the result as a probability value. The recognition unit 102a uses the output result as the predicted label. An example of output by the recognition unit 102a is shown below. (Output Example 1) Probability values for race: {Japanese, American} {0.2, 0.8} (Output Example 2) Presence or absence of disease: Probability value of {having diabetes} {0.8}
[0028] The regression unit 102b predicts age, blood pressure, BMI, etc., based on the input pseudo-data, and outputs the predicted values to the difference calculation unit 103. For example, the regression unit 102b outputs a scalar value representing age (e.g., 25) and a scalar value representing blood pressure (e.g., 120). The regression unit 102b uses these scalar values as predicted values.
[0029] The domains predicted by the recognition unit 102a and the regression unit 102b are predetermined based on the specified labels and values. For example, if the probability of having diabetes is set as the specified label, the recognition unit 102a will predict the presence or absence of diabetes for the input pseudodata. Similarly, if age is set as the specified value, the regression unit 102b will predict the age for the input pseudodata.
[0030] The difference calculation unit 103 calculates the difference between the specified label and the predicted label, and the difference between the specified value and the predicted value. The difference calculation unit 103 calculates the difference between the specified label and the predicted label using methods such as cross-entropy, temperature-dependent cross-entropy, KL divergence, L1 distance, and L2 distance. The difference calculation unit 103 also calculates the difference between the specified value and the predicted value, for example, the mean squared error between the specified value and the predicted value, or the mean absolute error between the specified value and the predicted value.
[0031] The difference calculation unit 103 adds up the difference between the specified label and the predicted label, and the difference between the specified value and the predicted value, and outputs the sum to the parameter update unit 104.
[0032] The parameter update unit 104 optimizes the parameters of the DNN constituting the pseudo-data generation unit 101 so that the sum of the difference between the specified label and the predicted label, and the difference between the specified value and the predicted value, is minimized. In this way, the pseudo-data generation unit 101 is trained to generate pseudo-data for the specified domain.
[0033] For example, given the specified label and value as "Probability of presence or absence of diabetes {1,0}, age {25}", the pseudo-data generation unit 101 generates pseudo-data such that the recognition unit 102a predicts the presence of diabetes and the regression unit 102b predicts an age of 25, in order to minimize the difference between the specified label and value and the predicted label and value. As a result, the pseudo-data generation unit 101 can arbitrarily generate pseudo-data for a person who has diabetes and is 25 years old.
[0034] [Generating pseudo-data] Figure 3 shows an example of using the trained pseudo-data generation unit 101. As shown in Figure 3, the trained pseudo-data generation unit 101 generates pseudo-data using random noise as input.
[0035] By performing the learning described above, the pseudo-data generation unit 101 can generate pseudo-data for unknown domains not covered by the actual data, allowing the user to obtain data encompassing a wide variety of domains.
[0036] In the above configuration, the pseudo-data generation unit 101 is an example of a generation means, the domain recognition unit 102 is an example of a prediction means, the difference calculation unit 103 is an example of a calculation means, and the parameter update unit 104 is an example of an update means.
[0037] [Learning Process] Next, we will explain the learning process that performs the learning described above. Figure 4 is a flowchart of the learning process by the learning device 10. This process is realized when the processor 12 shown in Figure 1 executes a pre-prepared program and operates as each element shown in Figure 2.
[0038] First, random noise, a specified label, and a specified value are input to the learning device 10 via the I / F 11 (step S101). The random noise is input to the pseudo-data generation unit 101. The specified label and specified value are input to the difference calculation unit 103.
[0039] Next, the pseudo-data generation unit 101 generates pseudo-data from random noise (step S102). The pseudo-data generation unit 101 outputs the pseudo-data to the domain recognition unit 102. Next, the domain recognition unit 102 makes predictions on the input pseudo-data and outputs the predicted labels and predicted values to the difference calculation unit 103 (step S103).
[0040] Next, the difference calculation unit 103 sums the difference between the specified label and the predicted label, and the difference between the specified value and the predicted value, and outputs it to the parameter update unit 104 (step S104). Next, the parameter update unit 104 optimizes the parameters of the DNN constituting the pseudo-data generation unit 101 so that the sum of the difference between the specified label and the predicted label, and the difference between the specified value and the predicted value is minimized (step S105). The processes in steps S102 to S105 are executed repeatedly, and for example, when the sum value falls below a predetermined threshold (step S106: Yes), the process ends.
[0041] [Differentiation] Next, a modified example of the first embodiment will be described.
[0042] Although the above explanation uses health checkup data as an example, the data generated by the pseudo-data generation unit 101 is not limited to this. The learning device of this embodiment can be applied to tabular data that includes items and their values, in addition to health checkup data. For example, the learning device of this embodiment may be trained to generate machine diagnostic data by the pseudo-data generation unit 101. In this case, the domains would include voltage, damage, oil leaks, etc.
[0043] <Second Embodiment> Next, a second embodiment will be described. In the second embodiment, a trained model is generated that takes real data and domain transformation labels as input and generates pseudo-data. The user can specify the domain they want to generate using the domain transformation labels.
[0044] Note that the learning device 20 according to the second embodiment has the same hardware configuration as the learning device 10 according to the first embodiment, so a description will be omitted. The learning device 20 is an example of a domain extension learning device.
[0045] Furthermore, the domains of the second embodiment include combinations of multiple domains. For example, in the second embodiment, "40 years old, Japanese, with diabetes" and "30 years old, without diabetes" are each treated as one domain. In addition, the domains of the second embodiment include at least one continuous variable and any number of categorical variables.
[0046] [Functional Configuration] Figure 5 is a block diagram showing the functional configuration of the learning device 20 according to the second embodiment. Functionally, the learning device 20 comprises a pseudo-data generation unit 201, a domain recognition unit 202, a difference calculation unit 203, and a parameter update unit 204. The pseudo-data generation unit 201 is the target of learning and is composed of a DNN or the like.
[0047] The learning device 20 receives real data, domain conversion labels, and specified domain information via the I / F 11. The real data and domain conversion labels are input to the pseudo-data generation unit 201. The specified domain information is input to the difference calculation unit 203.
[0048] The domain transformation label is a label that represents the difference between the domain of the target pseudo-data (destination) and the domain of the actual data (source). During training of the pseudo-data generation unit 201, a known domain is set as the destination domain.
[0049] Specifically, a domain transformation label is represented by the difference between the source and target continuous variables (such as age or BMI) and the target categorical variable (such as race, gender, or disease). The domain transformation label is set to include the difference of at least one continuous variable. For example, if the user wants to generate pseudo-data of "30 years old, without diabetes" from real data of "40 years old, with diabetes," the user sets the domain transformation label to "-10,0". Here, "-10" represents the difference of -10 years between 30 and 40 years old. Also, "0" is the label representing "no diabetes". Based on the real data and the domain transformation label, the learning device 20 learns the pseudo-data generation unit 201 to transform the real data into pseudo-data.
[0050] The specified domain information represents the typical features of the target domain. For example, if a user wants to generate pseudo-data for "30 years old, no diabetes," they would set the mean or median of the features of the actual data belonging to that domain as the typical features. Alternatively, the specified domain information can also be a label representing the target domain. For example, suppose the label representing "30 years old, no diabetes" is "0" and the label representing "30 years old, diabetes present" is "1." If a user wants to generate pseudo-data for "30 years old, diabetes present," they would set the label "1" representing "30 years old, diabetes present" as the specified domain information.
[0051] The pseudo-data generation unit 201 generates pseudo-data from the actual data and domain conversion labels. The pseudo-data generation unit 201 outputs the generated pseudo-data to the domain recognition unit 202.
[0052] The domain recognition unit 202 consists of a feature extractor or classifier that has been pre-trained to recognize domains from real data.
[0053] If the domain recognition unit 202 is given features as specified domain information, it extracts features from the input pseudo-data and outputs the extracted features to the difference calculation unit 203 as predicted domain information. For example, the domain recognition unit 202 outputs a 128-dimensional feature vector. On the other hand, if the domain recognition unit 202 is given labels as specified domain information, it outputs the probability of belonging to each label from the input pseudo-data and outputs the probability of belonging to each label to the difference calculation unit 203 as predicted domain information. For example, the domain recognition unit 202 outputs "0.2, 0.8" as the probability of belonging to label 0 (30 years old, no diabetes) and label 1 (30 years old, with diabetes).
[0054] The difference calculation unit 203 calculates the difference between the specified domain information and the predicted domain information and outputs it to the parameter update unit 204.
[0055] If feature quantities are provided as specified domain information, the difference calculation unit 203 calculates the difference between the specified domain information and the predicted domain information using methods such as cosine similarity, L1 distance, L2 distance, Chebyshev distance, and Minkowski distance. On the other hand, if domain labels are provided as specified domain information, the difference calculation unit 203 calculates the difference between the specified domain information and the predicted domain information using methods such as cross-entropy, temperature-controlled cross-entropy, KL divergence, L1 distance, and L2 distance.
[0056] The parameter update unit 204 optimizes the parameters of the DNN constituting the pseudo-data generation unit 201 so that the difference between the specified domain information and the predicted domain information is minimized. In this way, the pseudo-data generation unit 201 is trained to generate pseudo-data for the specified domain.
[0057] [Generating pseudo-data] Figure 6 shows an example of using the trained pseudo-data generation unit 201. As shown in Figure 6, the trained pseudo-data generation unit 201 generates pseudo-data using real data and domain transformation labels as input.
[0058] When training the pseudo-data generation unit 201, the user sets a known domain as the target domain for conversion. However, when generating data using the trained pseudo-data generation unit 201, the user can set either a known or unknown domain as the target domain for conversion. For example, if the age range of the actual data is 60s to 90s, the user can set the target domain to an unknown domain, less than 60, when generating the data.
[0059] As a result, the pseudo-data generation unit 201 can generate pseudo-data for unknown domains not covered by the actual data, allowing users to obtain data that encompasses a wide variety of domains.
[0060] In the above configuration, the pseudo-data generation unit 201 is an example of a generation means, the domain recognition unit 202 is an example of a prediction means, the difference calculation unit 203 is an example of a calculation means, and the parameter update unit 204 is an example of an update means.
[0061] [Learning Process] Next, we will explain the learning process that performs the learning described above. Figure 7 is a flowchart of the learning process by the learning device 20. This process is realized when the processor 12 shown in Figure 1 executes a pre-prepared program and operates as each element shown in Figure 5.
[0062] First, the learning device 20 receives real data, domain conversion labels, and specified domain information via the I / F 11 (step S201). The real data and domain conversion labels are input to the pseudo-data generation unit 201. The specified domain information is input to the difference calculation unit 203.
[0063] Next, the pseudo-data generation unit 201 generates pseudo-data from the actual data and domain conversion labels (step S202). The pseudo-data generation unit 201 outputs the generated pseudo-data to the domain recognition unit 202. Next, the domain recognition unit 202 obtains predicted domain information from the input pseudo-data and outputs it to the difference calculation unit 203 (step S203). Next, the difference calculation unit 203 calculates the difference between the specified domain information and the predicted domain information and outputs it to the parameter update unit 204 (step S204).
[0064] Next, the parameter update unit 204 optimizes the parameters of the DNN constituting the pseudo-data generation unit 201 so that the difference between the specified domain information and the predicted domain information is minimized (step S205). The processes in steps S202 to S205 are executed repeatedly, and the process terminates, for example, when the difference falls below a predetermined threshold (step S206: Yes).
[0065] [Differentiation] Next, a modified example of the second embodiment will be described.
[0066] Although the above explanation uses health checkup data as an example, the data generated by the pseudo-data generation unit 201 is not limited to this. The learning device of this embodiment can be applied to tabular data that includes items and their values, in addition to health checkup data. For example, the learning device of this embodiment may be trained to generate machine diagnostic data by the pseudo-data generation unit 201. In this case, the domains would include voltage, damage, oil leaks, or combinations thereof.
[0067] <Third Embodiment> Figure 8 is a block diagram showing the functional configuration of the domain extension learning device according to the third embodiment. The domain extension learning device 300 comprises a generation means 301, a prediction means 302, a calculation means 303, and an update means 304.
[0068] Figure 9 is a flowchart of the processing performed by the domain extension learning device of the third embodiment. The generation means 301 generates pseudo-health checkup data (step S301). The prediction means 302 predicts domains from the pseudo-health checkup data (step S302). The calculation means 303 calculates the difference between the predicted domains and the specified domains (step S303). The update means 304 updates the parameters of the generation means based on the difference (step S304).
[0069] According to the domain extension learning device 300 of the third embodiment, it is possible to generate pseudo-data for unknown domains. This allows users to acquire learning data encompassing a wide variety of domains, enabling them to optimize disease risk prediction models.
[0070] Some or all of the above embodiments may also be described as follows, but are not limited to the following:
[0071] (Note 1) A means for generating simulated health checkup data, A prediction means for predicting a domain from the aforementioned pseudo-health checkup data, A calculation means for calculating the difference between the predicted domain and the specified domain, Based on the difference, an update means for updating the parameters of the generation means, A domain extension learning device equipped with the following features.
[0072] (Note 2) The generation means generates the pseudo-health checkup data from random noise, The prediction means predicts a domain from the simulated health checkup data, outputs a predicted label and a predicted value, The calculation means is a domain extension learning device as described in Appendix 1, which obtains a specified label and a specified value as the specified domain, and calculates the difference between the predicted label and the specified label, and the difference between the predicted value and the specified value.
[0073] (Note 3) The prediction means comprises a recognition means and a regression means, The recognition means predicts categorical variables from the simulated health checkup data and outputs the predicted labels. The regression means is a domain extension learning device as described in Appendix 2, which predicts a continuous variable from the pseudo-health checkup data and outputs the predicted value.
[0074] (Note 4) The generation means generates the pseudo-health checkup data based on the actual health checkup data and the domain conversion label. The prediction means outputs prediction domain information from the simulated health checkup data, The calculation means obtains designated domain information, which is information about known domains, as the designated domain, and calculates the difference between the predicted domain information and the designated domain information, as described in Appendix 1.
[0075] (Note 5) The domain transformation label represents the difference between the domain of the target pseudo-health checkup data (the destination) and the domain of the source actual health checkup data, and includes the difference of at least one continuous variable and any number of destination categorical variables. The domain of the target pseudo-health checkup data to which the conversion is intended is a known domain, as described in Appendix 4 of the domain extension learning device.
[0076] (Note 6) The aforementioned categorical variable includes at least one of race, sex, and disease. The domain extension learning device according to Appendix 3 or Appendix 5, wherein the continuous variable includes at least one of age, BMI, and blood pressure value.
[0077] (Note 7) The specified domain information is a representative feature of the target domain. The prediction means is a domain extension learning device as described in Appendix 5, which extracts features from the pseudo-health checkup data and outputs the extracted features as prediction domain information.
[0078] (Note 8) The specified domain information is a label representing the target domain, The domain extension learning device described in Appendix 5, wherein the prediction means outputs the probability values of assignment of multiple labels from the pseudo-health checkup data and outputs the probability values of assignment as the prediction domain information.
[0079] (Note 9) The generation means is a domain extension learning device as described in Appendix 1, which is configured with a deep learning model.
[0080] (Note 10) A domain extension learning method performed by a computer, The generation process is performed to generate simulated health checkup data. A prediction process is performed to predict the domain from the aforementioned pseudo-health checkup data. The system performs a calculation to determine the difference between the predicted domain and the specified domain. A domain extension learning method that updates the parameters of the generation process based on the aforementioned difference.
[0081] (Note 11) The generation process is performed to generate simulated health checkup data. A prediction process is performed to predict the domain from the aforementioned pseudo-health checkup data. The system performs a calculation to determine the difference between the predicted domain and the specified domain. A program that causes a computer to perform a process to update the parameters of the generation process based on the aforementioned difference.
[0082] Although the present disclosure has been described above with reference to embodiments and examples, the present disclosure is not limited to the above embodiments and examples. Various modifications to the structure and details of the present disclosure can be understood by those skilled in the art within the scope of the present disclosure. [Explanation of Symbols]
[0083] 10, 20 Learning devices 101, 201 Pseudo-data generation unit 102, 202 Domain Recognition Unit 102a Recognition part 102b Regression section 103, 203 Difference calculation part 104, 204 Parameter update section
Claims
1. A means for generating simulated health checkup data, A prediction means for predicting a domain from the aforementioned pseudo-health checkup data, A calculation means for calculating the difference between the predicted domain and the specified domain, Based on the difference, an update means for updating the parameters of the generation means, A domain extension learning device equipped with the following features.
2. The generation means generates the pseudo-health checkup data from random noise, The prediction means predicts a domain from the simulated health checkup data, outputs a predicted label and a predicted value, The domain extension learning device according to claim 1, wherein the calculation means obtains a specified label and a specified value as the specified domain, and calculates the difference between the predicted label and the specified label, and the difference between the predicted value and the specified value.
3. The prediction means comprises a recognition means and a regression means, The recognition means predicts categorical variables from the simulated health checkup data and outputs the predicted labels. The domain extension learning device according to claim 2, wherein the regression means predicts a continuous variable from the pseudo-health checkup data and outputs the predicted value.
4. The generation means generates the pseudo-health checkup data based on the actual health checkup data and the domain conversion label. The prediction means outputs prediction domain information from the simulated health checkup data, The domain extension learning device according to claim 1, wherein the calculation means obtains designated domain information, which is information about a known domain, as the designated domain, and calculates the difference between the predicted domain information and the designated domain information.
5. The domain conversion label represents the difference between the domain of the target pseudo-health checkup data (the target of the conversion) and the domain of the actual health checkup data (the source of the conversion), and includes the difference of at least one continuous variable and any number of target categorical variables. The domain extension learning device according to claim 4, wherein the domain of the target pseudo-health checkup data that is the destination for the conversion is a known domain.
6. The aforementioned categorical variable includes at least one of race, sex, and disease. The domain extension learning device according to claim 3 or claim 5, wherein the continuous variable includes at least one of age, BMI, and blood pressure value.
7. The specified domain information is a representative feature of the target domain. The domain extension learning device according to claim 5, wherein the prediction means extracts features from the pseudo-health checkup data and outputs the extracted features as prediction domain information.
8. The specified domain information is a label representing the target domain, The domain extension learning device according to claim 5, wherein the prediction means outputs probability values of assignment for multiple labels from the pseudo-health checkup data, and outputs the probability values of assignment as the predicted domain information.
9. A domain extension learning method performed by a computer, The generation process is performed to generate simulated health checkup data. A prediction process is performed to predict the domain from the aforementioned pseudo-health checkup data. The system performs a calculation to determine the difference between the predicted domain and the specified domain. A domain extension learning method that updates the parameters of the generation process based on the aforementioned difference.
10. The generation process is performed to generate simulated health checkup data. A prediction process is performed to predict the domain from the aforementioned pseudo-health checkup data. The system performs a calculation to determine the difference between the predicted domain and the specified domain. A program that causes a computer to perform a process to update the parameters of the generation process based on the aforementioned difference.
Citation Information
Patent Citations
Tabular data generation system
JP7402359B1