Method for creating an extended dataset for analysis and computer program for creating an extended dataset for analysis

JP7918224B2Active Publication Date: 2026-09-09CHITOSE LAB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024064783
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2026-09-09
Estimated Expiration
2044-04-12

AI Technical Summary

Benefits of technology

【0017】 以上、本発明によって、サンプル数が少ない場合であっても、より精度の高い解析を行うことができる解析用拡張データセット作成方法及び解析用拡張データセット作成用コンピュータプログラムを提供することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007918224000001
    Figure 0007918224000001
  • Figure 0007918224000002
    Figure 0007918224000002
  • Figure 0007918224000003
    Figure 0007918224000003
Patent Text Reader

Abstract

To provide an analysis expansion dataset creation method capable of performing further highly accurate analysis even in the case where the number of samples is small, and a computer program for analysis expansion dataset creation.SOLUTION: According to the present invention, an analysis expansion dataset creation method includes the steps of: creating a primary dataset on the basis of raw data; and creating an analysis expansion dataset on the basis of the primary dataset. A computer program for analysis expansion dataset creation related to another standpoint of the present invention causes a computer to execute the steps of: creating the primary dataset on the basis of the raw data; and creating the analysis expansion dataset.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for creating an extended dataset for analysis and a computer program for creating an extended dataset for analysis. Background Art

[0002] Against the background of recent digitalization, the importance of data analysis processing has been increasing. By performing data analysis processing, it is possible to obtain more useful information and improve business efficiency and convenience in our daily lives. For example, in the medical field, accumulating a large amount of data for certain clinical cases and analyzing the data makes it possible to establish new treatment methods and utilize the results in preventive medicine.

[0003] Regarding data analysis technology, for example, the following Patent Document 1 discloses a technology of acquiring medical data and generating a learning dataset from the medical data. Prior Art Documents Patent Documents

[0004] Patent Document 1 Japanese Unexamined Patent Application Publication No. 2021-086558 Summary of the Invention Problems to be Solved by the Invention

[0005] However, the technology described in Patent Document 1 only generates a dataset for analysis after selecting necessary data from the acquired medical data. There is a problem that it is difficult to perform high-accuracy analysis when the number of samples of the acquired medical data is originally small, or when there is an extreme difference in the number of samples.

[0006] Therefore, in view of the above problems, the present invention aims to provide a method for creating an extended dataset for analysis and a computer program for creating an extended dataset for analysis that can perform analysis with higher accuracy even when the number of samples is small. [Means for solving the problem]

[0007] A method for creating an extended dataset for analysis according to one aspect of the present invention, which solves the above problems, comprises a primary dataset creation step of creating a primary dataset based on raw data, and an extended dataset creation step of creating an extended dataset for analysis based on the primary dataset.

[0008] Furthermore, although not limited to this perspective, the step of creating an extended dataset for analysis preferably includes at least one of the following steps: a stratification step for the primary dataset, an imbalance correction step, and a validation step.

[0009] Furthermore, although not limited to this perspective, it is preferable that the extended dataset creation step removes personally identifiable data from the primary dataset.

[0010] Furthermore, although not limited to this perspective, it is preferable that the augmented dataset creation step be performed using at least one of the following on the primary dataset: a generative adversarial network (GAN), a flow-based generative model, and a diffusion model.

[0011] Furthermore, a method for creating an extended dataset for analysis according to another aspect of the present invention comprises an extended data creation step of creating extended data based on raw data, and an extended dataset creation step for analysis of an extended dataset based on the extended data, thereby creating an extended dataset for analysis.

[0012] Furthermore, although not limited to this perspective, the step of creating augmented data for analysis preferably includes at least one of the following steps for the primary dataset: a stratification step, an imbalance correction step, and a validation step.

[0013] Furthermore, although not limited to this perspective, it is preferable that the augmented data creation step removes personally identifiable data from the primary dataset.

[0014] Furthermore, although not limited to this perspective, it is preferable that the augmentation data creation step be performed on the primary dataset using at least one of the following: a generative adversarial network (GAN), a flow-based generative model, and a diffusion model.

[0015] Furthermore, a computer program for creating an extended dataset for analysis according to another aspect of the present invention causes a computer to perform a primary dataset creation step of creating a primary dataset based on raw data, and an extended dataset creation step of creating an extended dataset for analysis based on the primary dataset.

[0016] Furthermore, a computer program for creating an extended dataset for analysis according to another aspect of the present invention causes a computer to perform an extended data creation step of creating extended data based on raw data, and an extended dataset creation step of creating an extended dataset for analysis based on the extended data. [Effects of the Invention]

[0017] In summary, the present invention provides a method for creating an extended dataset for analysis and a computer program for creating an extended dataset for analysis that enable more accurate analysis even when the number of samples is small. [Brief explanation of the drawing]

[0018] [Figure 1] FIG. 1 is a diagram showing a process flow of a method for creating an analysis data set according to Embodiment 1. [Figure 2] FIG. 2 is an image diagram of raw data according to Embodiment 1. [Figure 3] FIG. 3 is an image diagram of a raw data set according to Embodiment 1. [Figure 4] FIG. 4 is an image diagram of an extended data set for analysis according to Embodiment 1. [Figure 5] FIG. 5 is a diagram showing a process flow of a method for creating an analysis data set according to Embodiment 2. [Figure 6] FIG. 6 is a diagram showing a primary data set created in an example. [Figure 7] FIG. 7 is a diagram showing an extended data set for analysis created in an example. BEST MODE FOR CARRYING OUT THE INVENTION

[0019] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the present invention can be implemented in many different forms, and is not limited only to the specific examples described in the following embodiments and examples.

[0020] (Embodiment 1) FIG. 1 is a diagram showing a process flow of the method for creating an analysis data set (hereinafter referred to as "the present method") according to the present embodiment.

[0021] As shown in this drawing, the present method includes: (S1-1) a primary data set creation step of creating a primary data set based on raw data; and (S1-2) an extended analysis data set creation step of creating an extended data set for analysis based on the primary data set. According to the present method, there is an advantage that even when the number of raw data samples is small, it is possible to create an extended data set for analysis that enables highly accurate analysis.

[0022] Furthermore, this method is executed by an information processing device, i.e., a computer. Specifically, it is implemented by storing a program on a recording medium such as the computer's hard disk, and then loading this program into a volatile recording medium such as memory as needed and executing it. In other words, this method is implemented by an analysis-extended dataset creation program that causes the computer to perform (S1-1) a primary dataset creation step of creating a primary dataset based on raw data, and (S1-2) an analysis-extended dataset creation step of creating an analysis-extended dataset based on the primary dataset.

[0023] The computer used to run this program is not limited to those that have the above-mentioned functions, but it is preferable, though not limited, to include, a central processing unit (CPU), non-volatile recording media such as hard disks and flash memory, volatile recording media such as memory, a bus connecting these, input devices such as keyboards and mice, and display devices such as monitors, which are components of a typical computer.

[0024] The computer mentioned above may be a laptop or desktop computer, but it can also be a personal digital assistant (PDTA), specifically a smartphone or tablet, which have become increasingly popular in recent years. However, considering information processing capabilities, it is preferable to have a storage medium with sufficient memory capacity and a high-speed CPU rather than a PDTA. A typical PDTA integrates the CPU, display device, and other components into a single cover, and by placing a sensor on the display device to create a touch panel, it can also function as an input device, making it very easy to use. Therefore, it can be used if this aspect is important. In the case of a PDTA, the program that executes this method can be recorded and displayed as an application on the PDTA, and the method can be easily executed by launching this application.

[0025] Here, we will explain this method again. First, this method has a primary dataset creation step (S1-1) in which a primary dataset is created based on raw data.

[0026] Here, "raw data" refers to data obtained before creating a primary dataset, a collection of data acquired from multiple subjects, and raw data that has not undergone any processing for statistical or machine learning. In this embodiment, for example, it refers to a collection of data sets containing information on specific items acquired for each of multiple patients. Specifically, the raw data preferably contains information on specific items for each of the multiple subjects, but is not limited to that. For example, it may include data containing information on the subject's name (name data), data containing information on an identification number (identification number data), data containing information on an address (address data), data containing information on gender (gender data), data containing information on physical characteristics (e.g., weight, height, etc.) (physical data (weight data, height data)), and data containing information on the date of birth. This includes, but is not limited to, data such as birth date data, information about disease names and their history of occurrence and treatment (e.g., whether or not a person has a specific disease) (disease data), data about specific components in the blood (e.g., blood glucose, HbA1c, total protein, albumin, AST, ALT, γ-GTP, creatinine, eGFR, uric acid, HDL cholesterol, LDL cholesterol, triglycerides, red blood cells, hemoglobin, white blood cells, platelet count, etc.) (blood data), and data about genes (e.g., whether or not a person has a mutation in a specific gene, whether or not a gene is expressed) (gene data). Figure 2 shows an example of "raw data". This figure shows an example containing many sets of data in which, for each identification number (ID), sex (SEX), age (AGE), numerical values ​​of specific elements in the blood (blood data, BC1-4), numerical values ​​of expressed genes (gene data, GE1-4), and numerical values ​​related to diseases (disease data, disease) are recorded in the column direction. However, raw data contains information obtained directly from measurements, and may contain missing data points. Furthermore, even clearly abnormal values, such as false detections during measurement, may be recorded as data.

[0027] Furthermore, while the raw data in this example already contains only identification number data, data containing information that can identify an individual, such as name data, address data, and date of birth data (personally identifiable data), is unnecessary for the analysis process. However, since its leakage often causes harm to the individual concerned, it is preferable to delete such data in this step, specifically by deleting personally identifiable data from the raw data. This deletion of personally identifiable data may be done from the raw data, or it may be done after creating the primary dataset. However, deleting personally identifiable data after several processes have been performed to create the primary dataset will incur extra processing effort, so it is preferable to do this as early as possible.

[0028] Furthermore, in this method, the "primary dataset" is a dataset created initially based on raw data and is not an expanded dataset. More specifically, it is created based only on the number of data sets included in the raw data and is distinguished from the expanded dataset described later, where the number of sets increases. Figure 3 shows an example of a primary dataset. This figure shows an example of a dataset created based on the raw data shown in Figure 2. In the example of the primary dataset in this figure, data from columns that were included in Figure 2 but are considered unnecessary for the subsequent expansion process, or data from columns with severe value loss that are deemed difficult to use in the subsequent expansion process, have been deleted. Specifically, in the example in Figure 3, one blood data set (BC4) and one gene data set (GE3) from among multiple gene data sets present in Figure 2 have been deleted. In other words, the primary dataset is basically created by deleting and organizing predetermined data from the raw data.

[0029] Furthermore, it is preferable to perform data correction processing in this step. Here, "data correction processing" refers to processing to fill in missing values ​​in the raw data, or processing to correct values ​​in the data to a normal range if they are abnormal values ​​that are theoretically impossible to measure. In this case, the correction processing is not limited to the above, but if there is a representative value (representative value) in the data, it may be a process of inputting that representative value to fill in the gaps or convert it, or a process of applying dummy variable processing to the raw data and inputting or converting a value that is considered reasonable as a result may be performed.

[0030] Furthermore, it is preferable to perform data filtering in this step. Here, filtering refers to the process of deleting unnecessary data, or more specifically, the process of processing all unnecessary measurement data. As mentioned above, raw data contains data on a large amount of information, but not all of this data is necessary for the analysis, and in some cases, the accuracy of the analysis results may decrease depending on the selection. Therefore, it is possible to improve accuracy by filtering the data. This "filtering process" may simply be the process of deleting item data that the person performing this method clearly does not use, or it is possible to narrow down the necessary items by performing statistical processing on this raw data. Note that this data filtering process may also be performed on the primary dataset. Performing it on the raw dataset has the advantage of reducing the processing burden in subsequent stages, and performing it on the primary dataset may have the advantage of improving the accuracy of the analysis results, depending on the conditions.

[0031] Furthermore, this method includes (S1-2) a step of creating an extended dataset for analysis based on the primary dataset.

[0032] Furthermore, in this method, the "extended dataset for analysis" refers to a dataset created based on the primary dataset and used for analysis, and includes (is processed with) fictitious data sets. An image of this case is shown in Figure 4. The example shown in this figure shows an extended dataset for analysis created based on the primary dataset shown in Figure 3. Specifically, the primary dataset in Figure 3 has up to 110 data sets (identification numbers 1 to 110), but the extended dataset for analysis in Figure 4 adds fictitious data sets, specifically identification numbers 111 to 150, to the primary dataset. In other words, adding fictitious data increases the number of data sets, which has the advantage of enabling more detailed analysis, such as improving the learning accuracy through machine learning processing.

[0033] While there are various methods for adding to this set of data, it is preferable to create it using so-called generative AI rather than randomly adding it. Furthermore, it is preferable, but not limited to, to use at least one of the following algorithms: generative adversarial networks (GANs), flow-based generative models, and diffusion models.

[0034] Here, a "Generative Adversarial Network (GAN)" refers to a program that can learn features from multiple sets of data and generate pseudo-data. Details about GANs can be found in publications such as Goodfellow et al.'s paper published in 2014 (e.g., Goodfellow et al., “Generative adversarial nets,” in Proc. Int. Conf. Neural Inf. Process. Syst., 2014, pp.2672-2680), and this can be utilized.

[0035] Furthermore, "flow-based generative models" here refer to a type of generative AI that utilizes variable transformation rules for probability distributions. Details can be found in, for example, the literature by Ivan Kobyzev et al. (e.g., IEEE Transactions on Pattern Analysis and Machine Intelligence, arXiv:1908.09257v4 [stat.ML] 6 Jun 2020, “Normalizing Flows: An Introduction and Review of Current Methods”), and it is possible to utilize this approach.

[0036] Furthermore, a "diffusion model" here refers to a model that can generate similar data by learning the process of adding noise to the underlying data and destroying it. Details can be found, for example, in the literature by Jascha Sohl-Dickstein et al. (e.g., arXiv:1503.03585v8 [cs.LG] 18 Nov 2015, “Deep Unsupervised Learning using Nonequilibrium Thermodynamics”), and this can be utilized.

[0037] Furthermore, in this step, the number of hypothetical data sets to be generated (the number of data sets undergoing augmentation processing) can be adjusted as appropriate depending on the performance of the information processing device used. However, if the number of data sets in the primary dataset (for example, the number of identification number data) is around 30, the number can be increased by 100 times or more, and in some cases, up to 1000 times. While it is possible to increase the number by more than 10,000 times, as long as the performance of the information processing device allows, it is important to keep it within an appropriate range, because if it is too large, similar data will be created, and the effect of improving the accuracy in its analysis will saturate.

[0038] Furthermore, although not limited to this method, the step of creating an extended dataset for analysis preferably includes at least one of the following steps for the primary dataset: (S1-2-1) stratification, imbalance correction, and validation. Performing these steps makes it possible to perform extension processing by GANs, etc., with greater accuracy.

[0039] Here, the "stratification step" refers to the step of performing a stratification process on the data within the dataset, and "stratification" refers to the process of grouping data sets that have common attributes in the primary dataset and finding features by comparing them. This stratification process is not limited to any one of the following methods: correlation analysis, causal exploration, feature engineering, etc. By performing this stratification step, the accuracy and reliability of the augmented dataset for analysis generated by the data augmentation process can be improved.

[0040] Furthermore, the "imbalance correction step" refers to the step of applying imbalance correction processing to the data within the dataset, and "imbalance correction processing" refers to the process of correcting imbalances in the primary dataset when there are unbalanced data points. For example, if a balanced male-female ratio is desirable for analytical processing, but there is a significant bias between males and females, processing such as deleting data from the gender group that is heavily biased would be appropriate. This makes it possible to improve the reliability of data processing.

[0041] Furthermore, the term "verification step" here refers to the step of performing a verification process on the dataset, and "verification process" refers to the step of performing a verification process to confirm whether the stratification process or imbalance correction process described above is valid. This can be confirmed by performing the stratification process or imbalance correction process again after the initial processing to see if the same process has been performed, but is not limited to this.

[0042] In summary, this embodiment provides a method for creating an extended dataset for analysis and a computer program for creating an extended dataset for analysis, which enable the generation of highly accurate hypothetical data even with a small sample size, and which allows for more accurate analysis by performing machine learning or other analyses based on this data. The effects of this will become clear from the examples described later.

[0043] (Embodiment 2) In Embodiment 1 described above, a primary dataset is created from raw data, and an extended dataset for analysis is created based on this. However, in this embodiment, the difference is that extended data is created from raw data, and then the extended dataset for analysis is created. A detailed explanation follows below, but the same configurations and processes as in Embodiment 1 will be omitted from the explanation.

[0044] Figure 5 is a diagram showing the processing flow of the method for creating an extended dataset for analysis according to this embodiment (hereinafter referred to as "this method").

[0045] As shown in this figure, the method for creating an extended dataset for analysis according to this embodiment (hereinafter referred to as "this method") comprises (S2-1) an extended data creation step of creating extended data based on raw data, and (S2-2) an extended dataset creation step for analysis of creating an extended dataset for analysis based on the extended data.

[0046] Furthermore, this method, like Embodiment 1 described above, can be implemented by recording a program for creating an extended dataset for analysis to perform the above steps on a computer's recording medium and then executing it. The explanation is the same as in Embodiment 1 described above, so it will be omitted.

[0047] First, this method includes (S2-1) an extended data creation step in which extended data is created based on raw data.

[0048] In this embodiment, "raw data" is the same as that described in Embodiment 1 above. On the other hand, "extended data" is data created based on the raw data, and includes not only a large number of data sets that are initially included in the raw data, but also a large number of hypothetical data sets. By creating extended data based on the raw data, it is possible to increase the number of data even with a small number of samples, and the same effects as in Embodiment 1 above can be obtained. The extension processing used in this case is the same as in Embodiment 1 above, and it is also preferable to include stratification processing, etc.

[0049] Furthermore, this method includes (S2-2) a step of creating an extended dataset for analysis based on the extended data. The extended dataset for analysis is the same as in Embodiment 1 above. The process of creating an extended dataset for analysis based on the extended data is the same as in Embodiment 1 above, but the same process as in (S1-1) the primary dataset creation step of creating a primary dataset based on raw data in Embodiment 1 above can be adopted.

[0050] As described above, this embodiment also provides a method for creating an extended dataset for analysis and a computer program for creating an extended dataset for analysis that can perform analysis with higher accuracy even when the number of samples is small, similar to Embodiment 1. [Examples]

[0051] Here, we actually created an extended dataset for analysis from the raw data and confirmed the accuracy of that dataset. Specifically, we used sample data on heart disease that is publicly available on the internet as raw data to verify its effectiveness.

[0052] The sample data used in this embodiment includes age data (Age) containing information about age, sex data (Sex) containing information about gender, and chest pain type data (ChestPainType) containing information about the type of chest pain (TA: typical angina, ATA: atypical angina, NAP: non-anginal pain, ASY: (Asymptomatic), Resting blood pressure data (RestingBP) (mmHg) including information on resting blood pressure, Serum cholesterol data (Cholesterol) (mm / dl) including information on serum cholesterol, Fasting blood glucose data (FastingBS) (1: FastingBS > 120 mg / dl, 0: Otherwise), Resting electrocardiogram results data (RestingECG) including information on resting electrocardiogram results (Normal: Normal, ST: ST-T wave abnormality, LVH: Tendency towards cardiac hypertrophy according to Estes criteria), Maximum heart rate data (MaxHR) including information on maximum heart rate, Exercise Angina data (ExerciseAngina) including information on exercise-induced angina (Y: Yes, N: No), Depression tendency data (Oldpeak) including information on depressive tendencies, Exercise cardiac peak slope data (ST_Slope) including information on the slope of the cardiac peak during exercise (Up: Uphill, Flat: Flat, Down: The data includes information about the presence or absence of heart disease, specifically heart disease data (HeartDisease) (1: heart disease, 0: normal). There were 511 sets of this raw data.

[0053] First, the raw data was corrected to check for missing values ​​and obvious outliers, making it suitable for analysis. Then, based on this, a filtering process was applied to narrow down the data to the variables necessary to predict the presence or absence of heart disease, resulting in the primary dataset. This primary dataset is shown in Figure 6.

[0054] Next, using the primary dataset created above, we focused on Heart Disease and expanded it using Generative Adversarial Networks (GANs) so that the sample sizes for healthy individuals (0 responses) and those with heart disease (1 response) were 0:286 and 1:286, respectively. We then adjusted the sample size to create an expanded dataset for analysis. The results are shown in Figure 7. This expanded dataset contains 572 pairs, representing an addition of 71 pairs.

[0055] (Comparative prediction accuracy) First, using the primary dataset described above, we performed predictions using a gradient boosting classification algorithm, with the target variable being heart disease data (HeartDisease) and the explanatory variables being variables other than heart disease data (HeartDisease). As a result, we confirmed that the prediction accuracy was 0.7282 and that the confusion matrix was as follows. [69,6] [22,6]

[0056] This means that (1) the number of samples predicted to be healthy and actually healthy was 69, (2) the number of samples predicted to have heart disease but actually healthy was 6, (3) the number of samples predicted to be healthy but actually had heart disease was 22, and (4) the number of samples predicted to have heart disease and actually had heart disease was 6. From this, it can be said that this prediction model will achieve a certain level of accuracy if the answer to any question is "healthy". This means that in the operation of prediction models in the real world, predicting that someone is healthy when they actually have heart disease can lead to actual harm due to incorrect judgment. It can also be considered that this is due to a bias in machine learning (the model is learning mostly from healthy data) because the number of healthy people (410) and people with heart disease (101) is unbalanced.

[0057] The sample data used in this study consisted of 410 healthy individuals and 101 individuals with heart disease, resulting in a bias of approximately 4:1. To correct this bias, one possible correction method is to reduce the data from the larger group (healthy individuals). However, this would reduce the overall sample size used for training the predictive model, making it impossible to guarantee sufficient training accuracy.

[0058] (Prediction accuracy in examples) In response, predictions were made using the same method and parameters as above on the expanded dataset created for analysis. As a result, the prediction accuracy was 0.8314, showing improvement. The confusion matrix was as follows. [68,18] [11,75]

[0059] This means that (1) the number of samples predicted to be healthy and actually healthy was 68, (2) the number of samples predicted to have heart disease but actually healthy was 18, (3) the number of samples predicted to be healthy but actually had heart disease was 11, and (4) the number of samples predicted to have heart disease and actually had heart disease was 75. In other words, it can be inferred that by expanding the data, the imbalance in the number of healthy people and people with heart disease was corrected, and the learning accuracy improved as a result.

[0060] In summary, this embodiment has confirmed the effects of the present invention. Specifically, by applying generative AI augmentation processing such as GANs to raw data, it becomes possible to increase the number of hypothetical data sets, thereby providing an analytical dataset that can improve prediction accuracy beyond that of real data. This is particularly promising because it allows for analysis of sets that cannot be analyzed using only real data, by augmenting them. [Industrial applicability]

[0061] The present invention has industrial applicability as a method for creating an extended dataset for analysis and a computer program for creating an extended dataset for analysis.

Claims

1. A primary dataset creation step to create a primary dataset consisting of a set of data based on raw data obtained for each of multiple subjects, including identification number data, gender data, age data, and blood data, and at least one of genetic data and disease data. A method for creating an extended dataset for analysis, comprising: a step of creating an extended dataset for analysis that includes a set of fictitious data based on the primary dataset; At least one of the steps of the primary dataset creation step and the extended dataset creation step for analysis includes a stratification step, an imbalance correction step and a validation step, The stratification step involves using at least one of the following methods—correlation analysis, causal exploration, and feature engineering—to group multiple sets of data that share common attributes, and then performing a stratification process to identify features by comparing the grouped sets of data. The aforementioned imbalance correction step is a step of performing an imbalance correction process to correct the bias between sets of data when there are unbalanced sets of data among the sets of data. A computer-operated method for creating an extended dataset for analysis, wherein the verification step is a step of performing a process to verify whether the stratification process or the imbalance correction process is appropriate for the primary dataset or the extended dataset for analysis after the stratification process or the imbalance correction process has been applied.

2. The method for creating an extended dataset for analysis according to claim 1, wherein the primary dataset creation step includes data correction processing and data filtering processing.

3. The method for creating an extended dataset for analysis according to claim 1, wherein personally identifiable data is deleted in at least one of the steps of creating the primary dataset and creating the extended dataset for analysis.

4. The method for creating an extended dataset for analysis according to claim 1, wherein the step of creating an extended dataset for analysis is performed using at least one of a generative adversarial network (GAN), flow-based generative models, and a diffusion model on the primary dataset.

5. An enhanced data creation step that creates enhanced data including sets of data based on raw data obtained for each of multiple subjects, including identification number data, gender data, age data, and blood data, and at least one of genetic data and disease data, as well as sets of fictitious data. A method for creating an extended dataset for analysis, comprising: a step of creating an extended dataset for analysis consisting of a set of data based on the aforementioned extended data, At least one of the steps of the augmented data creation step and the augmented dataset creation step for analysis includes a stratification step, an imbalance correction step and a verification step, The stratification step involves using at least one of the following methods—correlation analysis, causal exploration, and feature engineering—to group multiple sets of data that share common attributes, and then performing a stratification process to identify features by comparing the grouped sets of data. The aforementioned imbalance correction step is a step of performing an imbalance correction process to correct the bias between sets of data when there are unbalanced sets of data among the sets of data. The verification step is a step of performing a process to verify whether the stratification process or the imbalance correction process is appropriate for the extended data or the extended dataset for analysis after the stratification process or the imbalance correction process has been applied. A method for creating extended datasets for computer analysis.

6. The method for creating an extended dataset for analysis according to claim 5, wherein the extended data creation step involves deleting personally identifiable data from the raw data.

7. The method for creating an extended dataset for analysis according to claim 5, wherein the extended data creation step is performed using at least one of a generative adversarial network (GAN), flow-based generative models, and a diffusion model on the raw data.

8. On the computer, A primary dataset creation step to create a primary dataset consisting of a set of data based on raw data obtained for each of multiple subjects, including identification number data, gender data, age data, and blood data, and at least one of genetic data and disease data. A computer program for creating an extended dataset for analysis, which causes the program to perform the step of creating an extended dataset for analysis that includes a set of fictitious data based on the primary dataset, At least one of the steps of the primary dataset creation step and the extended dataset creation step for analysis includes a stratification step, an imbalance correction step and a validation step, The stratification step involves using at least one of the following methods—correlation analysis, causal exploration, and feature engineering—to group multiple sets of data that share common attributes, and then performing a stratification process to identify features by comparing the grouped sets of data. The aforementioned imbalance correction step is a step of performing an imbalance correction process to correct the bias between sets of data when there are unbalanced sets of data among the sets of data. The verification step is a step of performing a process to verify whether the stratification process or the imbalance correction process is appropriate for the primary dataset or the extended dataset for analysis after the stratification process or the imbalance correction process has been applied. A computer program for creating extended datasets for analysis.

9. On the computer, An enhanced data creation step that creates enhanced data including sets of data based on raw data obtained for each of multiple subjects, including identification number data, gender data, age data, and blood data, and at least one of genetic data and disease data, as well as sets of fictitious data. A computer program for creating an extended dataset for analysis, which causes the program to perform an analytical extended dataset creation step, which creates an analytical extended dataset consisting of a set of data based on the aforementioned extended data, At least one of the steps of the augmented data creation step and the augmented dataset creation step for analysis includes a stratification step, an imbalance correction step and a verification step, The stratification step involves using at least one of the following methods—correlation analysis, causal exploration, and feature engineering—to group multiple sets of data that share common attributes, and then performing a stratification process to identify features by comparing the grouped sets of data. The aforementioned imbalance correction step is a step of performing an imbalance correction process to correct the bias between sets of data when there are unbalanced sets of data among the sets of data. The verification step is a step of performing a process to verify whether the stratification process or the imbalance correction process is appropriate for the extended data or the extended dataset for analysis after the stratification process or the imbalance correction process has been applied. A computer program for creating extended datasets for analysis.

Citation Information

Patent Citations

  • Data selection device, learning device, and program

    JP2021086558A

  • Training data generation program, device and method

    JP2023175296A

  • Processing device, processing method, and program

    WO2023228405A1

  • Model management device, model management system, and model management method

    WO2023238544A1