Information processing device, information processing method, and program

The information processing device addresses the challenge of evaluating synthetic data quality and confidentiality by generating and comparing composite datasets, ensuring effective machine learning performance.

JP7855772B1Active Publication Date: 2026-05-08AIOI INSURANCE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
AIOI INSURANCE CO LTD
Filing Date
2025-07-30
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies face challenges in evaluating both the confidentiality and quality of synthetic data used for machine learning, leading to decreased performance due to excessive information alteration or reduction, without advanced expertise.

Method used

An information processing device that generates multiple composite datasets, performs machine learning using these datasets and the original dataset, evaluates learning accuracy, and presents evaluation results for comparison, enabling easy assessment of confidentiality and usefulness.

Benefits of technology

Facilitates easy evaluation of both confidentiality and usefulness of synthetic data as machine learning data, allowing users to select highly accurate synthetic data without specialized knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007855772000001_ABST
    Figure 0007855772000001_ABST
Patent Text Reader

Abstract

This invention provides an information processing device that allows for easy evaluation of both the confidentiality of synthesized data and its usefulness as machine learning data. [Solution] The system comprises: a data generation unit that generates multiple synthetic datasets based on an original dataset; a machine learning execution unit that automatically performs machine learning on an AI model using each synthetic dataset and the original dataset as training data; a learning accuracy evaluation unit that evaluates and outputs the accuracy of the machine learning; an accuracy evaluation unit that evaluates and outputs the accuracy of each synthetic dataset; and an evaluation result presentation unit that presents the accuracy evaluation results of each synthetic dataset and the accuracy evaluation results of the machine learning so that they can be compared.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] For AI (Artificial Intelligence) development, it is necessary to perform machine learning using a large amount of high-quality data. On the other hand, when performing AI development using data owned and managed by other companies, if the data contains personal information or confidential information, it may be difficult to obtain the data, and AI development may be inhibited.

[0003] As a countermeasure when personal information or the like is included in data, for example, Patent Document 1 discloses an apparatus that generates secondary data in which information is changed or deleted so as not to exceed a preset anonymity evaluation value from primary data including personal information.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, when performing machine learning using synthetic data in which information has been changed or deleted from the viewpoint of anonymity, the performance of machine learning may decrease because the information has changed too much or the amount of information has become too small. Therefore, it is desirable to use data that maintains both anonymity and quality as machine learning data. However, it has been difficult to evaluate both anonymity and the quality of machine learning data without advanced expertise.

[0006] Therefore, the present invention aims to provide an information processing device that can easily evaluate both the confidentiality of synthesized data and its usefulness as machine learning data. [Means for solving the problem]

[0007] An information processing device according to one aspect of the present invention comprises: a data generation unit that generates a plurality of composite datasets based on an original dataset; a machine learning execution unit that automatically performs machine learning of an AI model using each composite dataset and the original dataset as training data; a learning accuracy evaluation unit that evaluates and outputs the accuracy of the machine learning; an accuracy evaluation unit that evaluates and outputs the accuracy of each composite dataset; and an evaluation result presentation unit that presents the evaluation result of the accuracy of the machine learning for each composite dataset so that it can be compared with the evaluation result of the accuracy of the machine learning.

[0008] An information processing method according to one aspect of the present invention comprises the steps of: a computer generating a plurality of synthetic datasets based on an original dataset; a computer automatically performing machine learning on an AI model using each of the synthetic datasets and the original dataset as training data; a computer evaluating and outputting the accuracy of the machine learning; a computer evaluating and outputting the degree of confidentiality of each synthetic dataset; and a computer presenting the evaluation results of the degree of confidentiality of each synthetic dataset and the evaluation results of the accuracy of the machine learning in a way that allows for comparison.

[0009] A program according to one aspect of the present invention causes a computer to function as a data generation unit that generates multiple synthetic datasets based on an original dataset; a machine learning execution unit that automatically performs machine learning of an AI model using each synthetic dataset and the original dataset as training data; a learning accuracy evaluation unit that evaluates and outputs the accuracy of the machine learning; an accuracy evaluation unit that evaluates and outputs the accuracy of each synthetic dataset; and an evaluation result presentation unit that presents the evaluation results of the accuracy of the machine learning for each synthetic dataset so that they can be compared. [Effects of the Invention]

[0010] According to the present invention, it is possible to provide an information processing device that can easily evaluate both the confidentiality of synthesized data and its usefulness as machine learning data. [Brief explanation of the drawing]

[0011] [Figure 1] This figure shows the configuration of the information processing system 1 according to this embodiment. [Figure 2] A block diagram showing the configuration of the information processing device 10 according to this embodiment. [Figure 3] A block diagram showing the configuration of the user terminal 20 according to this embodiment. [Figure 4] This block diagram shows a functional module of a program executed by the processor 11 of the information processing device 10 according to this embodiment. [Figure 5] A flowchart of the generation and evaluation process of synthesized data by the information processing system 1 according to this embodiment. [Figure 6] A diagram illustrating the original dataset according to this embodiment. [Figure 7] This diagram illustrates the process from generating synthetic data to evaluating the accuracy of machine learning according to this embodiment. [Figure 8] A figure illustrating an example of automated machine learning according to this embodiment. [Figure 9] A diagram illustrating the detailed procedure of automated machine learning according to this embodiment. [Figure 10] This figure shows an example of the learning accuracy evaluated by the learning accuracy evaluation unit 103 according to this embodiment. [Figure 11] A flowchart illustrating the procedure for evaluating the level of confidentiality by the confidentiality evaluation unit 104 according to this embodiment. [Figure 12] This figure shows an example of how the evaluation results of a composite dataset are displayed by the evaluation result presentation unit 105 according to this embodiment. [Figure 13] This figure shows an example of how the evaluation results of a composite dataset are displayed when evaluated using four evaluation axes by the evaluation result presentation unit 105 according to this embodiment. [Figure 14] This figure shows an example of how the evaluation results of a composite dataset are displayed when evaluated using four evaluation axes by the evaluation result presentation unit 105 according to this embodiment. [Modes for carrying out the invention]

[0012] Next, embodiments for carrying out the present invention will be described in detail with reference to the drawings. Figure 1 is a diagram illustrating the configuration of an information processing system 1 including an information processing device 10 according to an embodiment of the present invention. As shown in Figure 1, the information processing system 1 comprises an information processing device 10 and a user terminal 20. The information processing device 10 is connected to the user terminal 20 via a communication network N such as the Internet.

[0013] The information processing system 1 has a function of generating synthetic data from original data including personal information and confidential information. Here, the synthetic data does not include the same upper part as the original data, but is data in which the characteristics of data such as statistical information are maintained. Further, the information processing system 1 has a function of evaluating the confidentiality and usefulness of the generated synthetic data. Here, the usefulness means the usefulness as machine learning data, that is, the learning accuracy of the optimized model by performing machine learning using the data. The information processing device 10 may be a general-purpose computer, may be composed of one computer, or may be composed of a plurality of computers distributed on the communication network N. The information processing device 10 may be installed in an enterprise or the like that provides the information processing system 1, or may be constructed on the cloud. Note that the functions of the information processing device 10 in the present embodiment may be implemented in the user terminal 20.

[0014] FIG. 2 is a block diagram showing the configuration of the information processing device 10. As shown in FIG. 2, the information processing device 10 includes a processor 11, a main memory 12, an input / output interface 13, a communication interface 14, and a storage device 15. The storage device 15 is a computer-readable recording medium such as a semiconductor memory (for example, a volatile memory or a non-volatile memory) or a disk medium (for example, a magnetic recording medium or a magneto-optical recording medium). Programs to be executed by the processor 11 and various data are stored in the storage device 15. The programs are read from the storage device 15 into the main memory 12 and are interpreted and executed by the processor 11, whereby various functions are executed.

[0015] The user terminal 20 is a terminal used by the user to utilize the information processing system 1. The user terminal 20 can use any terminal device capable of exchanging data with the information processing apparatus 10 via the communication network N, such as a personal computer (PC), a notebook PC, a tablet terminal, a smartphone, etc. FIG. 3 is a block diagram showing the configuration of the user terminal 20. As shown in FIG. 3, the user terminal 20 includes a processor 21, an input device 22 such as a keyboard, a mouse, various operation buttons, and a touch panel, a display device 23 such as a liquid crystal display, a communication interface 24 for connecting to the communication network N, and a storage device 25 such as a disk drive or a semiconductor memory (ROM, RAM, etc.). Various programs executed by the processor 21 and various data may be stored in the storage device 25. A dedicated application for connecting to the information processing apparatus 10 and utilizing the services of the information processing system 1 may be installed in the user terminal 20.

[0016] FIG. 4 is a block diagram showing functional modules of a program executed by the processor 11 of the information processing apparatus 10. As shown in FIG. 4, the functional modules executed by the processor 11 of the information processing apparatus 10 include a data generation unit 101, a machine learning execution unit 102, a learning accuracy evaluation unit 103, a secrecy evaluation unit 104, and an evaluation result presentation unit 105.

[0017] Next, the generation and evaluation process of synthetic data by the information processing system 1 will be described using the flowchart of FIG. 5. First, when the user inputs an original data set including personal information via the user terminal 20, the data generation unit 101 generates a plurality of synthetic data sets based on the original data set (step ST101).

[0018] Figure 6 illustrates an example of an original dataset. The example in Figure 6 shows data for cases related to damage insurance claims in the event of an accident. Each data entry includes information such as the claimed amount, the duration from the accident to the claim, the age of the parties involved, and the prefecture where the accident occurred. For each case, it is indicated as "Y" (Yes) or "N" (No) whether it was a fraudulent case. Users can input the original dataset by specifying the storage location of the original dataset on the user terminal 20 or by uploading the original dataset. The data generation unit 101 processes the individual data entries in the input original dataset so that they do not contain any information identical to the original data. For example, it may change the age by 2-3 years or change the prefecture to a neighboring prefecture.

[0019] The data generation unit 101 may, for example, use a machine learning model (data generation model) to learn the statistical features of the original dataset and generate a synthetic dataset that retains the learned features. The machine learning model may, for example, employ a model that generates synthetic data using a method that applies differential privacy.

[0020] Next, the machine learning execution unit 102 automatically performs machine learning on the AI ​​model using the original dataset and the generated synthetic dataset (step ST102). Furthermore, the learning accuracy evaluation unit 103 evaluates the accuracy of the learning model optimized by machine learning (step ST103).

[0021] Figure 7 will be used to explain in detail the process from synthetic data generation to evaluation of the accuracy of the machine learning model. As shown in Figure 7, the original dataset OD and the synthetic datasets SD1, SD2, ... generated by different methods in the data generation unit 101 are used as training datasets, and the machine learning execution unit 102 performs automated machine learning. For example, the original dataset exemplified in Figure 6 and the generated synthetic datasets are input as training datasets with correct answers of whether or not an insurance claim is fraudulent, and the model is trained to determine whether or not an insurance claim is fraudulent. When machine learning is completed, the optimized model using each dataset is output along with information on the evaluation of the learning accuracy by the learning accuracy evaluation unit 103. The user can check the evaluation of the learning accuracy for each dataset on the user terminal 20.

[0022] Furthermore, while the example in Figure 7 uses one model for training, as shown in Figure 8, it is also possible to train multiple different AI models using each dataset. That is, machine learning is performed on multiple different models using one training dataset (original dataset or synthetic dataset) as input data, and the optimized model and training accuracy are output for each model. The training accuracy evaluation unit 103 may automatically select the model with the highest training accuracy from among them and output the selected model and the evaluation information of the training accuracy as the evaluation result of training using the dataset. The multiple models to be trained may be selected by the user from those pre-implemented in the information processing device 10.

[0023] Figure 9 illustrates the detailed procedure for automated machine learning. As shown in Figure 9, a single dataset is divided into training data TD, validation data VD, and test data TD, and training is first performed using the training data TD. At this time, models with multiple patterns of hyperparameters (parameters set before training) are prepared and each is trained to optimize the parameters. After training, the training accuracy of the models with optimized parameters is evaluated using the validation data VD. The training accuracy evaluation unit 103 selects the model with the highest evaluation hyperparameters, evaluates the accuracy of the selected model using the test data TD, and outputs the evaluation result of training using the dataset.

[0024] The learning accuracy evaluated by the learning accuracy evaluation unit 103 may be one or a combination of the correct response rate, detection rate (recall), precision (accuracy), and F-score, which are defined using the confusion matrix shown in Figure 10.

[0025] Next, the confidentiality evaluation unit 104 evaluates the confidentiality of each composite dataset (step ST104). Figure 11 is a flowchart illustrating the procedure for confidentiality evaluation by the confidentiality evaluation unit 104. The method shown in Figure 11 calculates the proportion of data in the composite dataset that matches data in the original dataset, and uses this as the confidentiality evaluation value. First, the confidentiality evaluation unit 104 retrieves one data from the original dataset (step ST201). Next, it retrieves one data from the composite dataset (step ST202). The confidentiality evaluation unit 104 determines whether the two retrieved data match (step ST203). If they match (Yes), it increments counter k by 1 (step ST204). The confidentiality evaluation unit 104 performs the same process for all combinations of original data and all composite data (steps ST205, ST206). Finally, it calculates the sum of counter k against the total number of composite data and outputs it as the confidentiality evaluation value (step ST207). However, the method for evaluating the degree of confidentiality is not limited to this method.

[0026] Next, the evaluation result presentation unit 105 presents the evaluation results of the confidentiality level and the machine learning accuracy level of each synthetic dataset on the user terminal 20 so that they can be compared (step ST105). Figure 12 is a diagram showing an example of how the evaluation results of synthetic datasets are displayed by the evaluation result presentation unit 105. As shown in Figure 12, points representing each synthetic dataset and the original dataset are plotted on a graph where the vertical axis represents the confidentiality level evaluation result and the horizontal axis represents the learning accuracy evaluation result. The higher the point in the graph, the higher the evaluation of both the data set and the data set. The similarity to the original dataset can also be visually grasped. Alternatively, when the user points the pointer to a point on the screen, the name of the data generation method, learning accuracy (utility), and confidentiality level (security) of that dataset may be displayed, as shown in Figure 12.

[0027] Users can examine the characteristics of each composite data set using the graph illustrated in Figure 12 and select the composite data set that is suitable for their application.

[0028] In the above embodiment, the confidentiality of the synthesized data and its usefulness as machine learning data are evaluated, but the evaluation criteria are not limited to these two. For example, data fidelity (the degree to which the statistical characteristics of the original data are reproduced) and fairness (whether the information content of specific attributes is increased or whether the bias of the original data is corrected) may also be added as evaluation criteria. Examples of methods for evaluating fidelity and fairness are as follows.

[0029] (Method for evaluating fidelity) 1. Comparison of statistical characteristics This compares the mean, variance, and distribution shape of the original and composite data. (Example) If the average age in the original data is 35, check if the average age in the composite data is also around 35. 2. Maintaining correlation Verify whether the relationships between data items are maintained. (Example) If there is a positive correlation between annual income and age in the original data, check whether a similar relationship exists in the synthesized data.

[0030] (Method for evaluating fairness) 1. Statistical Parity This involves verifying whether the data characteristics are fairly distributed across different groups (attributes). Specifically, it involves measuring the differences in descriptive statistics (mean, variance, etc.) for each group. (Example) Verify whether the hiring rates for men and women are extremely skewed in the composite data. 2.Disparate Impact This evaluates whether the performance of machine learning models significantly degrades in specific groups. The primary method involves comparing the difference in true positive rates (the rate at which a case is correctly identified as "positive") between groups. (Example) Check whether the AI ​​used for loan assessments shows an unfairly low approval rate for a particular racial group.

[0031] The evaluation can be performed using the following procedure. (1) Define the group to be evaluated (gender, age group, region, etc.). (2) Calculate the data distribution and statistics for each group. (3) Quantify the differences between the groups. (4) Set an acceptable range of differences and determine whether the criteria are met.

[0032] Furthermore, the evaluation result presentation unit 105 may allow users to select any two evaluation axes from the four independent evaluation axes (usefulness, confidentiality, fidelity, and fairness) to compare the evaluation results. Figures 13 and 14 show examples of display of evaluation results when a composite dataset is evaluated using the four evaluation axes described above. As shown in Figures 13 and 14, by operating the pull-down lists 51 and 52, any two of the four evaluation axes can be selected and set as the X and Y axes. For example, in the example in Figure 13, confidentiality is set as the X axis and usefulness as the Y axis. In graph 53, points representing datasets DS1 to DS5 are plotted based on their respective evaluation values ​​for confidentiality and usefulness. When a point (for example, DS5) is selected on the graph, the values ​​of each evaluation index for the selected dataset may be displayed in the area 54 on the right. Also, in the example in Figure 14, fidelity is set as the X axis and fairness as the Y axis. In this way, any combination of the four evaluation axes (six possibilities) can be selected, allowing users to check the evaluation results with the optimal combination according to their purpose. Furthermore, as shown in Figures 13 and 14, a recommended range may be displayed for each combination of indicators.

[0033] While evaluation metrics other than the four listed above may be used, it is desirable that the evaluation axes be independent of each other in order to minimize the impact of changes in one metric on the values ​​of other metrics.

[0034] (Method for generating composite data) The data generation unit 101 synthesizes data with different values ​​but the same data structure as the original dataset to generate a synthetic dataset. Synthetic data can be obtained by generating data that retains the features of the original data learned by the data generation model. Examples of data generation models include, but are not limited to, statistical value-based models that generate sets of data that can reconstruct probability distributions and cross-tabulation tables, as well as Basian Networks, Copula, GANs, and Diffusion Models.

[0035] On the other hand, there is a risk that the original data may be reconstructed from the synthetic data, or that information about the original data may be inferred. To address these risks, it is desirable to generate synthetic data using methods that satisfy the safety indicator known as differential privacy. Specifically, for example, by adding noise to the original data when the data generation model is being trained, it is possible to make it difficult to reconstruct the original data.

[0036] As described above, according to this embodiment, by simply inputting the original dataset, the information processing device 10 generates a synthetic dataset with personal information concealed, performs machine learning on the AI ​​model to evaluate the learning accuracy, and displays it together with the evaluation of the degree of concealment. Therefore, even without specialized knowledge of data processing or machine learning, users can select highly accurate synthetic data suitable for their application.

[0037] Furthermore, by performing machine learning on multiple models and selecting and evaluating the model with the highest learning accuracy for each synthetic dataset, it is possible to find highly practical combinations of synthetic datasets and models.

[0038] The embodiments described above are provided to facilitate understanding of the present invention and are not intended to limit its interpretation. The flowcharts, sequences, and specific examples of elements of the embodiments described in the embodiments are not limited to those exemplified and can be modified as appropriate. Furthermore, it is possible to partially substitute or combine the configurations shown in different embodiments. [Explanation of Symbols]

[0039] 1... Information processing system, 10... Information processing device, 11... Processor, 12... Main memory, 13... Input / output interface, 14... Communication interface, 15... Storage device, 20... User terminal, 21... Processor, 22... Input device, 23... Display device, 24... Communication interface, 25... Storage device, 101... Data generation unit, 102... Machine learning execution unit, 103... Learning accuracy evaluation unit, 104... Confidentiality evaluation unit, 105... Evaluation result presentation unit

Claims

1. A data generation unit that generates multiple composite datasets based on the original dataset, A machine learning execution unit that automatically performs machine learning on an AI model using each of the synthesized datasets and the original dataset as training data, A learning accuracy evaluation unit that evaluates and outputs the accuracy of the aforementioned machine learning, For each composite dataset, a confidentiality evaluation unit evaluates and outputs the degree of confidentiality based on the proportion of composite data in the composite dataset that matches the data in the original dataset, An information processing device comprising: an evaluation result presentation unit that displays the evaluation result of the confidentiality level of each synthetic dataset and the evaluation result of the machine learning accuracy on the same screen.

2. The aforementioned machine learning execution unit, For each set of training data, machine learning is performed using multiple different AI models. The aforementioned learning accuracy evaluation unit, The information processing apparatus according to claim 1, wherein for each set of training data, the learning accuracy of the AI ​​model with the highest machine learning accuracy is adopted as the evaluation result.

3. The aforementioned machine learning execution unit, For each training dataset, machine learning is performed using an AI model with multiple different hyperparameters. The aforementioned learning accuracy evaluation unit, The information processing apparatus according to claim 1, wherein for each training data set, the hyperparameter that yields the highest evaluation of machine learning accuracy is selected, and the learning accuracy of the AI ​​model with the hyperparameter set is adopted as the evaluation result.

4. The aforementioned learning accuracy evaluation unit, The information processing apparatus according to claim 1, which performs an evaluation of at least one of the following: accuracy rate, detection rate, precision, and F-score.

5. The evaluation result presentation unit is, The information processing apparatus according to claim 1, wherein points representing each composite dataset are displayed on a graph having a first axis showing the value of the evaluation result of the degree of confidentiality and a second axis showing the value of the evaluation result of the accuracy of the machine learning.

6. The data generation unit, The information processing apparatus according to claim 1, which learns the statistical features of the original dataset using a data generation model and generates a synthetic dataset such that the learned features are preserved.

7. The data generation model generates synthetic data using a method that applies differential privacy, as described in claim 6.

8. The evaluation result presentation unit is, Furthermore, the information processing apparatus according to claim 1, which presents the results of the fidelity evaluation and / or fairness evaluation of each composite dataset so that they can be compared.

9. The evaluation result presentation unit is, The information processing apparatus according to claim 1, which presents the evaluation results of each composite dataset in a manner that allows comparison of any two selected indicators from three or more independent evaluation indicators as evaluation axes.

10. The process involves a computer generating multiple synthetic datasets based on an original dataset, The process involves a computer automatically performing machine learning on an AI model using each of the synthesized datasets and the original dataset as training data, The process involves a computer evaluating the accuracy of the machine learning and outputting the result, The process involves a computer evaluating the degree of confidentiality for each synthetic dataset based on the proportion of synthetic data in that dataset that matches the data in the original dataset, and outputting the result. An information processing method comprising the step of a computer displaying the results of the confidentiality evaluation of each synthetic dataset and the results of the machine learning accuracy evaluation on the same screen.

11. Computers, A data generation unit that generates multiple composite datasets based on the original dataset, A machine learning execution unit that automatically performs machine learning on an AI model using each of the synthesized datasets and the original dataset as training data, A learning accuracy evaluation unit that evaluates and outputs the accuracy of the aforementioned machine learning, For each composite dataset, a confidentiality evaluation unit evaluates and outputs the degree of confidentiality based on the proportion of composite data in the composite dataset that matches the data in the original dataset, A program that functions as an evaluation result presentation unit, displaying the evaluation results of the confidentiality level for each synthetic dataset and the evaluation results of the machine learning accuracy on the same screen.

Citation Information

Patent Citations

  • Information processor, method for processing information, computer program, and learning system

    JP2022064115A

  • Teacher data management device, teacher data management method, and program for teacher data management

    JP2024157440A

  • Learning device, learning system, learning method, and recording medium

    WO2025104827A1

  • Process controller

    JP1985083101A