Information processing device and program

WO2026203203A1PCT designated stage Publication Date: 2026-10-01NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012494
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2026-10-01

Smart Images

  • Figure JP2025012494_01102026_PF_FP_ABST
    Figure JP2025012494_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device comprising: an input unit that inputs the probability distribution of original data, the probability distribution of synthetic data, a statistical distance function, and a statistical distance; and a calculation unit that changes the probability distribution of the synthetic data until the distance between the probability distribution of the synthetic data and the probability distribution of the original data in the statistical distance function reaches the statistical distance, to thereby obtain the changed probability distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and program

[0001] This invention relates to a technology for evaluating the quality of tabular data.

[0002] When utilizing personal data stored in a database as tabular data, it is necessary to consider the privacy of the individuals contained in that data. As a method of protecting individual privacy, techniques for generating synthetic data are attracting attention.

[0003] Synthetic data refers to artificial data generated to resemble the data being analyzed, using statistical values, machine learning models, etc. In the case of privacy-protected synthetic data, privacy protection processing is performed during the data generation process.

[0004] Dwork Cynthia. Differential privacy. Automata, languages ​​and programming, pp. 1-12, 2006.

[0005] To guarantee the quality of privacy-protected synthetic data, it is necessary to maintain a good balance between data security and usefulness. The evaluations of security and usefulness are as follows:

[0006] Security assessment involves evaluating how well personal information contained in the data is protected. Standard security assessment methods include differential privacy (Non-Patent Document 1).

[0007] The evaluation of usefulness involves assessing whether the synthesized data and the original data have similar properties, and various metrics (evaluation criteria) are used. However, because there are no standard evaluation criteria, it is difficult to compare the merits of different synthesis methods.

[0008] This invention has been made in view of the above points, and aims to provide a technology that enables the appropriate setting of evaluation criteria used when evaluating the quality of synthetic data generated for privacy protection.

[0009] According to the disclosed technology, an information processing device is provided, comprising: an input unit for inputting the probability distribution of the original data, the probability distribution of the composite data, a statistical distance function, and a statistical distance; and a calculation unit for determining the probability distribution when the probability distribution of the composite data is changed until the distance between the statistical distance function and the probability distribution of the original data becomes the statistical distance.

[0010] The disclosure technology makes it possible to appropriately set evaluation criteria used when assessing the quality of synthetic data generated for privacy protection.

[0011] This is a diagram illustrating the configuration of the information processing device 100. This diagram is for explaining the outlines of Examples 1 and 2. This is a diagram showing algorithm 1 corresponding to the processing procedure of evaluator 1. This is a diagram showing an example of output from evaluator 1. This is a diagram showing an example of output from evaluator 1. This is a diagram showing an example of output from evaluator 1. This is a diagram showing algorithm 2 corresponding to the processing procedure of evaluator 2. This is a diagram showing an example of output from evaluator 2. This is a diagram showing an example of output from evaluator 2. This is a diagram showing an example of the hardware configuration of the device.

[0012] Hereinafter, embodiments of the present invention (this embodiment) will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the embodiments described below.

[0013] (Summary of Embodiments) When evaluating the quality of synthetic data, in order to say that "synthetic data is similar to the original data," it is necessary to clarify what statistical distance function should be used and what the value of that distance should be. The technology according to this embodiment can provide a method for doing so. By using this method, it becomes possible to understand which changes in statistical quantities the statistical distance function is susceptible to. Furthermore, it can provide an intuitive understanding of the statistical distance for the similarity between two tabular datasets.

[0014] Specifically, the information processing device 100, described later, outputs the change in distribution when the statistical distance between the synthesized data and the original data is reduced. The influence of the statistical quantity can be inferred from the process of change.

[0015] Hereinafter, the configuration and operation of the information processing apparatus 100 will be described in detail.

[0016] (Example of Apparatus Configuration) FIG. 1 shows an example of a functional configuration of an information processing apparatus 100 that operates as an evaluator 1 and an evaluator 2, which will be described later, in the present embodiment. As shown in FIG. 1, the information processing apparatus 100 includes an input unit 110, an arithmetic unit 120, an output unit 130, and a data storage unit 140. Note that the information processing apparatus 100 may also be referred to as an information processing system 100. The information processing apparatus 100 may be an apparatus (which may also be referred to as a system) configured of a plurality of computing devices (computers).

[0017] The information processing apparatus 100 may have a function of performing secure computation. Secure computation is a computation method that performs computation on data while the data remains encrypted.

[0018] In the information processing apparatus 100, data is input from the input unit 110, and the arithmetic unit 120 performs computation for obtaining a probability distribution to be described later. The output unit 130 outputs the computation result obtained by the arithmetic unit 120. Note that the computation result obtained by the arithmetic unit 120 may be stored in the data storage unit 140 without being output from the output unit 130.

[0019] For example, predetermined parameters are stored in the data storage unit 140. The arithmetic unit 120 reads the parameters from the data storage unit 140 and uses them for computation.

[0020] Note that the information processing apparatus 100 may have the functions of both the evaluator 1 and the evaluator 2, or the information processing apparatus 100 serving as the evaluator 1 and the information processing apparatus 100 serving as the evaluator 2 may be provided separately. Furthermore, an information processing apparatus that functions as an evaluator may also be referred to as an evaluation apparatus.

[0021] Hereinafter, the operation of the evaluator 1 will be described as a first embodiment, and the operation of the evaluator 2 will be described as a second embodiment.

[0022] (Summary of Embodiments 1 and 2) First, the summary of Embodiments 1 and 2 will be described. In each embodiment, the original data and synthetic data are, for example, table-format data indicating attributes of a plurality of people.

[0023] The probability distribution in each embodiment corresponds to distribution data obtained by representing tabular format data as a histogram and normalizing the data such that the sum of the values equals 1. For example, suppose in data regarding 10 people, 3 people correspond to attribute value A, 3 people correspond to attribute value B, and 4 people correspond to attribute value C, then the probability distribution with attribute values as random variables is "attribute value A = 0.3, attribute value B = 0.3, attribute value C = 0.4".

[0024] Figure 2 is a diagram illustrating the processing overview of Embodiments 1 and 2. As shown in Figure 2, in Embodiment 1 that uses Evaluator 1, "original probability distribution P, probability distribution Q of synthetic data, distance function d t , and statistical distance θ" are input to Evaluator 1. The synthetic data is data generated based on original data corresponding to the original probability distribution P.

[0025] Evaluator 1 changes distribution Q until the statistical distance between distribution P and distribution Q reaches θ, and outputs a probability distribution Q having a distance θ with respect to P for the distance function d t . By means of Evaluator 1, a distance function d that is a candidate for an evaluation criterion used, for example, to evaluate the quality of synthetic data t and the distance θ can be determined.

[0026] In Embodiment 2 that uses Evaluator 2, "original probability distribution P, distance function d t , statistical distance θ, and statistic stat" are input to Evaluator 2. The statistic stat is information indicating what is to be maximized or minimized. In Embodiment 2, the inputs are the distance function d determined as a candidate using Evaluator 1 t and the distance θ. However, the inputs are not limited to those determined as candidates using Evaluator 1.

[0027] Evaluator 2 searches, by changing distribution P, for a distribution Q that maximizes or minimizes the statistic stat within a range where the distance from distribution P is θ, and outputs the distribution Q.

[0028] Hereinafter, Embodiment 1 (Evaluator 1) and Embodiment 2 (Evaluator 2) will be described in more detail.

[0029] (Example 1: Operation of Evaluator 1) First, Example 1 will be described. In Example 1, evaluator 1 calculates distribution Q using statistical distance function d t to obtain Q when Q is changed until the distance between Q and distribution P becomes θ.

[0030] The above processing corresponds to obtaining Q in the following formula (1) * .

[0031] Formula (1) means obtaining Q that makes the absolute value of "d t (P, Q)−θ" equal to 0, which is the minimum value thereof. Note that Δ K-1 is "Δ K-1 : = {x∈[0,1] K |Σ i=1 K x i = 1}", which is regarded as a set of probability distributions. "[0,1] K " represents a set of K "numbers from 0 to 1 inclusive".

[0032] Evaluator 1 approximately obtains the distribution Q, which is the solution of formula (1), using a gradient descent method. FIG. 3 is a diagram showing Algorithm 1 corresponding to the processing procedure of Evaluator 1. The operation of Evaluator 1 will be described with reference to FIG. 3. In the information processing apparatus 100, the following input operation is performed by the input unit 110, the output operation is performed by the output unit 130, and the calculation for obtaining the target Q is performed by the calculation unit 120.

[0033] To Evaluator 1, an original distribution P, a target distribution Q, a target distance function d t , a statistical distance θ, and a correction micro width γ are input. Evaluator 1 performs the following calculation to output the distribution Q that is the solution of formula (1).

[0034] Evaluator 1 executes steps 2 to 5 while "d t (P, Q) > θ" holds true.

[0035] In step 2, Evaluator 1 uses "∇ t d(P, Q)", which is the derivative of "d Q d t (P, Q)−θ" with respect to Q, to correct Q so that "d t (P, Q)−θ" approaches 0. In the first stage of correction, dt Since we assume that (P, Q) > θ, d t Q is modified in a direction that decreases (P, Q).

[0036] In steps 3-4, evaluator 1 rounds up each component of Q to 0 if it is negative. In step 5, evaluator 1 normalizes so that the sum of the components of Q is 1. evaluator 1 then performs the following operation: t The process terminates when (P, Q) = θ, and the obtained Q is output.

[0037] In this embodiment, the distance function d t The value of is normalized to [0,1], making it easier to compare different distance functions. t Let the distance function d be arbitrary. t It is possible to use it.

[0038] The following describes specific input and output examples.

[0039] (Example 1: Example 1) The input in Example 1 is as follows.

[0040] • Original probability distribution P: exponential distribution • Probability distribution Q of composite data: uniform distribution • Statistical distance function d t :TVD distance, L ∞ Distance / Statistical distance θ: 0.1, 0.2, 0.3, 0.5 Note that TVD (Total Variation Distance) is L 1 This is half the distance. Figure 4 shows the output from evaluator 1 for each of the above inputs. Figure 4 shows the shape of the distribution at each distance as the probability distribution Q approaches the distribution P for each statistical distance function.

[0041] (Example 1: Example 2) The input in Example 2 is as follows.

[0042] • Original probability distribution P: relationship attribute of the Adult dataset • Probability distribution Q of the synthesized data: synthesized data generated by the DP-CTGAN method • Statistical distance function d t :TVD distance, L 2 distance, L ∞Distance, KL divergence, statistical distance θ: 0.1, 0.2, 0.3, 0.5. Note that DP-CTGAN is a method that applies DP-GAN, which guarantees differential privacy in GANs, to CTGAN by using DP-SGD during the training of the discriminative model. The outputs from evaluator 1 for each of the above inputs are shown in Figures 5 and 6. Figures 5 and 6 show the shape of the distribution at each distance when the probability distribution Q approaches the distribution P for each statistical distance function.

[0043] As shown in Examples 1 and 2 above, by using the evaluator 1, it becomes possible to understand the differences in statistical distance functions and the differences in the shape of the distributions resulting from differences in statistical distances. This makes it possible to determine candidate distances that should be used as evaluation criteria for the statistical distance function.

[0044] For example, in the examples shown in Figures 5 and 6, the KL divergence, TVD distance, and L have similar distribution shapes. 2 Distances θ = 0.1 and 0.2 can be determined as possible candidates.

[0045] (Example 2: Operation of Evaluator 2) Next, Example 2 will be described. In Example 2, the evaluator 2 calculates a distribution Q, which is a distribution obtained by changing distribution P such that the statistic stat in the distribution is minimized (or maximized) within a range of distance θ from distribution P.

[0046] The above process is Q in equation (2) below. * This is equivalent to finding [something].

[0047] In equation (2), stat(Q) can be, for example, the maximum value, minimum value, or variance of the distribution. "argmin stat(Q)" can be, for example, "maximizing the maximum value at Q," "minimizing the minimum value at Q," "maximizing the variance of Q," or "minimizing the variance of Q." Regarding "argmin," if you want to maximize a certain function, you can apply "argmin" to stat(Q) which is the function with a negative sign.

[0048] The evaluator 2 approximates the distribution Q, which is the solution to equation (2), using the gradient descent method. Figure 7 is a diagram showing algorithm 2, which corresponds to the processing procedure of the evaluator 2. The operation of the evaluator 2 will be explained with reference to Figure 7. In the information processing device 100, the following input operations are performed by the input unit 110, the output operations are performed by the output unit 130, and the calculation to find the target Q is performed by the calculation unit 120.

[0049] Evaluator 2 uses the original distribution P and the target distance function d. t The statistical distance θ, the target statistic stat, and the small width γ for correction are input. The evaluator 2 performs the following calculation and outputs the distribution Q which is the solution to equation (2).

[0050] In step 1, evaluator 2 sets Q to P.

[0051] Evaluator 2 is "d t Perform steps 3-6 while (P, Q) ≤ θ.

[0052] In step 3, evaluator 2 is the derivative of "stat(Q)" with respect to Q, which is "∇ Q Using "stat(Q)", we modify Q so that "stat(Q)" is minimized. Initially, d t Since (P, Q) = 0, d t Q is modified in a direction that increases (P, Q).

[0053] In steps 4-5, evaluator 2 rounds up each component of Q to zero if it is negative. In step 6, evaluator 2 normalizes Q so that the sum of its components is 1.

[0054] Evaluator 2 is "d t The process terminates when (P, Q) = θ, and the calculated Q is output.

[0055] In the above process, candidate distance functions d are selected from the results of evaluator 1. t The distance θ is selected and used.

[0056] Furthermore, Q is difficult to search in high-dimensional spaces. *This will be efficiently explored using the gradient descent method. Any differentiable statistic can be used as the statistic stat.

[0057] The following describes specific input and output examples.

[0058] (Example 2: Example 1) The input in Example 1 is as follows.

[0059] • Original probability distribution P: exponential distribution • Statistical distance function d t TVD distance, statistical distance θ: 0.1, statistic stat: max. In other words, evaluator 2 finds Q by varying P so that the maximum value of the probability distribution P is maximized within the range of distance θ = 0.1. Figure 8 shows the output from evaluator 2 for the above input.

[0060] (Example 2: Example 2) The input in Example 2 is as follows.

[0061] • Original probability distribution P: relationship attribute of the Adult dataset • statistical distance function d t :TVD distance, L 2 Distance, KL divergence, statistical distance θ: 0.1, 0.3, statistic stat: max. In other words, for each of the three statistical distance functions, the evaluator 2 finds Q by varying P so as to maximize the maximum value of the probability distribution P within the range of distance θ = 0.1 and 0.3, respectively. The output from the evaluator 2 for the above input is shown in Figure 9.

[0062] As shown in Examples 1 and 2 above, using evaluator 2 makes it possible to compare the effects of statistics based on distance functions using actual datasets. In Example 2 above, the large change in the maximum value of the KL divergence indicates that the maximum value of the distribution does not have a significant impact on the KL divergence.

[0063] (Example Hardware Configuration) The information processing device 100 described in this embodiment can be realized, for example, by having a computer execute a program. This computer may be a physical computer or a virtual machine on the cloud.

[0064] In other words, the device can be realized by using hardware resources such as the CPU and memory built into a computer to execute a program corresponding to the processing performed by the device. The program can be recorded on a computer-readable recording medium (such as portable memory), saved, and distributed. It can also be provided via a network, such as the Internet or email.

[0065] Figure 10 shows an example of the hardware configuration of the computer described above. The computer in Figure 10 has a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., all of which are interconnected by bus B. The computer may also be equipped with a GPU.

[0066] The program that enables processing on the computer is provided on a recording medium 1001, such as a CD-ROM or memory card. When the recording medium 1001 containing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001; it may also be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files and data.

[0067] The memory device 1003 reads and stores a program from the auxiliary storage device 1002 when a program startup command is received. The CPU 1004 implements the functions related to the memory device 1003 according to the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) etc., based on a program. The input device 1007 consists of a keyboard and mouse, buttons, or a touch panel, etc., and is used to input various operation commands. The output device 1008 outputs the calculation results.

[0068] Furthermore, the functions of the elements disclosed herein may be implemented using circuits or processing circuitry that include general-purpose processors, special-purpose processors, integrated circuits, ASICs (Application Specific Integrated Circuits), FPGAs (Field Programmable Gate Arrays), conventional circuits, and / or combinations thereof that are programmed using one or more programs stored in one or more memories, or otherwise configured to perform the disclosed functions. A processor is considered processing circuitry or circuitry because it includes transistors and other circuits. A processor may be a programmed processor that executes programs stored in memory. In this disclosure, a circuit, unit, or means is hardware that performs the enumerated functions, or hardware programmed to perform the enumerated functions. Hardware may be any hardware disclosed herein that is programmed or configured to perform the enumerated functions.

[0069] The system includes memory for storing computer programs, which include computer instructions. These computer instructions provide logic and routines that enable hardware (e.g., processing circuitry or circuitry) to perform the methods disclosed herein. The computer programs can be implemented in commonly known forms, such as computer-readable storage media, computer program products, memory devices, recording media such as CD-ROMs and DVDs, and / or memory in FPGAs and ASICs.

[0070] (Regarding the effects of the embodiment) In the conventional technology, there are no clear criteria for evaluating synthesized data, making it difficult to judge the synthesis method. On the other hand, by using the technology according to this embodiment, it becomes possible to gain insight into the impact of evaluation criteria on statistical values ​​and the values ​​that should be used as criteria for evaluation functions.

[0071] For example, when a statistical distance function and distance are selected using evaluator 1, evaluator 2 outputs the most unfavorable distribution for a given statistic, allowing the user to reconsider the function and distance settings based on the results.

[0072] The following additional information is disclosed regarding the embodiments described above.

[0073] <Notes> (Note 1) An information processing device comprising: an input unit for inputting the probability distribution of the original data, the probability distribution of the composite data, a statistical distance function, and a statistical distance; and a calculation unit for determining the probability distribution when the probability distribution of the composite data is changed until the distance between it and the probability distribution of the original data in the statistical distance function becomes the statistical distance. (Note 2) An information processing device comprising: an input unit for inputting the probability distribution of the original data, a statistical distance function, a statistical distance, and a statistical quantity; and a calculation unit for determining the probability distribution when the probability distribution of the original data is changed such that the statistical quantity becomes the maximum or minimum within a range where the distance between it and the probability distribution of the original data at the time of input in the statistical distance function becomes the statistical distance. (Note 3) The information processing device according to Note 1 or 2, wherein the calculation unit uses the gradient descent method to determine the target probability distribution. (Note 4) A non-temporary storage medium storing a program for causing a computer to function as each part of the information processing device according to Note 1 or 2.

[0074] Although this embodiment has been described above, the present invention is not limited to this specific embodiment, and various modifications and changes are possible within the scope of the gist of the invention as described in the claims.

[0075] 100 Information processing device 110 Input unit 120 Calculation unit 130 Output unit 140 Data storage unit 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device

Claims

1. An information processing device comprising: an input unit for inputting the probability distribution of the original data, the probability distribution of the composite data, a statistical distance function, and a statistical distance; and a calculation unit for determining the probability distribution when the probability distribution of the composite data is changed until the distance between the statistical distance function and the probability distribution of the original data becomes the statistical distance.

2. An information processing device comprising: an input unit for inputting the probability distribution of the original data, a statistical distance function, a statistical distance, and a statistical quantity; and a calculation unit for determining the probability distribution when the probability distribution of the original data is changed such that the statistical quantity is maximized or minimized within a range where the distance in the statistical distance function between the original data's probability distribution at the time of input and the probability distribution of the original data is the statistical distance.

3. The information processing apparatus according to claim 1 or 2, wherein the calculation unit uses the gradient descent method to determine the target probability distribution.

4. A program for causing a computer to function as a component of the information processing apparatus described in claim 1 or 2.