Information processor and information processing method

The information processing apparatus addresses the risk of user information leakage by integrating user data from multiple operators using noise addition and irreversible conversion queries, ensuring that the data satisfies ε-local differential privacy and preventing unauthorized access.

JP2025088689APending Publication Date: 2025-06-11KDDI CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024082102
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-06-11

AI Technical Summary

Technical Problem

There is a risk of user information leakage during the collection process, as conventional methods anonymize user information after it is collected from multiple operators.

Method used

An information processing apparatus and method that integrate user data from multiple operators by generating noise addition and irreversible conversion queries, ensuring that noise is added to the data with a predetermined probability and that data identification information is irreversibly converted, thereby preventing unauthorized access.

Benefits of technology

The solution effectively prevents the leakage of non-anonymized user information during collection by ensuring that integrated user data satisfies ε-local differential privacy, thus enhancing data privacy and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025088689000001_ABST
    Figure 2025088689000001_ABST
Patent Text Reader

Abstract

To prevent leakage of user information, for which anonymization processing is not performed, in the process of collecting the user information.SOLUTION: An information processor 1 comprises: a generation section 132 that generates a first noise giving query for giving a first data group noise and a second noise giving query for giving a second data group noise so that noise is given to a plurality of pieces of data included in an integrated data group with a predetermined probability when the integrated data group is generated by integrating the first data group and the second data group, a first irreversible conversion query which is a query for irreversibly converting a plurality of pieces of data identification information included in the first data group, and a second irreversible conversion query which is a query for irreversibly converting a plurality of pieces of data identification information included in the second data group; and a transmission section 133 that transmits the first noise giving query and the first irreversible conversion query to a first device 2 and transmits the second noise giving query and the second irreversible conversion query to a second device 3.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus and an information processing method.

Background Art

[0002] Conventionally, user information, which is information about users, has been collected from multiple operators and data analysis has been performed. In this case, in order to protect the privacy of users, at least a part of the user information collected from multiple operators is anonymized. For example, Patent Document 1 discloses a system that performs irreversible conversion or the like on data that is a combination key for combining multiple user information, combines the personal information of users corresponding to each of the multiple operators using the converted combination key, and additionally performs an anonymization process on the combined data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, since user information is collected from each of multiple operators and then anonymized, there is a risk that user information that has not been anonymized may leak during the process of collecting user information.

[0005] Therefore, the present invention has been made in view of these points, and an object thereof is to prevent user information that has not been anonymized from leaking during the process of collecting user information.

Means for Solving the Problems

[0006] The information processing apparatus according to the first aspect of the present invention includes a first data group including a plurality of first records associating data identification information for identifying data with first data, and a second data group including a plurality of second records associating the data identification information with second data. For each of the plurality of first data included in the first data group, a first noise addition query, which is a query for adding noise so that noise is added to a plurality of data included in the integrated data group with a predetermined probability when the first data group and the second data group are integrated into an integrated data group, and a first irreversible conversion query, which is a query for irreversibly converting the plurality of data identification information included in the first data group by a predetermined method, are generated by a first generation unit. For each of the plurality of second data included in the second data group, a second noise addition query, which is a query for adding noise so that noise is added to a plurality of data included in the integrated data group with the predetermined probability, and a second irreversible conversion query, which is a query for irreversibly converting the plurality of data identification information included in the second data group by the predetermined method, are generated by a second generation unit. A transmission unit transmits the first noise addition query and the first irreversible conversion query generated by the first generation unit to a first device corresponding to the provider of the first data group, and transmits the second noise addition query and the second irreversible conversion query generated by the second generation unit to a second device corresponding to the provider of the second data group. A data group acquisition unit acquires a converted first data group including a plurality of first records associating the data identification information converted based on the first irreversible conversion query with the plurality of first data to which noise has been added based on the first noise addition query, and a converted second data group including a plurality of second records associating the data identification information converted based on the second irreversible conversion query with the plurality of second data to which noise has been added based on the second noise addition query. Based on the data identification information of each of the plurality of first records included in the converted first data group acquired by the data group acquisition unit and the data identification information of each of the plurality of second records included in the converted second data group acquired by the data group acquisition unit,An integration unit that generates the integrated data group obtained by integrating the first data group after the conversion and the second data group after the conversion.

[0007] The first record includes a plurality of first data corresponding to each of n 1 attributes, the second record includes a plurality of second data corresponding to each of n 2 attributes, and the plurality of data included in the integrated data group satisfy ε-local differential privacy with a parameter ε indicating the strength of privacy. The first generation unit is configured such that when noise is added to the first data of each of the n 1 attributes, the first data of each of the n 1 attributes with noise added satisfies ε 1 -local differential privacy (where the parameter ε 1 indicating the strength of privacy is ε 1 = ε / (n 1 + n 2 )) and generates the first noise addition query for adding noise. The second generation unit is configured such that when noise is added to the second data of each of the n 2 attributes, the second data of each of the n 2 attributes with noise added satisfies ε 2 -local differential privacy (where the parameter ε 2 indicating the strength of privacy is ε 2 = ε / (n 1 + n 2 )) and may generate the second noise addition query for adding noise.

[0008] There are k data groups (where k is an integer of 3 or more) associated with the data identification information, and for each of the multiple k-th data included in the k-th data group, when the k data groups are integrated into the integrated data group, a k-th noise addition query that adds noise to the multiple data included in the integrated data group with a predetermined probability, and a k-th irreversible conversion query that irreversibly converts the multiple data identification information included in the k-th data group by a predetermined method, and further includes a k-th generation unit that generates the k-th irreversible conversion query, and the k-th record included in the k-th data group includes multiple k-th data corresponding to each of n k attributes, and the k-th generation unit, when noise is added to the k-th data of each of the n k attributes, the k-th data of each of the n k attributes to which noise is added satisfies ε k -local differential privacy (where the parameter ε k indicating the strength of privacy is ε k =ε / (n 1 +n 2 +···+n k ))), the k-th noise addition query that adds noise may be generated so as to satisfy the condition.

[0009] The information processing device acquires first item information indicating items corresponding to each of the multiple attributes constituting the first record from the first device, and acquires second item information indicating items corresponding to each of the multiple attributes constituting the second record from the second device, specifies the n 1 which is the number of attributes corresponding to the first data based on the acquired first item information, specifies the n 2 which is the number of attributes corresponding to the second data based on the acquired second item information, and based on the specified n 1 and the n 2 , a first parameter ε 1 indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added, and a second parameter ε 2It may have a determination unit that determines

[0010] At least one of the first generation unit and the second generation unit reduces the number of values that the data corresponding to at least one attribute among the plurality of data corresponding to the plurality of attributes included in the data group can take, and after reducing the number of values that the data can take, it may generate the noise addition query that adds the noise to each of the plurality of data.

[0011] The first generation unit generates a first update query, which is a query for updating the converted first data group by swapping the first data included in the first record included in the converted first data group with the first data included in other first records at a first ratio. The second generation unit generates a second update query, which is a query for updating the converted second data group by swapping the second data included in the second record included in the converted second data group with the second data included in other second records at a second ratio. The transmission unit may transmit the first update query generated by the first generation unit to the first device and transmit the second update query generated by the second generation unit to the second device.

[0012] The information processing apparatus may have an update unit that updates the integrated data group by swapping the data included in the records included in the integrated data group with the data included in other records at a third ratio.

[0013] The first generation unit generates the first irreversible conversion query that adds random data to each of the plurality of data identification information included in the first data group and then performs irreversible conversion by the predetermined method. The second generation unit generates the second irreversible conversion query that adds the same random data as the random data added to the data identification information included in the first data group corresponding to the data identification information to each of the plurality of data identification information included in the second data group and then performs irreversible conversion by the predetermined method.

[0014] The integration unit may further process the integrated data group to generate statistical data, and perform correction to remove the noise using the predetermined probability used for adding the noise to the statistical data.

[0015] The first data group and the second data group include identification data that can be used to identify newly added records. When the first generation unit regenerates the first irreversible conversion query, it generates the first noise addition query for adding the noise to the first records newly added based on the identification data. When the second generation unit regenerates the second irreversible conversion query, it may generate the second noise addition query for adding the noise to the second records newly added based on the identification data.

[0016] The integration unit may integrate the converted first data group and the converted second data group based on the data identification information of each of the plurality of first records included in the converted first data group and the data identification information of each of the plurality of second records included in the converted second data group, exclude the data identification information, and generate the integrated data group.

[0017] The information processing method according to the second aspect of the present invention is a query for adding noise to each of a plurality of first data included in a first data group including a plurality of first records associating data identification information for identifying data with the first data, and a second data group including a plurality of second records associating the data identification information with second data, so that noise is added to a plurality of data included in the integrated data group when the first data group and the second data group are integrated into an integrated data group, which is a first noise addition query, and a query for irreversibly converting a plurality of the data identification information included in the first data group by a predetermined method, which is a first irreversible conversion query, and generating steps; for each of the plurality of second data included in the second data group, a second noise addition query that is a query for adding noise so that noise is added to a plurality of data included in the integrated data group with the predetermined probability, and a second irreversible conversion query that is a query for irreversibly converting a plurality of the data identification information included in the second data group by the predetermined method, and generating steps; transmitting the generated first noise addition query and the first irreversible conversion query to a first device corresponding to a provider of the first data group, and transmitting the generated second noise addition query and the second irreversible conversion query to a second device corresponding to a provider of the second data group; a converted first data group including a plurality of first records associating the data identification information converted based on the first irreversible conversion query with a plurality of first data to which noise is added based on the first noise addition query; and a converted second data group including a plurality of second records associating the data identification information converted based on the second irreversible conversion query with a plurality of second data to which noise is added based on the second noise addition query, and obtaining steps; and generating the integrated data group obtained by integrating the converted first data group and the converted second data group based on the data identification information of each of the plurality of first records included in the obtained converted first data group and the data identification information of each of the plurality of second records included in the obtained converted second data group.

Advantages of the Invention

[0018] According to the present invention, in the process of collecting user information, it is possible to prevent the outflow of user information that has not been anonymized.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0020] [Outline of Information Processing System S] FIG. 1 is a diagram for explaining the outline of the information processing system S. The information processing system S includes an information processing apparatus 1, a first apparatus 2 that manages a first data group, and a second apparatus 3 that manages a second data group, and generates an integrated data group by integrating the first data group and the second data group after anonymizing the user information included in the first data group and the second data group.

[0021] The information processing apparatus 1 is operated by, for example, an aggregation operator that provides a service for aggregating data and providing the aggregated data, and is communicably connected to external apparatuses such as the first apparatus 2 and the second apparatus 3 via a communication network (not shown) such as the Internet or a mobile phone line.

[0022] The first device 2 is operated by, for example, a first operator, and manages a first data group including a plurality of first records associating a data ID as data identification information for identifying data with first data. The second device 3 is operated by, for example, a second operator, and manages a second data group including a plurality of second records associating a common data ID with the data ID included in the first data group and second data.

[0023] The information processing apparatus 1 generates a first noise addition query for adding noise to each of the plurality of first data included in the first data group, and a first irreversible conversion query which is a query for irreversibly converting the plurality of data IDs included in the first data group by a predetermined method. The first noise addition query is, for example, a query for adding noise to each of the plurality of first data included in the first data group such that noise is added to the plurality of data included in the integrated data group when the first data group and the second data group are integrated into an integrated data group with a predetermined probability. The query is, for example, assumed to be an SQL (Structured Query Language) statement executable in a relational database management system.

[0024] The information processing apparatus 1 generates a second noise addition query for adding noise to each of the plurality of second data included in the second data group, and a second irreversible conversion query which is a query for irreversibly converting the plurality of data IDs included in the second data group by a predetermined method. Similar to the first noise addition query, the second noise addition query is a query for adding noise to each of the plurality of second data included in the second data group such that noise is added to the plurality of data included in the integrated data group when the first data group and the second data group are integrated into an integrated data group with a predetermined probability.

[0025] The information processing apparatus 1 transmits the generated first noise addition query and the first irreversible conversion query to the first device 2, and transmits the generated second noise addition query and the second irreversible conversion query to the second device 3.

[0026] The first device 2 executes a first noise addition query and a first irreversible conversion query received from the information processing device 1, and associates a data ID converted based on the first irreversible conversion query with a plurality of first data to which noise has been added based on the first noise addition query, generating a converted first data group including a plurality of first records. The first device 2 transmits the generated converted first data group to the information processing device 1.

[0027] The second device 3 executes a second noise addition query and a second irreversible conversion query received from the information processing device 1, and associates a data ID converted based on the second irreversible conversion query with a plurality of second data to which noise has been added based on the second noise addition query, generating a converted second data group including a plurality of second records. The second device 3 transmits the generated converted second data group to the information processing device 1.

[0028] In this way, in each of the first device 2 and the second device 3, after anonymizing the data group, the data group can be transmitted to the information processing device 1. Therefore, in the process of collecting user information, it is possible to prevent the leakage of user information that has not been anonymized.

[0029] The information processing device 1 generates an integrated data group by integrating the converted first data group and the converted second data group based on the data IDs of the plurality of first records included in the converted first data group received from the first device 2 and the data IDs of the plurality of second records included in the converted second data group received from the second device 3.

[0030] In this way, noise will be added to a plurality of data included in the integrated data group with a predetermined probability. Also, since the data ID is converted by the irreversible conversion query, it becomes difficult to identify an individual based on the converted data ID. Thereby, the information processing device 1 can ensure the privacy of the user information included in the integrated data.

[0031] [Functional Configuration of Information Processing Apparatus 1] Next, the functional configuration of the information processing apparatus 1 will be described. FIG. 2 is a diagram showing the functional configuration of the information processing apparatus 1.

[0032] As shown in FIG. 2, the information processing apparatus 1 includes a communication unit 11, a storage unit 12, and a control unit 13. The communication unit 11 is a communication interface for transmitting and receiving data to and from the first device 2, the second device 3, etc. via a communication network.

[0033] The storage unit 12 is a storage medium for storing various types of data, and includes a ROM (Read Only Memory), a RAM (Random Access Memory), a hard disk, an SSD (Solid State Drive), a flash memory, etc. The storage unit 12 stores the programs executed by the control unit 13. The storage unit 12 stores a program that causes the control unit 13 to function as a determination unit 131, a generation unit 132, a transmission unit 133, a data group acquisition unit 134, and an integration unit 135.

[0034] The control unit 13 is, for example, a CPU (Central Processing Unit). The control unit 13 functions as a determination unit 131, a generation unit 132, a transmission unit 133, a data group acquisition unit 134, and an integration unit 135 by executing the programs stored in the storage unit 12.

[0035] Hereinafter, when explaining the functions of the control unit 13, the first data group and the second data group will be described. FIG. 3 is a diagram showing an example of the first data group and the second data group. In FIG. 3, (A) shows the first data group, and (B) shows the second data group.

[0036] The first data group is a data group managed by a first operator, and is stored in a database provided in the first device 2 or a database provided in a server accessible by the first device 2. As shown in FIG. 3, the first data group includes a data ID as data identification information for identifying data, and n 1It includes a plurality of first records that associate a plurality of first data corresponding to each of the attributes. The data ID is, for example, a common user ID assigned to the user by a first operator and a second operator.

[0037] In the example shown in FIG. 3, the first data group is a data group that associates sales in stores operated by a first operator with the age of the user, and includes first data of items such as "age", "product category food", "product category daily necessities", and "purchase ranking" corresponding to each of the plurality of attributes. The first data group is assumed to indicate one table or a table generated by concatenating a plurality of tables, but is not limited thereto, and may be a view that refers to one or more tables.

[0038] The second data group is a data group managed by a second operator and is stored in a database provided in the second device 3 or a database provided in a server accessible by the second device 3. As shown in FIG. 3, the second data group includes a data ID as data identification information for identifying data and a plurality of second records that associate a plurality of second data corresponding to each of n 2 attributes. In the example shown in FIG. 3, the second data group is a data group that associates the age of the user with the visit history to facilities, and includes second data of items such as "gender", "visit place supermarket", and "visit place park" corresponding to each of the plurality of attributes. The second data group is assumed to indicate one table or a table generated by concatenating a plurality of tables, but is not limited thereto, and may be a view that refers to one or more tables.

[0039] The first data group and the second data group include information of the same user, and it is assumed that the data ID of the same user is common in the first data group and the second data group. Thereby, the first record included in the first data group and the second record included in the second data group can be concatenated using the data ID as a key.

[0040] Next, the functions of the control unit 13 will be described. The determination unit 131 determines the probability of adding noise to the first data group and the probability of adding noise to the second data group such that noise is added to a plurality of data included in the integrated data group obtained by integrating the first data group and the second data group with a predetermined probability.

[0041] The determination unit 131 determines a first parameter ε 1 and a second parameter ε 2 indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added.

[0042] When the determination unit 131 determines the first parameter ε 1 and the second parameter ε 2 local differential privacy will be described. First, let any data pair in a certain data group be x1 and x2. And when, for the data x, the function for adding random noise is R(x) and its output is y, the function R is defined to satisfy local differential privacy when the following formula (1) holds.

[0043]

Equation

[0044] Here, Pr[] is a random variable. Also, e is the natural logarithm, and ε is a parameter indicating the strength of privacy. Also, the local differential privacy with the privacy strength of ε is called ε-local differential privacy.

[0045] Examples of data processing that satisfy ε-local differential privacy include the following processing examples. For example, when the data x can take k values, based on the following formula (2), for the input of the data x, the data y is output.

[0046]

Equation

[0047] The determination unit 131 acquires first item information indicating items corresponding to each of a plurality of attributes constituting the first record from the first device 2, and acquires second item information indicating items corresponding to each of a plurality of attributes constituting the second record from the second device 3. The item information is information indicating items to be included in the integrated data among the plurality of items included in the first record. In the example shown in FIG. 3, the determination unit 131 acquires first item information indicating four items, namely, "age", "product category food", "product category daily necessities", and "purchase ranking". Further, the determination unit 131 acquires first item information indicating three items, namely, "gender", "visited place supermarket", and "visited place park".

[0048] 1 identifies n, which is the number of attributes corresponding to the first data, based on the acquired first item information, and identifies n, which is the number of attributes corresponding to the second data, based on the acquired second item information. 2 1 1 2 2 1 1 2 2 1 1 2 2 1 1 2 2 1 1 2 2 The sum of n, which is the number of attributes corresponding to the first data, and n, which is the number of attributes corresponding to the second data, is the number of attributes included in the integrated data group. The determination unit 131 determines a first parameter ε and a second parameter ε indicating the strength of privacy in local differential privacy satisfied by the first data and the second data of each of the plurality of attributes after noise is applied, based on the identified numbers of attributes n and n. For example, as shown in the following formula (3), the determination unit 131 determines the first parameter ε and the second parameter ε such that the data of each of the plurality of attributes satisfies (ε / n + n)-local differential privacy.

[0049]

Equation

[0050] As a result, ((ε / n 1 +n 2 ) - local differential privacy is applied, and the integrated data group obtained by aggregating the data of n 1 +n 2 attributes will satisfy ε-local differential privacy.

[0051] The generation unit 132 functions as a first generation unit and generates a first noise-added query. The first noise-added query adds noise to each of the plurality of first data included in the first data group such that noise is added to the plurality of data included in the integrated data group when the first data group and the second data group are integrated into an integrated data group with a predetermined probability. The predetermined probability is the probability that ε-local differential privacy is satisfied and is determined by the parameter ε indicating the strength of privacy. For example, the predetermined probability is calculated using Equation (2). As shown in Equation (2), the smaller ε is, the higher the probability that the data x is converted to another value.

[0052] For example, when noise is added to the first data of each of the n 1 attributes indicated by the first item information, the generation unit 132 generates a first noise-added query that adds noise to the first data of each of the n 1 attributes with noise added so that the first data satisfies ε 1 -local differential privacy. Here, ε 1 is the first parameter determined by the determination unit 131.

[0053] In addition, the generation unit 132 generates a first irreversible conversion query that is a query for irreversibly converting the data ID as the plurality of data identification information included in the first data group by a predetermined method. The predetermined method is, for example, a method of irreversibly converting the data ID using a hash function, but is not limited thereto, and other methods may be used as long as they are irreversible conversion methods.

[0054] Further, the generation unit 132 functions as a second generation unit and generates a second noise-adding query. The second noise-adding query is a query for adding noise to each of the plurality of second data included in the second data group such that noise is added to each of the plurality of data included in the integrated data group with a predetermined probability. For example, when noise is added to each of the n 2 attributes indicated by the second item information in the second data, the generation unit 132 generates a second noise-adding query for adding noise so that the second data for each of the n 2 attributes to which noise is added satisfies ε 2 -local differential privacy. Here, ε 2 is the second parameter determined by the determination unit 131.

[0055] Also, the generation unit 132 generates a second irreversible transformation query, which is a query for irreversibly transforming a plurality of data IDs included in the second data group by a predetermined method in the same manner as the first irreversible transformation query.

[0056] Note that the generation unit 132 may reduce the number of values that the data corresponding to at least one of the plurality of attributes included in the first data group and the second data group can take, and after reducing the number of values that the data can take, generate a noise-adding query for adding noise to each of the plurality of data. For example, when the data with the attribute "age" indicates the actual age of each of the plurality of users, the generation unit 132 generates a noise-adding query including a process of reducing the values that the data can take by changing the data to data indicating age groups such as "10s" and "20s". By doing so, the privacy of the user can be enhanced.

[0057] Further, the generation unit 132 generates a first irreversible conversion query that adds random data to each of the plurality of data IDs included in the first data group and then performs irreversible conversion by a predetermined method, and adds the same random data as the random data added to the data ID included in the first data group corresponding to the data ID to each of the plurality of data IDs included in the second data group, and then generates a second irreversible conversion query that performs irreversible conversion by a predetermined method. By doing so, the information processing apparatus 1 can reduce the risk that the converted data ID is decoded into the data ID before conversion.

[0058] Further, the generation unit 132 may generate a first update query, which is a query for updating the converted first data group by swapping the first data included in the first record included in the converted first data group with the first data included in other first records at a first ratio. Further, the generation unit 132 may generate a second update query, which is a query for updating the converted second data group by swapping the second data included in the second record included in the converted second data group with the second data included in other second records at a second ratio.

[0059] Here, the first ratio and the second ratio may be the same or different. Further, the first ratio and the second ratio may be changed according to the number of values that the data can take. For example, when the number of values that the data can take is large, the ratio at which the data is swapped may be increased.

[0060] Also, after the integrated data group is generated by the integration unit 135 described later, new records may be added to each of the first data group and the second data group, and it may be required to generate an integrated data group with the new records added. When noise addition is repeated a plurality of times for all of the first data groups and all of the second data groups, data groups of a plurality of variations corresponding to the same data group are generated. In this case, by analyzing the data groups of the plurality of variations, it becomes easier to infer the content of the data group before anonymization, and there arises a problem that the privacy risk such as an increase in the identifiability of the user increases. In contrast, the generation unit 132 may generate a noise addition query for adding noise only to the newly added records.

[0061] In this case, the first data group and the second data group include specific data that can be used to identify the newly added records. The specific data is, for example, date data indicating a date or a flag indicating whether a record is included in the integrated data. Then, when regenerating the first irreversible conversion query, the generation unit 132 generates a first noise addition query for adding noise to the newly added first record based on the specific data, and when regenerating the second irreversible conversion query, generates a second noise addition query for adding noise to the newly added second record based on the specific data. By doing so, the information processing apparatus 1 can suppress an increase in the privacy risk when providing the integrated data group.

[0062] The transmission unit 133 transmits the first noise-added query and the first irreversible conversion query generated by the generation unit 132 to the first device 2 corresponding to the provider of the first data group. Further, the transmission unit 133 transmits the second noise-added query and the second irreversible conversion query generated by the generation unit 132 to the second device 3 corresponding to the provider of the second data group. For example, the transmission unit 133 transmits the first noise-added query and the first irreversible conversion query to the first device 2 via an Internet VPN (Virtual Private Network) provided in advance between the information processing device 1 and the first device 2. Similarly, the transmission unit 133 transmits the second noise-added query and the second irreversible conversion query to the second device 3 via, for example, a second VPN provided in advance between the information processing device 1 and the second device 3.

[0063] Also, when the first update query and the second update query are generated by the generation unit 132, the transmission unit 133 transmits the first update query to the first device 2 and transmits the second update query to the second device 3.

[0064] By executing the query received from the information processing device 1, the first device 2 generates a converted first data group including a plurality of first records associating the converted data ID as data identification information converted from the data ID based on the first irreversible conversion query and a plurality of first data with noise added based on the first noise-added query. For example, when the first device 2 receives the first update query from the information processing device 1, the first device 2 executes the first update query before adding noise to the plurality of first data based on the first noise-added query. Thereafter, the first device 2 transmits the converted first data group to the information processing device 1 via, for example, the first VPN. Note that a device different from the first device 2 may transmit the converted first data group to the information processing device 1.

[0065] The second device 3 generates a converted second data group including a plurality of second records in which the converted data ID as data identification information converted from the data ID based on the second irreversible conversion query and a plurality of second data to which noise is added based on the second noise addition query are associated, by executing the second noise addition query and the second irreversible conversion query received from the information processing device 1. For example, when the second device 3 receives a second update query from the information processing device 1, the second device 3 executes the second update query before adding noise to the plurality of second data based on the second noise addition query. Thereafter, the second device 3 transmits the converted second data group to the information processing device 1 via, for example, a second VPN. Note that a device different from the second device 3 may transmit the converted second data group to the information processing device 1.

[0066] The data group acquisition unit 134 acquires the converted first data group and the converted second data group. For example, the data group acquisition unit 134 acquires the converted first data group and the converted second data group by receiving the converted first data group transmitted from the first device 2 and receiving the converted second data group transmitted from the second device 3.

[0067] FIG. 4 is a diagram showing an example of the converted first data group and the converted second data group. In FIG. 4, (A) shows the converted first data group, and (B) shows the converted second data group. Also, in FIG. 4, it can be confirmed that the same data ID included in the first data group and the second data group is converted into the same character string. Also, in FIG. 4, it can be confirmed that the data surrounded by the thick-framed cells is converted.

[0068] The integration unit 135 generates an integrated data group by integrating the converted first data group and the converted second data group based on the data IDs (converted data IDs) of the plurality of first records included in the converted first data group acquired by the data group acquisition unit 134 and the data IDs (converted data IDs) of the plurality of second records included in the converted second data group acquired by the data group acquisition unit 134. Specifically, the integration unit 135 generates an integrated data group by combining the first data group and the second data group using the converted data ID as a key. FIG. 5 is a diagram showing an example of the integrated data group. As shown in FIG. 5, it can be confirmed that the first data and the second data associated with the converted data ID included in both the first data group and the second data group are associated with each other.

[0069] The integration unit 135 generates an integrated data group including the converted data ID, the converted first data group, and the converted second data group, but is not limited thereto. The integration unit 135 integrates the converted first data group and the converted second data group based on the data IDs of the plurality of first records included in the converted first data group and the data IDs of the plurality of second records included in the converted second data group, and may generate an integrated data group by excluding the data ID. By doing so, the integrated data does not include the data ID, so the risk of restoring the first record and the second record from the integrated data based on the data ID can be reduced.

[0070] In addition, the integration unit 135 may further process the generated integrated data group to generate statistical data. Then, the integration unit 135 may perform correction to remove noise using a predetermined probability used for adding noise to the generated statistical data. For example, when calculating a statistical value using the integrated data, the integration unit 135 uses at least one of the values of the privacy strength parameters ε, ε 1 and ε 2 to statistically correct the statistical value.

[0071] For example, let P be the transition matrix when adding noise to the data of a certain attribute included in the integrated data, and let the elements included in the transition matrix P be p i,j . p i,j indicates the probability that the value i of a certain attribute randomly transitions to the value j, and is determined using, for example, the expression obtained by replacing x with i and y with j in the above-described expression (2). The integration unit 135 corrects the distribution Q = (q 1 , …, q d ) T of a certain attribute obtained by processing the integrated data into the distribution Q' using the transition matrix P and the following expression (4).

[0072] [Equation]

[0073] Here, when the data of a certain attribute is included in the first data group, the first parameter ε 1 is applied to ε included in the expression (2), and when the data of a certain attribute is included in the second data group, the second parameter ε 2 is applied to ε included in the expression (2) to construct the transition matrix P. Further, when only ε is known, the first parameter ε 1 and the second parameter ε 2 are derived using the expression (3), and it is assumed that the transition matrix P is constructed in the same manner. By doing so, the information processing apparatus 1 can generate statistical data with a high probability corresponding to the first data group and the second data group before noise is added.

[0074] Note that the integrated data group integrated by the integration unit 135 may be transmitted by the transmission unit 133 to the first device 2 and the second device 3. By doing so, in the first business operator, data analysis can be performed based on the second data collected by the second business operator, and in the second business operator, data analysis can be performed based on the first data collected by the first business operator.

[0075] [Operation Sequence] Next, the processing flow of the information processing apparatus 1 will be described. FIG. 6 is a sequence diagram showing the processing flow until the information processing apparatus 1 generates an integrated data group.

[0076] First, the determination unit 131 acquires first item information indicating items corresponding to each of a plurality of attributes constituting the first record from the first device 2 (S1), and acquires second item information indicating items corresponding to each of a plurality of attributes constituting the second record from the second device 3 (S2).

[0077] Subsequently, the determination unit 131 specifies the number of attributes corresponding to the data based on the acquired first item information and the acquired second item information (S3). Specifically, the determination unit 131 determines the number n of attributes corresponding to the first data 1 and the number n of attributes corresponding to the second data. 2 Then, the determination unit 131 determines the first parameter ε indicating the privacy strength in the local differential privacy satisfied by the first data and the second data of each of the plurality of attributes after noise is added, based on the specified numbers n of attributes 1 and n 2 (S4). 1 and the second parameter ε 2 (S4).

[0078] Subsequently, the generation unit 132 generates a first noise addition query, a second noise addition query, a first irreversible transformation query, and a second irreversible transformation query (S5). The generation unit 132 generates the first noise addition query based on the acquired first item information and the first parameter ε determined by the determination unit 131 1 and generates the second noise addition query based on the acquired second item information and the second parameter ε determined by the determination unit 131. 2 The generation unit 132 also generates a first irreversible transformation query for irreversibly transforming the data ID included in the first data, and a second irreversible transformation query for irreversibly transforming the data ID included in the second data.

[0079] Subsequently, the transmission unit 133 transmits the first noise addition query and the first irreversible conversion query to the first device 2 (S6), and transmits the second noise addition query and the second irreversible conversion query to the second device 3 (S7).

[0080] The first device 2 generates a first data group after conversion by executing the query received from the information processing device 1 (S8). The second device 3 generates a second data group after conversion by executing the query received from the information processing device 1 (S9). The first device 2 transmits the first data group after conversion to the information processing device 1 (S10), and the second device 3 transmits the second data group after conversion to the information processing device 1 (S11). The data group acquisition unit 134 receives the first data group after conversion transmitted from the first device 2 and receives the second data group after conversion transmitted from the second device 3.

[0081] The integration unit 135 generates an integrated data group by integrating the first data group after conversion and the second data group after conversion based on the converted data IDs of the plurality of first records included in the first data group after conversion received by the data group acquisition unit 134 and the converted data IDs of the plurality of second records included in the second data group after conversion acquired by the data group acquisition unit 134 (S12).

[0082] [Modification Example 1] In the above-described embodiment, the information processing device 1 generates two noise addition queries and two irreversible conversion queries corresponding to the first data group and the second data group, but is not limited thereto. The information processing device 1 may generate a noise addition query and an irreversible conversion query corresponding to three or more data groups.

[0083] For example, assume that there are k data groups (where k is an integer of 3 or more) associated with data IDs as data identification information, and the kth record included in the kth data group includes a plurality of kth data corresponding to each of n k attributes respectively.

[0084] In this case, the generation unit 132 functions as the k-th generation unit, and for each of the plurality of k-th data included in the k-th data group, when k data groups are integrated into an integrated data group, a query that adds noise to the plurality of data included in the integrated data group with a predetermined probability, i.e., the k-th noise addition query, and a query that irreversibly transforms the plurality of data identification information included in the k-th data group by a predetermined method, i.e., the k-th irreversible transformation query, are generated. Then, the generation unit 132 k When noise is added to the k-th data of each of the n k attributes, the generation unit 132 generates the k-th noise addition query that adds noise so that the k-th data of each of the n k attributes with added noise satisfies ε k -local differential privacy. However, the parameter ε k indicating the strength of privacy is ε 1 = ε / (n 2 + n k + ··· + n

[0085] )).

[0086] Also, the generation unit 132 generates the k-th irreversible transformation query that irreversibly transforms the plurality of users included in the k-th data group by a predetermined method. The transmission unit 133 transmits the generated k-th noise addition query and the k-th irreversible transformation query to the k-th device.

[0087] [Modification Example 2] In the above-described embodiment, the generation unit 132 generates a first update query for swapping the first data included in the first record included in the converted first data group with the first data included in another first record, and a second update query for swapping the second data included in the second record included in the converted second data group with the second data included in another second record. The first device 2 executes the first update query, and the second device 3 executes the second update query. However, the present invention is not limited to this. The information processing apparatus 1 may execute data swapping.

[0088] In this case, the control unit 13 includes an update unit that updates the integrated data group by swapping the data included in the record included in the integrated data group with the data included in another record at a third ratio. For example, the update unit updates the integrated data group by swapping some of the data corresponding to each of the plurality of items included in the record included in the integrated data group with the data of the same item included in another record at a third ratio.

[0089] Further, the update unit updates the converted first data group by swapping the first data included in the first record included in the converted first data group acquired by the data group acquisition unit 134 with the first data included in another first record at a first ratio, and updates the converted second data group by swapping the second data included in the second record included in the converted second data group acquired by the data group acquisition unit 134 with the second data included in another second record at a second ratio. For example, the update unit updates the converted first data group and the converted second data group by executing the first update query and the second update query generated by the generation unit 132. Then, the integration unit 135 generates integrated data by integrating the updated first data group and the updated second data group. By doing so, the information processing apparatus 1 can reduce the load related to the conversion of the data group in the first device 2 and the second device 3.

[0090] [Effect by Information Processing Apparatus 1] As described above, when the information processing apparatus 1 according to the present embodiment integrates the first data group and the second data group into an integrated data group, noise is added to the plurality of data included in the integrated data group with a predetermined probability. A first noise addition query for adding noise to the first data group, a second noise addition query for adding noise to the second data group, a first irreversible conversion query that is a query for irreversibly converting a plurality of data identification information included in the first data group, and a second irreversible conversion query that is a query for irreversibly converting a plurality of data identification information included in the second data group are generated. The first noise addition query and the first irreversible conversion query are transmitted to the first apparatus 2, and the second noise addition query and the second irreversible conversion query are transmitted to the second apparatus 3. Then, the information processing apparatus 1 includes a converted first data group including a plurality of first records in which the data identification information converted based on the first irreversible conversion query and the plurality of first data to which noise is added based on the first noise addition query are associated, a converted second data group including a plurality of second records in which the data identification information converted based on the second irreversible conversion query and the plurality of second data to which noise is added based on the second noise addition query are associated, and based on the data identification information included in these data groups, these data groups are integrated, the data identification information is excluded, and an integrated data group is generated. By doing so, the information processing apparatus 1 can prevent user information that has not been anonymized from leaking during the process of collecting user information.

[0091] Note that the present invention enables contribution to Goal 9, "Build the infrastructure for industry and innovation," of the Sustainable Development Goals (SDGs) led by the United Nations.

[0092] As described above, the present invention has been explained using embodiments. However, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist. For example, all or part of the device can be configured by functionally or physically dispersing and integrating it in any unit. Also, new embodiments resulting from any combination of a plurality of embodiments are included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination have the effects of the original embodiments combined.

Explanation of Reference Numerals

[0093] 1 Information processing device 2 First device 3 Second device 11 Communication unit 12 Storage unit 13 Control unit 131 Decision unit 132 Generation unit 133 Transmission unit 134 Data group acquisition unit 135 Integration unit S Information processing system

Claims

1. a first generating unit that generates a first noise-adding query for each of a plurality of first data included in a first data group including a plurality of first records associating first data with data identification information for identifying data and a second data group including a plurality of second records associating the data identification information with second data; the first noise-adding query is a query for adding noise to a plurality of data included in the integrated data group by integrating the first data group and the second data group into an integrated data group such that noise is added with a predetermined probability to the plurality of data included in the integrated data group; and a first irreversible conversion query is a query for irreversibly converting a plurality of the data identification information included in the first data group by a predetermined method; a second generating unit that generates a second noise-adding query, which is a query for adding noise to each of the plurality of second data included in the second data group in such a manner that noise is added to the plurality of data included in the integrated data group with the predetermined probability, and a second irreversible conversion query, which is a query for irreversibly converting the plurality of data identification information included in the second data group by the predetermined method; a transmission unit that transmits the first noise-added query and the first irreversible conversion query generated by the first generation unit to a first device corresponding to a provider of the first data group, and transmits the second noise-added query and the second irreversible conversion query generated by the second generation unit to a second device corresponding to a provider of the second data group; a data group acquiring unit that acquires a converted first data group including a plurality of first records that associate the data identification information converted based on the first irreversible conversion query with a plurality of first data to which noise has been added based on the first noise-adding query, and a converted second data group including a plurality of second records that associate the data identification information converted based on the second irreversible conversion query with a plurality of second data to which noise has been added based on the second noise-adding query; an integration unit that generates an integrated data group by integrating the converted first data group and the converted second data group based on the data identification information of each of a plurality of first records included in the converted first data group acquired by the data group acquisition unit and the data identification information of each of a plurality of second records included in the converted second data group acquired by the data group acquisition unit; An information processing device having the above configuration.

2. The first record is n 1 a plurality of first data corresponding to each of the attributes; The second record is n 2 a plurality of second data corresponding to the respective attributes; The plurality of data included in the integrated data group satisfies ε-local differential privacy, where ε is a parameter indicating the strength of privacy, The first generation unit 1 When noise is added to the first data of each attribute, 1 The first data of each attribute is ε 1 - Local differential privacy (where the parameter ε indicates the strength of privacy) 1 is ε 1 = ε / (n 1 +n 2 ) generating the first noise-added query such that noise is added to satisfy The second generation unit 2 When noise is added to the second data of each of the attributes, 2 The second data of each attribute is ε 2 - Local differential privacy (where the parameter ε indicates the strength of privacy) 2 is ε 2 = ε / (n 1 +n 2 ) generating the second noise-added query such that noise is added to satisfy The information processing device according to claim 1 .

3. There are k data groups associated with the data identification information (where k is an integer equal to or greater than 3); The kth generation unit generates a kth noise-adding query, which is a query for adding noise to each of the kth data included in the kth data group so that noise is added to the multiple data included in the integrated data group with a predetermined probability when the k data groups are integrated to form the integrated data group, and a kth irreversible conversion query, which is a query for irreversibly converting the multiple pieces of data identification information included in the kth data group by a predetermined method; The k-th record included in the k-th data group is k a plurality of k-th data items corresponding to the k attributes, The k generation unit k When noise is added to the k-th data of each attribute, k The k-th data of each attribute is ε k - Local differential privacy (where the parameter ε indicates the strength of privacy) k is ε k = ε / (n 1 +n 2 +...+n k ) to generate the kth noise-added query, The information processing device according to claim 2 .

4. First item information indicating items corresponding to each of a plurality of attributes constituting the first record is obtained from the first device, and second item information indicating items corresponding to each of a plurality of attributes constituting the second record is obtained from the second device, and n, which is the number of attributes corresponding to the first data, is calculated based on the obtained first item information. 1 and determining the number of attributes corresponding to the second data based on the acquired second item information. 2 The number of attributes identified is determined as n. 1 and the n 2 Based on the above, a first parameter ε indicating the strength of privacy in local differential privacy satisfied by the first data and the second data after noise is added is calculated. 1 and the second parameter ε 2 A determination unit for determining The information processing device according to claim 2 .

5. at least one of the first generation unit and the second generation unit reduces a number of possible values ​​of data corresponding to at least one attribute among a plurality of data corresponding to each of a plurality of attributes included in a data group, and generates the noise-adding query that adds the noise to each of the plurality of data after reducing the number of possible values ​​of the data. The information processing device according to claim 1 .

6. the first generation unit generates a first update query that is a query that updates the first data group after the conversion by replacing first data included in the first record included in the first data group after the conversion with the first data included in another first record at a first ratio; the second generation unit generates a second update query that is a query that updates the converted second data group by replacing second data included in the second records included in the converted second data group with the second data included in other second records at a second ratio; The transmission unit transmits the first update query generated by the first generation unit to the first device, and transmits the second update query generated by the second generation unit to the second device. The information processing device according to claim 1 .

7. an updating unit that updates the integrated data set by replacing data included in records included in the integrated data set with data included in other records at a third ratio; The information processing device according to claim 1 .

8. the first generation unit generates the first irreversible conversion query by adding random data to each of the plurality of pieces of data identification information included in the first data group and then performing irreversible conversion using the predetermined method; the second generation unit generates the second irreversible conversion query by performing irreversible conversion using the predetermined method after adding random data that is the same as the random data added to the data identification information included in the first data group and that corresponds to the data identification information, to each of the plurality of data identification information included in the second data group. The information processing device according to claim 1 .

9. the integration unit further processes the integrated data group to generate statistical data, and corrects the statistical data to remove the noise using the predetermined probability used to impart the noise. The information processing device according to claim 1 .

10. the first data group and the second data group include identification data that can be used to identify a newly added record; When regenerating the first irreversible conversion query, the first generation unit generates the first noise-added query that adds the noise to a newly added first record based on the identification data; When regenerating the second irreversible conversion query, the second generation unit generates the second noise-added query that adds the noise to a newly added second record based on the identification data. The information processing device according to claim 1 .

11. the integrating unit integrates the converted first data group and the converted second data group based on the data identification information of each of a plurality of first records included in the converted first data group and the data identification information of each of a plurality of second records included in the converted second data group, and generates the integrated data group by excluding the data identification information. The information processing device according to claim 1 .

12. Executed by the information processing device, a step of generating a first noise-adding query for each of a plurality of first data included in a first data group including a plurality of first records associating first data with data identification information for identifying data and a second data group including a plurality of second records associating the data identification information with second data, the first data group being one of the first data group and the second data group being one of the second data groups, the first noise-adding query being a query for adding noise to the plurality of data included in the integrated data group by integrating the first data group and the second data group into an integrated data group such that noise is added with a predetermined probability to the plurality of data included in the integrated data group, and a first irreversible conversion query being a query for irreversibly converting the plurality of data identification information included in the first data group by a predetermined method; generating a second noise-adding query that is a query for adding noise to each of the plurality of second data included in the second data group in such a manner that noise is added to the plurality of data included in the integrated data group with the predetermined probability, and a second irreversible conversion query that is a query for irreversibly converting the plurality of data identification information included in the second data group by the predetermined method; transmitting the generated first noise-added query and the generated first lossy conversion query to a first device corresponding to a provider of the first data group, and transmitting the generated second noise-added query and the generated second lossy conversion query to a second device corresponding to a provider of the second data group; acquiring a converted first data group including a plurality of first records associating the data identification information converted based on the first irreversible conversion query with a plurality of first data to which noise has been added based on the first noise-adding query, and acquiring a converted second data group including a plurality of second records associating the data identification information converted based on the second irreversible conversion query with a plurality of second data to which noise has been added based on the second noise-adding query; generating an integrated data group by integrating the converted first data group and the converted second data group based on the data identification information of each of a plurality of first records included in the acquired converted first data group and the data identification information of each of a plurality of second records included in the acquired converted second data group; An information processing method comprising the steps of:

Citation Information

Patent Citations

  • Coordination server program, business operator server program, and data coordinated system

    JP2021117679A