Information processor and information processing method

The information processing apparatus addresses the risk of unanonymized user information leakage by integrating noise addition and irreversible conversion queries into the data collection process from multiple operators, ensuring robust privacy protection.

JP2025088298AActive Publication Date: 2025-06-11KDDI CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023202916
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-06-11
Estimated Expiration
2043-11-30

AI Technical Summary

Technical Problem

Conventional methods for collecting user information from multiple operators anonymize data only after collection, risking leakage of unanonymized user information during the collection process.

Method used

An information processing apparatus and method that integrate user data from multiple operators by generating noise addition and irreversible conversion queries for each data group, ensuring anonymization and privacy protection during data integration.

Benefits of technology

Prevents the leakage of unanonymized user information by ensuring that noise is added and data is irreversibly converted with a predetermined probability, thus maintaining user privacy throughout the data collection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025088298000001_ABST
    Figure 2025088298000001_ABST
Patent Text Reader

Abstract

To provide an information processor and an information processing method for preventing leakage of user information, for which anonymization processing is not performed, in the process of collecting the user information.SOLUTION: An information processor 1 comprises: a generation section 132 that generates a first noise giving query for giving a first data group noise and a second noise giving query for giving a second data group noise so that noise is given to a plurality of pieces of data included in an integrated data group with a predetermined probability when the integrated data group is generated by integrating the first data group and the second data group, a first irreversible conversion query which is a query for irreversibly converting a plurality of pieces of data identification information included in the first data group, and a second irreversible conversion query which is a query for irreversibly converting a plurality of pieces of data identification information included in the second data group; and a transmission section 133 that transmits the first noise giving query and the first irreversible conversion query to a first device 2 and transmits the second noise giving query and the second irreversible conversion query to a second device 3.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus and an information processing method.

Background Art

[0002] Conventionally, it has been practiced to collect user information, which is information about users, from multiple operators and perform data analysis. In this case, in order to protect the privacy of users, at least a part of the user information collected from multiple operators is anonymized. For example, in Patent Document 1, irreversible conversion or the like is performed on data that is a joining key for joining multiple pieces of user information, and the personal information of users corresponding to each of the multiple operators is joined using the converted joining key, and a system is disclosed in which additional anonymization processing is performed on the joined data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, since anonymization processing is performed after collecting user information from each of multiple operators, there is a risk that user information on which anonymization processing has not been performed may leak during the process of collecting user information.

[0005] Therefore, the present invention has been made in view of these points, and an object thereof is to prevent user information on which anonymization processing has not been performed from leaking during the process of collecting user information.

Means for Solving the Problems

[0006] The information processing apparatus according to the first aspect of the present invention includes a first data group including a plurality of first records associating data identification information for identifying data with first data, and a second data group including a plurality of second records associating the data identification information with second data. For each of the plurality of first data included in the first data group, a first noise addition query, which is a query for adding noise to the plurality of data included in the integrated data group such that noise is added to the plurality of data included in the integrated data group with a predetermined probability when the first data group and the second data group are integrated into an integrated data group, and a first irreversible conversion query, which is a query for irreversibly converting the plurality of data identification information included in the first data group by a predetermined method, are generated by a first generation unit. For each of the plurality of second data included in the second data group, a second noise addition query, which is a query for adding noise to the plurality of data included in the integrated data group such that noise is added to the plurality of data included in the integrated data group with the predetermined probability, and a second irreversible conversion query, which is a query for irreversibly converting the plurality of data identification information included in the second data group by the predetermined method, are generated by a second generation unit. A transmission unit transmits the first noise addition query and the first irreversible conversion query generated by the first generation unit to a first device corresponding to the provider of the first data group, and transmits the second noise addition query and the second irreversible conversion query generated by the second generation unit to a second device corresponding to the provider of the second data group. A data group acquisition unit acquires a converted first data group including a plurality of first records associating the data identification information converted based on the first irreversible conversion query with the plurality of first data to which noise is added based on the first noise addition query, and a converted second data group including a plurality of second records associating the data identification information converted based on the second irreversible conversion query with the plurality of second data to which noise is added based on the second noise addition query. Based on the data identification information of each of the plurality of first records included in the converted first data group acquired by the data group acquisition unit and the data identification information of each of the plurality of second records included in the converted second data group acquired by the data group acquisition unit,An integration unit that generates the integrated data group obtained by integrating the first data group after the conversion and the second data group after the conversion.

[0007] The first record includes a plurality of first data corresponding to each of the n 1 attributes, the second record includes a plurality of second data corresponding to each of the n 2 attributes, the plurality of data included in the integrated data group satisfy ε-local differential privacy with a parameter ε indicating the strength of privacy, and the first generation unit is such that when noise is added to the first data of each of the n 1 attributes, the first data of each of the n 1 attributes to which noise is added satisfies ε 1 -local differential privacy (where the parameter ε 1 indicating the strength of privacy is ε 1 = ε / (n 1 + n 2 )) and generates the first noise addition query for adding noise, and the second generation unit is such that when noise is added to the second data of each of the n 2 attributes, the second data of each of the n 2 attributes to which noise is added satisfies ε 2 -local differential privacy (where the parameter ε 2 indicating the strength of privacy is ε 2 = ε / (n 1 + n 2 )) and may generate the second noise addition query for adding noise.

[0008] There are k data groups (where k is an integer greater than or equal to 3) associated with the data identification information. For each of the multiple k-th data included in the k-th data group, when the k data groups are integrated into the integrated data group, a k-th noise addition query that adds noise to the multiple data included in the integrated data group with a predetermined probability, and a k-th irreversible conversion query that irreversibly converts the multiple data identification information included in the k-th data group by a predetermined method. The k-th generation unit further has a k-th generation unit that generates a k-th record included in the k-th data group, where the k-th record includes multiple k-th data corresponding to each of n attributes. When noise is added to the k-th data of each of the n attributes, the k-th generation unit adds noise so that the k-th data of each of the n attributes to which noise is added satisfies ε-local differential privacy (where the parameter ε indicating the strength of privacy is ε = ε / (n + n + ··· + n)). k The k-th record included in the k-th data group includes multiple k-th data corresponding to each of n attributes. The k-th generation unit, when noise is added to the k-th data of each of the n attributes, adds noise so that the k-th data of each of the n attributes to which noise is added satisfies ε-local differential privacy (where the parameter ε indicating the strength of privacy is ε = ε / (n + n + ··· + n)). k The information processing apparatus obtains first item information indicating items corresponding to each of the multiple attributes constituting the first record from the first device, and obtains second item information indicating items corresponding to each of the multiple attributes constituting the second record from the second device. Based on the obtained first item information, it identifies n, which is the number of attributes corresponding to the first data, and based on the obtained second item information, it identifies n, which is the number of attributes corresponding to the second data. Based on the identified n and n, it determines the first parameter ε and the second parameter ε indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added. k -local differential privacy (where the parameter ε indicating the strength of privacy is ε k = ε / (n k + n k + ··· + n 1 + n 2 + ··· + n k )) may be satisfied.

[0009] The information processing apparatus obtains first item information indicating items corresponding to each of the multiple attributes constituting the first record from the first device, and obtains second item information indicating items corresponding to each of the multiple attributes constituting the second record from the second device. Based on the obtained first item information, it identifies n, which is the number of attributes corresponding to the first data, and based on the obtained second item information, it identifies n, which is the number of attributes corresponding to the second data. Based on the identified n and n, it determines the first parameter ε and the second parameter ε indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added. 1 The information processing apparatus obtains first item information indicating items corresponding to each of the multiple attributes constituting the first record from the first device, and obtains second item information indicating items corresponding to each of the multiple attributes constituting the second record from the second device. Based on the obtained first item information, it identifies n, which is the number of attributes corresponding to the first data, and based on the obtained second item information, it identifies n, which is the number of attributes corresponding to the second data. Based on the identified n and n, it determines the first parameter ε and the second parameter ε indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added. 2 The information processing apparatus obtains first item information indicating items corresponding to each of the multiple attributes constituting the first record from the first device, and obtains second item information indicating items corresponding to each of the multiple attributes constituting the second record from the second device. Based on the obtained first item information, it identifies n, which is the number of attributes corresponding to the first data, and based on the obtained second item information, it identifies n, which is the number of attributes corresponding to the second data. Based on the identified n and n, it determines the first parameter ε and the second parameter ε indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added. 1 and the n 2 Based on n and n, the first parameter ε and the second parameter ε indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added are determined. 1 and the second parameter ε 2It may have a determination unit that determines

[0010] At least one of the first generation unit and the second generation unit reduces the number of values that the data corresponding to at least one attribute among the plurality of data corresponding to the plurality of attributes included in the data group can take, and after reducing the number of values that the data can take, a noise addition query for adding the noise to each of the plurality of data may be generated.

[0011] The first generation unit generates a first update query, which is a query for updating the converted first data group by replacing the first data included in the first record included in the converted first data group with the first data included in another first record at a first ratio. The second generation unit generates a second update query, which is a query for updating the converted second data group by replacing the second data included in the second record included in the converted second data group with the second data included in another second record at a second ratio. The transmission unit may transmit the first update query generated by the first generation unit to the first device and transmit the second update query generated by the second generation unit to the second device.

[0012] The information processing apparatus may have an update unit that updates the integrated data group by replacing the data included in the record included in the integrated data group with the data included in another record at a third ratio.

[0013] The first generation unit generates a first irreversible conversion query that adds random data to each of the plurality of data identification information included in the first data group and then performs irreversible conversion by the predetermined method. The second generation unit generates a second irreversible conversion query that adds the same random data as the random data added to the data identification information included in the first data group corresponding to the data identification information to each of the plurality of data identification information included in the second data group and then performs irreversible conversion by the predetermined method.

[0014] The integration unit may further process the integrated data group to generate statistical data, and perform correction to remove the noise using the predetermined probability used for adding the noise to the statistical data.

[0015] The first data group and the second data group include identification data that can be used to identify newly added records. When the first generation unit regenerates the first irreversible conversion query, it generates the first noise addition query for adding the noise to the first records newly added based on the identification data. When the second generation unit regenerates the second irreversible conversion query, it may generate the second noise addition query for adding the noise to the second records newly added based on the identification data.

[0016] The integration unit may integrate the converted first data group and the converted second data group based on the data identification information of each of the plurality of first records included in the converted first data group and the data identification information of each of the plurality of second records included in the converted second data group, exclude the data identification information, and generate the integrated data group.

[0017] The information processing method according to the second aspect of the present invention includes a first data group including a plurality of first records associating data identification information for identifying data with first data, which is executed by an information processing apparatus, and a second data group including a plurality of second records associating the data identification information with second data. For each of the plurality of first data included in the first data group, a first noise addition query, which is a query for adding noise so that noise is added to a plurality of data included in the integrated data group with a predetermined probability when the first data group and the second data group are integrated into an integrated data group, and a first irreversible conversion query, which is a query for irreversibly converting a plurality of the data identification information included in the first data group by a predetermined method, are generated. For each of the plurality of second data included in the second data group, a second noise addition query, which is a query for adding noise so that noise is added to a plurality of data included in the integrated data group with the predetermined probability, and a second irreversible conversion query, which is a query for irreversibly converting a plurality of the data identification information included in the second data group by the predetermined method, are generated. The generated first noise addition query and the first irreversible conversion query are transmitted to a first apparatus corresponding to the provider of the first data group, and the generated second noise addition query and the second irreversible conversion query are transmitted to a second apparatus corresponding to the provider of the second data group. A converted first data group including a plurality of first records associating the data identification information converted based on the first irreversible conversion query with a plurality of first data to which noise is added based on the first noise addition query, and a converted second data group including a plurality of second records associating the data identification information converted based on the second irreversible conversion query with a plurality of second data to which noise is added based on the second noise addition query are acquired. Based on the data identification information of each of the plurality of first records included in the acquired converted first data group and the data identification information of each of the plurality of second records included in the acquired converted second data group, the integrated data group obtained by integrating the converted first data group and the converted second data group is generated.

Advantages of the Invention

[0018] According to the present invention, in the process of collecting user information, it is possible to prevent the outflow of user information that has not been anonymized.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0020] [Outline of Information Processing System S] FIG. 1 is a diagram for explaining the outline of the information processing system S. The information processing system S includes an information processing apparatus 1, a first apparatus 2 that manages a first data group, and a second apparatus 3 that manages a second data group, and generates an integrated data group by integrating the first data group and the second data group after anonymizing the user information included in the first data group and the second data group.

[0021] The information processing apparatus 1 is operated by, for example, an aggregation operator that provides a service for aggregating data and providing the aggregated data, and is communicably connected to external apparatuses such as the first apparatus 2 and the second apparatus 3 via a communication network (not shown) such as the Internet or a mobile phone line.

[0022] The first device 2 is operated by, for example, a first operator, and manages a first data group including a plurality of first records associating a data ID as data identification information for identifying data with first data. The second device 3 is operated by, for example, a second operator, and manages a second data group including a plurality of second records associating a common data ID with the data ID included in the first data group and second data.

[0023] The information processing apparatus 1 generates a first noise addition query for adding noise to each of the plurality of first data included in the first data group, and a first irreversible conversion query which is a query for irreversibly converting a plurality of data IDs included in the first data group by a predetermined method. The first noise addition query is, for example, a query for adding noise to each of the plurality of first data included in the first data group such that noise is added to a plurality of data included in the integrated data group when the first data group and the second data group are integrated into an integrated data group, with a predetermined probability. The query is, for example, assumed to be an SQL (Structured Query Language) statement executable in a relational database management system.

[0024] The information processing apparatus 1 generates a second noise addition query for adding noise to each of the plurality of second data included in the second data group, and a second irreversible conversion query which is a query for irreversibly converting a plurality of data IDs included in the second data group by a predetermined method. Similar to the first noise addition query, the second noise addition query is a query for adding noise to each of the plurality of second data included in the second data group such that noise is added to a plurality of data included in the integrated data group when the first data group and the second data group are integrated into an integrated data group, with a predetermined probability.

[0025] The information processing apparatus 1 transmits the generated first noise addition query and the first irreversible conversion query to the first device 2, and transmits the generated second noise addition query and the second irreversible conversion query to the second device 3.

[0026] The first device 2 executes a first noise addition query and a first irreversible conversion query received from the information processing device 1, and associates a data ID converted based on the first irreversible conversion query with a plurality of first data to which noise has been added based on the first noise addition query, generating a converted first data group including a plurality of first records. The first device 2 transmits the generated converted first data group to the information processing device 1.

[0027] The second device 3 executes a second noise addition query and a second irreversible conversion query received from the information processing device 1, and associates a data ID converted based on the second irreversible conversion query with a plurality of second data to which noise has been added based on the second noise addition query, generating a converted second data group including a plurality of second records. The second device 3 transmits the generated converted second data group to the information processing device 1.

[0028] In this way, in each of the first device 2 and the second device 3, after performing anonymization processing on the data group, the data group can be transmitted to the information processing device 1. Therefore, in the process of collecting user information, it is possible to prevent the outflow of user information that has not undergone anonymization processing.

[0029] The information processing device 1 generates an integrated data group by integrating the converted first data group and the converted second data group based on the data IDs of the plurality of first records included in the converted first data group received from the first device 2 and the data IDs of the plurality of second records included in the converted second data group received from the second device 3.

[0030] In this way, noise will be added to a plurality of data included in the integrated data group with a predetermined probability. Also, since the data ID is converted by the irreversible conversion query, it becomes difficult to identify an individual based on the converted data ID. As a result, the information processing device 1 can ensure the privacy of the user information included in the integrated data.

[0031] [Functional Configuration of Information Processing Apparatus 1] Next, the functional configuration of the information processing apparatus 1 will be described. FIG. 2 is a diagram showing the functional configuration of the information processing apparatus 1.

[0032] As shown in FIG. 2, the information processing apparatus 1 includes a communication unit 11, a storage unit 12, and a control unit 13. The communication unit 11 is a communication interface for transmitting and receiving data to and from the first device 2, the second device 3, etc. via a communication network.

[0033] The storage unit 12 is a storage medium for storing various data, and includes a ROM (Read Only Memory), a RAM (Random Access Memory), a hard disk, an SSD (Solid State Drive), a flash memory, etc. The storage unit 12 stores a program executed by the control unit 13. The storage unit 12 stores a program that causes the control unit 13 to function as a determination unit 131, a generation unit 132, a transmission unit 133, a data group acquisition unit 134, and an integration unit 135.

[0034] The control unit 13 is, for example, a CPU (Central Processing Unit). The control unit 13 functions as a determination unit 131, a generation unit 132, a transmission unit 133, a data group acquisition unit 134, and an integration unit 135 by executing a program stored in the storage unit 12.

[0035] Hereinafter, when explaining the functions of the control unit 13, a first data group and a second data group will be described. FIG. 3 is a diagram showing an example of the first data group and the second data group. In FIG. 3, (A) shows the first data group, and (B) shows the second data group.

[0036] The first data group is a data group managed by a first operator, and is stored in a database provided in the first device 2 or a database provided in a server accessible by the first device 2. As shown in FIG. 3, the first data group includes a data ID as data identification information for identifying data, and n 1It includes a plurality of first records that associate a plurality of first data corresponding to each of the attributes. The data ID is, for example, a common user ID assigned to the user by the first operator and the second operator.

[0037] In the example shown in FIG. 3, the first data group is a data group that associates sales in a store operated by the first operator with the age of the user, and includes first data of items such as "age", "product category food", "product category daily necessities", and "purchase ranking" corresponding to each of the plurality of attributes. The first data group is assumed to represent one table or a table generated by concatenating a plurality of tables, but is not limited thereto, and may be a view that refers to one or more tables.

[0038] The second data group is a data group managed by the second operator and is stored in a database provided in the second device 3 or a database provided in a server accessible by the second device 3. As shown in FIG. 3, the second data group includes a data ID as data identification information for identifying data, and n 2 It includes a plurality of second records that associate a plurality of second data corresponding to each of the attributes. In the example shown in FIG. 3, the second data group is a data group that associates the age of the user with the visit history to the facility, and includes second data of items such as "gender", "visit location supermarket", and "visit location park" corresponding to each of the plurality of attributes. The second data group is assumed to represent one table or a table generated by concatenating a plurality of tables, but is not limited thereto, and may be a view that refers to one or more tables.

[0039] The first data group and the second data group include information of the same user, and it is assumed that the data IDs of the same user are common in the first data group and the second data group. Thereby, the first record included in the first data group and the second record included in the second data group can be concatenated using the data ID as a key.

[0040] Next, the functions of the control unit 13 will be described. The determination unit 131 determines the probability of adding noise to the first data group and the probability of adding noise to the second data group so that noise is added to a plurality of data included in the integrated data group obtained by integrating the first data group and the second data group with a predetermined probability.

[0041] The determination unit 131 determines a first parameter ε 1 and a second parameter ε 2 indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added.

[0042] When the determination unit 131 determines the first parameter ε 1 and the second parameter ε 2 Local differential privacy will be described. First, let any data pair in a certain data group be x1 and x2. And when, for the data x, the function for adding random noise is R(x) and its output is y, the function R is defined to satisfy local differential privacy when the following formula (1) holds.

[0043]

Equation

[0044] Here, Pr[] is a random variable. Also, e is the natural logarithm, and ε is a parameter indicating the strength of privacy. Also, the local differential privacy with the privacy strength of ε is called ε-local differential privacy.

[0045] Examples of data processing that satisfy ε-local differential privacy include the following processing examples. For example, when the data x can take k values, based on the following formula (2), for the input of the data x, the data y is output.

[0046]

Equation

[0047] The determination unit 131 acquires first item information indicating items corresponding to each of a plurality of attributes constituting the first record from the first device 2, and acquires second item information indicating items corresponding to each of a plurality of attributes constituting the second record from the second device 3. The item information is information indicating items to be included in the integrated data among the plurality of items included in the first record. In the example shown in FIG. 3, the determination unit 131 acquires first item information indicating four items, namely, "age", "product category food", "product category daily necessities", and "purchase ranking". Further, the determination unit 131 acquires first item information indicating three items, namely, "gender", "visited place supermarket", and "visited place park".

[0048] The determination unit 131 specifies n, which is the number of attributes corresponding to the first data, based on the acquired first item information 1 and specifies n, which is the number of attributes corresponding to the second data, based on the acquired second item information 2 . n, which is the number of attributes corresponding to the first data 1 and n, which is the number of attributes corresponding to the second data 2 The sum of is the number of attributes included in the integrated data group. The determination unit 131 determines a first parameter ε 1 and a second parameter ε 2 indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data of each of the plurality of attributes after noise is added, based on the specified number of attributes n 1 and the second parameter ε 2 . For example, as shown in the following formula (3), the determination unit 131 determines the first parameter ε 1 and the second parameter ε 2 such that the data of each of the plurality of attributes satisfies (ε / n 1 +n 2 )-local differential privacy is applied

[0049]

Equation

[0050] As a result, ((ε / n 1 + n 2 ))-local differential privacy is applied, and the integrated data group obtained by aggregating the data of n 1 + n 2 attributes will satisfy ε-local differential privacy.

[0051] The generation unit 132 functions as a first generation unit and generates a first noise-added query. The first noise-added query adds noise to each of the plurality of first data included in the first data group such that noise is added to the plurality of data included in the integrated data group when the first data group and the second data group are integrated into an integrated data group, with a predetermined probability. The predetermined probability is the probability that ε-local differential privacy is satisfied, and is determined by the parameter ε indicating the strength of privacy. For example, the predetermined probability is calculated using Equation (2). As shown in Equation (2), the smaller ε is, the higher the probability that the data x is converted to another value.

[0052] For example, when noise is added to each of the n 1 attributes indicated by the first item information in the first data, the generation unit 132 generates a first noise-added query that adds noise to the first data of each of the n 1 attributes with noise added so as to satisfy ε 1 -local differential privacy. Here, ε 1 is the first parameter determined by the determination unit 131.

[0053] In addition, the generation unit 132 generates a first irreversible conversion query, which is a query for irreversibly converting the data ID as the plurality of data identification information included in the first data group by a predetermined method. The predetermined method is, for example, a method of irreversibly converting the data ID using a hash function, but is not limited thereto, and other methods may be used as long as they are irreversible conversion methods.

[0054] In addition, the generation unit 132 functions as a second generation unit and generates a second noise-adding query. The second noise-adding query is a query that adds noise to each of the plurality of second data included in the second data group such that noise is added to each of the plurality of data included in the integrated data group with a predetermined probability. For example, when noise is added to each of the n attributes indicated by the second item information, the generation unit 132 adds noise to each of the second data of the n attributes to which noise has been added so that the second data satisfies ε-local differential privacy. Here, ε is the second parameter determined by the determination unit 131. 2 When noise is added to each of the second data of the n attributes, the generation unit 132 adds noise to each of the second data of the n attributes to which noise has been added so that the second data satisfies ε-local differential privacy. Here, ε is the second parameter determined by the determination unit 131. 2 When noise is added to each of the second data of the n attributes, the generation unit 132 adds noise to each of the second data of the n attributes to which noise has been added so that the second data satisfies ε-local differential privacy. Here, ε is the second parameter determined by the determination unit 131. 2 - Generates a second noise-adding query that adds noise so as to satisfy local differential privacy. Here, ε is the second parameter determined by the determination unit 131. 2 is the second parameter determined by the determination unit 131.

[0055] In addition, the generation unit 132 generates a second irreversible transformation query, which is a query that irreversibly transforms a plurality of data IDs included in the second data group by a predetermined method in the same manner as the first irreversible transformation query.

[0056] Note that the generation unit 132 may reduce the number of values that the data corresponding to at least one of the plurality of attributes included in the first data group and the second data group can take, and after reducing the number of values that the data can take, generate a noise-adding query that adds noise to each of the plurality of data. For example, when the data with the attribute "age" indicates the actual age of each of the plurality of users, the generation unit 132 changes the data to data indicating age groups such as "teens" and "twenties" for the data, thereby generating a noise-adding query including a process of reducing the values that the data can take. By doing so, the privacy of the user can be enhanced.

[0057] Further, the generation unit 132 generates a first irreversible conversion query that adds random data to each of the plurality of data IDs included in the first data group and then performs irreversible conversion by a predetermined method, and adds the same random data as the random data added to the data ID included in the first data group corresponding to the data ID to each of the plurality of data IDs included in the second data group, and then generates a second irreversible conversion query that performs irreversible conversion by a predetermined method. By doing so, the information processing apparatus 1 can reduce the risk that the converted data ID is decoded into the data ID before conversion.

[0058] Further, the generation unit 132 may generate a first update query, which is a query for updating the converted first data group by swapping the first data included in the first record included in the converted first data group with the first data included in other first records at a first ratio. Further, the generation unit 132 may generate a second update query, which is a query for updating the converted second data group by swapping the second data included in the second record included in the converted second data group with the second data included in other second records at a second ratio.

[0059] Here, the first ratio and the second ratio may be the same or different. Further, the first ratio and the second ratio may be changed according to the number of values that the data can take. For example, when the number of values that the data can take is large, the ratio at which the data is swapped may be increased.

[0060] Also, after the integrated data group is generated by the integration unit 135 described later, new records may be added to each of the first data group and the second data group, and it may be required to generate an integrated data group with new records added. When noise addition is repeated a plurality of times for all of the first data groups and all of the second data groups, data groups of a plurality of variations corresponding to the same data group are generated. In this case, by analyzing the data groups of the plurality of variations, it becomes easier to infer the content of the data group before anonymization, and there arises a problem that privacy risks such as an increase in the identifiability of the user increase. In contrast, the generation unit 132 may generate a noise addition query that adds noise only to the newly added records.

[0061] In this case, the first data group and the second data group include specific data that can be used to identify the newly added records. The specific data is, for example, date data indicating a date or a flag indicating whether a record is included in the integrated data. Then, when the first irreversible conversion query is regenerated, the generation unit 132 generates a first noise addition query that adds noise to the newly added first record based on the specific data, and when the second irreversible conversion query is regenerated, the generation unit 132 generates a second noise addition query that adds noise to the newly added second record based on the specific data. By doing so, the information processing apparatus 1 can suppress an increase in privacy risk when providing the integrated data group.

[0062] The transmission unit 133 transmits the first noise-added query and the first irreversible transformation query generated by the generation unit 132 to the first device 2 corresponding to the provider of the first data group. Further, the transmission unit 133 transmits the second noise-added query and the second irreversible transformation query generated by the generation unit 132 to the second device 3 corresponding to the provider of the second data group. The transmission unit 133 transmits the first noise-added query and the first irreversible transformation query to the first device 2 via, for example, an Internet VPN (Virtual Private Network) provided by a first cloud service provided in advance between the information processing device 1 and the first device 2. Similarly, the transmission unit 133 transmits the second noise-added query and the second irreversible transformation query to the second device 3 via, for example, a second VPN provided in advance between the information processing device 1 and the second device 3.

[0063] Further, when the first update query and the second update query are generated by the generation unit 132, the transmission unit 133 transmits the first update query to the first device 2 and transmits the second update query to the second device 3.

[0064] The first device 2 generates a converted first data group including a plurality of first records in which converted data IDs as data identification information converted from the data IDs based on the first irreversible transformation query and a plurality of first data with noise added based on the first noise-added query are associated by executing the query received from the information processing device 1. When the first device 2 receives the first update query from the information processing device 1, for example, the first device 2 executes the first update query before adding noise to the plurality of first data based on the first noise-added query. Thereafter, the first device 2 transmits the converted first data group to the information processing device 1 via, for example, the first VPN. Note that a device different from the first device 2 may transmit the converted first data group to the information processing device 1.

[0065] The second device 3 generates a converted second data group including a plurality of second records associating the converted data ID as data identification information converted from the data ID based on the second irreversible conversion query and a plurality of second data with noise added based on the second noise addition query by executing the second noise addition query and the second irreversible conversion query received from the information processing device 1. For example, when the second device 3 receives a second update query from the information processing device 1, the second device 3 executes the second update query before adding noise to the plurality of second data based on the second noise addition query. Thereafter, the second device 3 transmits the converted second data group to the information processing device 1 via, for example, a second VPN. Note that a device different from the second device 3 may transmit the converted second data group to the information processing device 1.

[0066] The data group acquisition unit 134 acquires the converted first data group and the converted second data group. For example, the data group acquisition unit 134 acquires the converted first data group and the converted second data group by receiving the converted first data group transmitted from the first device 2 and receiving the converted second data group transmitted from the second device 3.

[0067] FIG. 4 is a diagram showing an example of the converted first data group and the converted second data group. In FIG. 4, (A) shows the converted first data group, and (B) shows the converted second data group. Also, in FIG. 4, it can be confirmed that the same data ID included in the first data group and the second data group is converted into the same character string. Also, in FIG. 4, it can be confirmed that the data surrounded by the thick-frame cells is converted.

[0068] The integration unit 135 generates an integrated data group by integrating the converted first data group and the converted second data group based on the data IDs (converted data IDs) of the plurality of first records included in the converted first data group acquired by the data group acquisition unit 134 and the data IDs (converted data IDs) of the plurality of second records included in the converted second data group acquired by the data group acquisition unit 134. Specifically, the integration unit 135 generates an integrated data group by combining the first data group and the second data group using the converted data ID as a key. FIG. 5 is a diagram showing an example of the integrated data group. As shown in FIG. 5, it can be confirmed that the first data and the second data associated with the converted data ID included in both the first data group and the second data group are associated with each other.

[0069] The integration unit 135 generates an integrated data group including the converted data ID, the converted first data group, and the converted second data group, but is not limited thereto. The integration unit 135 integrates the converted first data group and the converted second data group based on the data IDs of the plurality of first records included in the converted first data group and the data IDs of the plurality of second records included in the converted second data group, and may generate an integrated data group by excluding the data ID. By doing so, the integrated data does not include the data ID, so the risk of restoring the first record and the second record from the integrated data based on the data ID can be reduced.

[0070] Further, the integration unit 135 may further process the generated integrated data group to generate statistical data. Then, the integration unit 135 may perform correction to remove noise using a predetermined probability used for adding noise to the generated statistical data. For example, when calculating a statistical value using the integrated data, the integration unit 135 uses at least one of the values of the privacy strength parameters ε, ε 1 and ε 2 to statistically correct the statistical value.

[0071] For example, let P be the transition matrix when adding noise to the data of a certain attribute included in the integrated data, and let the element included in the transition matrix P be p i,j . p i,j represents the probability that the value i of a certain attribute randomly transitions to the value j, and is determined, for example, using the formula obtained by replacing x with i and y with j in the above-described formula (2). The integration unit 135 corrects the distribution Q = (q 1 , …, q d ) T of a certain attribute obtained by processing the integrated data to the distribution Q' using the transition matrix P and the following formula (4).

[0072] [Equation]

[0073] Here, when the data of a certain attribute is included in the first data group, the first parameter ε 1 is applied to ε included in formula (2), and when the data of a certain attribute is included in the second data group, the second parameter ε 2 is applied to ε included in formula (2) to construct the transition matrix P. Also, when only ε is known, the first parameter ε 1 and the second parameter ε 2 are derived using formula (3), and the transition matrix P is similarly constructed. By doing so, the information processing apparatus 1 can generate statistical data with a high probability corresponding to the first data group and the second data group before noise is added.

[0074] Note that the integrated data group integrated by the integration unit 135 may be transmitted by the transmission unit 133 to the first device 2 and the second device 3. By doing so, in the first operator, data analysis can be performed based on the second data collected by the second operator, and in the second operator, data analysis can be performed based on the first data collected by the first operator.

[0075] [Operation Sequence] Next, the processing flow of the information processing apparatus 1 will be described. FIG. 6 is a sequence diagram showing the processing flow until the information processing apparatus 1 generates an integrated data group.

[0076] First, the determination unit 131 acquires first item information indicating items corresponding to each of a plurality of attributes constituting the first record from the first device 2 (S1), and acquires second item information indicating items corresponding to each of a plurality of attributes constituting the second record from the second device 3 (S2).

[0077] Subsequently, the determination unit 131 specifies the number of attributes corresponding to the data based on the acquired first item information and the acquired second item information (S3). Specifically, the determination unit 131 determines the number n of attributes corresponding to the first data 1 and the number n of attributes corresponding to the second data. 2 Then, the determination unit 131 determines the first parameter ε indicating the privacy strength in the local differential privacy satisfied by the first data and the second data of each of the plurality of attributes after noise is added based on the specified numbers n of attributes 1 and n. 2 And the second parameter ε. 1 (S4). 2

[0078] Subsequently, the generation unit 132 generates a first noise addition query, a second noise addition query, a first irreversible transformation query, and a second irreversible transformation query (S5). The generation unit 132 generates a first noise addition query based on the acquired first item information and the first parameter ε determined by the determination unit 131, and generates a second noise addition query based on the acquired second item information and the second parameter ε determined by the determination unit 131. Further, the generation unit 132 generates a first irreversible transformation query for irreversibly transforming the data ID included in the first data and a second irreversible transformation query for irreversibly transforming the data ID included in the second data. 1 2

[0079] Subsequently, the transmission unit 133 transmits the first noise addition query and the first irreversible conversion query to the first device 2 (S6), and transmits the second noise addition query and the second irreversible conversion query to the second device 3 (S7).

[0080] The first device 2 generates a first data group after conversion by executing the query received from the information processing device 1 (S8). The second device 3 generates a second data group after conversion by executing the query received from the information processing device 1 (S9). The first device 2 transmits the first data group after conversion to the information processing device 1 (S10), and the second device 3 transmits the second data group after conversion to the information processing device 1 (S11). The data group acquisition unit 134 receives the first data group after conversion transmitted from the first device 2 and receives the second data group after conversion transmitted from the second device 3.

[0081] The integration unit 135 generates an integrated data group by integrating the first data group after conversion and the second data group after conversion based on the converted data IDs of the plurality of first records included in the first data group after conversion received by the data group acquisition unit 134 and the converted data IDs of the plurality of second records included in the second data group after conversion acquired by the data group acquisition unit 134 (S12).

[0082] [Modification Example 1] In the above-described embodiment, the information processing device 1 generates two noise addition queries and two irreversible conversion queries corresponding to the first data group and the second data group, but is not limited thereto. The information processing device 1 may generate a noise addition query and an irreversible conversion query corresponding to three or more data groups.

[0083] For example, assume that there are k data groups (where k is an integer of 3 or more) associated with data IDs as data identification information, and the kth record included in the kth data group includes a plurality of kth data corresponding to n k attributes respectively.

[0084] In this case, the generation unit 132 functions as the k-th generation unit, and for each of the plurality of k-th data included in the k-th data group, when k data groups are integrated into an integrated data group, a k-th noise addition query that adds noise to the plurality of data included in the integrated data group with a predetermined probability, and a k-th irreversible conversion query that irreversibly converts the plurality of data identification information included in the k-th data group by a predetermined method are generated. Then, the generation unit 132 k When noise is added to the k-th data of each of the n k attributes, the k-th noise addition query that adds noise so that the k-th data of each of the n k attributes to which noise is added satisfies ε k -local differential privacy is generated. However, the parameter ε k indicating the strength of privacy is ε 1 = ε / (n 2 + n k + ··· + n

[0085] )).

[0086] Also, the generation unit 132 generates a k-th irreversible conversion query that irreversibly converts the plurality of users included in the k-th data group by a predetermined method. The transmission unit 133 transmits the generated k-th noise addition query and the k-th irreversible conversion query to the k-th device.

[0087] [Modification Example 2] In the above-described embodiment, the generation unit 132 generates a first update query for swapping the first data included in the first record included in the converted first data group with the first data included in another first record, and a second update query for swapping the second data included in the second record included in the converted second data group with the second data included in another second record. The first device 2 executes the first update query, and the second device 3 executes the second update query. However, the present invention is not limited to this. The information processing apparatus 1 may execute data swapping.

[0088] In this case, the control unit 13 includes an update unit that updates the integrated data group by swapping the data included in the records included in the integrated data group with the data included in other records at a third ratio. For example, the update unit updates the integrated data group by swapping some of the data corresponding to each of a plurality of items included in the records included in the integrated data group with the data of the same item included in other records at a third ratio.

[0089] Further, the update unit updates the converted first data group by swapping the first data included in the first record included in the converted first data group acquired by the data group acquisition unit 134 with the first data included in another first record at a first ratio, and updates the converted second data group by swapping the second data included in the second record included in the converted second data group acquired by the data group acquisition unit 134 with the second data included in another second record at a second ratio. For example, the update unit updates the converted first data group and the converted second data group by executing the first update query and the second update query generated by the generation unit 132. Then, the integration unit 135 generates integrated data by integrating the updated first data group and the updated second data group. By doing so, the information processing apparatus 1 can reduce the load related to the conversion of the data group in the first device 2 and the second device 3.

[0090] [Effect by Information Processing Apparatus 1] As described above, when the integrated data group is formed by integrating the first data group and the second data group, the information processing apparatus 1 according to the present embodiment generates a first noise addition query for adding noise to the first data group so that noise is added to a plurality of data included in the integrated data group with a predetermined probability, a second noise addition query for adding noise to the second data group, a first irreversible conversion query which is a query for irreversibly converting a plurality of data identification information included in the first data group, and a second irreversible conversion query which is a query for irreversibly converting a plurality of data identification information included in the second data group. Then, the information processing apparatus 1 transmits the first noise addition query and the first irreversible conversion query to the first apparatus 2, and transmits the second noise addition query and the second irreversible conversion query to the second apparatus 3. Then, the information processing apparatus 1 obtains a converted first data group including a plurality of first records in which the data identification information converted based on the first irreversible conversion query is associated with the plurality of first data to which noise is added based on the first noise addition query, and a converted second data group including a plurality of second records in which the data identification information converted based on the second irreversible conversion query is associated with the plurality of second data to which noise is added based on the second noise addition query. Based on the data identification information included in these data groups, these data groups are integrated, and the data identification information is excluded to generate an integrated data group. By doing so, the information processing apparatus 1 can prevent user information that has not been anonymized from leaking during the process of collecting user information.

[0091] Note that the present invention makes it possible to contribute to Goal 9, "Build the infrastructure for industry and innovation," of the Sustainable Development Goals (SDGs) led by the United Nations.

[0092] As described above, the present invention has been described using embodiments. However, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist. For example, all or part of the device can be configured by functionally or physically dispersing and integrating it in any unit. Also, new embodiments resulting from any combination of a plurality of embodiments are included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination have the effects of the original embodiments combined.

Explanation of Signs

[0093] 1 Information processing device 2 First device 3 Second device 11 Communication unit 12 Storage unit 13 Control unit 131 Decision unit 132 Generation unit 133 Transmission unit 134 Data group acquisition unit 135 Integration unit S Information processing system

Claims

1. A first data group including a plurality of first records associating data identification information for identifying data with first data, and a plurality of second records associating the data identification information with second data. For each of the plurality of first data included in the first data group, a first noise addition query that adds noise so that noise is added to the plurality of data included in the integrated data group with a predetermined probability when the first data group and the second data group are integrated into an integrated data group, and a first irreversible conversion query that irreversibly converts the plurality of data identification information included in the first data group by a predetermined method. A first generation unit that generates; For each of the plurality of second data included in the second data group, a second noise addition query that adds noise so that noise is added to the plurality of data included in the integrated data group with the predetermined probability, and a second irreversible conversion query that irreversibly converts the plurality of data identification information included in the second data group by the predetermined method. A second generation unit that generates; A transmission unit that transmits the first noise addition query and the first irreversible conversion query generated by the first generation unit to a first device corresponding to a provider of the first data group, and transmits the second noise addition query and the second irreversible conversion query generated by the second generation unit to a second device corresponding to a provider of the second data group; A converted first data group including a plurality of first records associating the data identification information converted based on the first irreversible conversion query with the plurality of first data to which noise has been added based on the first noise addition query, and a converted second data group including a plurality of second records associating the data identification information converted based on the second irreversible conversion query with the plurality of second data to which noise has been added based on the second noise addition query. A data group acquisition unit that acquires; An integration unit that generates the integrated data group by integrating the converted first data group and the converted second data group based on the data identification information of each of the plurality of first records included in the converted first data group acquired by the data group acquisition unit and the data identification information of each of the plurality of second records included in the converted second data group acquired by the data group acquisition unit; An information processing apparatus having

2. The first record is n 1 includes a plurality of first data corresponding to each of the n attributes, The second record is n 2 including a plurality of second data corresponding to each of the n attributes, The plurality of data included in the integrated data group satisfy ε-local differential privacy with a parameter ε indicating the strength of privacy, The first generation unit is the n 1 When noise is added to the first data of each of the n attributes, the first data of each of the n attributes with noise added is ε 1 - local differential privacy (where the parameter ε indicating the strength of privacy is ε 1 = ε / (n 1 is ε 1 = ε / (n 1 + n 2 )) to generate the first noise addition query that adds noise so as to satisfy the condition, The second generation unit 2 When noise is added to the second data of each of the attributes, 2 The second data of each attribute is ε 2 - Local differential privacy (where the parameter ε indicates the strength of privacy) 2 is ε 2 = ε / (n 1 +n 2 ) generating the second noise-added query such that noise is added to satisfy The information processing apparatus according to claim 1.

3. There are k data groups (where k is an integer of 3 or more) associated with the data identification information, A k-th noise addition query that adds noise to a plurality of data included in the integrated data group with a predetermined probability when the k data groups are integrated into the integrated data group for each of the plurality of k-th data included in the k-th data group, and a k-th irreversible conversion query that irreversibly converts a plurality of the data identification information included in the k-th data group by a predetermined method, and a k-th generation unit that generates the above is further provided, The k-th record included in the k-th data group has n k pieces of k-th data corresponding to each of the attributes, The k-th generation unit is the n k When noise is added to the k-th data of each of the n k attributes, the k-th data of each of the n k attributes with noise added satisfies ε k -local differential privacy (where the parameter ε k indicating the strength of privacy is ε 1 = ε / (n 2 + n k +... + n )), and generates the k-th noise-adding query that adds noise so as to satisfy the condition. The information processing apparatus according to claim 2.

4. Acquire first item information indicating items corresponding to each of a plurality of attributes constituting the first record from the first device, and acquire second item information indicating items corresponding to each of a plurality of attributes constituting the second record from the second device, and determine n which is the number of attributes corresponding to the first data based on the acquired first item information 1 and determine n which is the number of attributes corresponding to the second data based on the acquired second item information 2 and determine n which is the number of the determined attributes 1 and the n 2 Based on n and the n, determine a first parameter ε indicating the strength of privacy in the local differential privacy satisfied by the first data and the second data after noise is added 1 and a second parameter ε 2 and having a determination unit for determining The information processing apparatus according to claim 2.

5. At least one of the first generation unit and the second generation unit reduces the number of values that the data corresponding to at least one attribute among the plurality of data corresponding to each of the plurality of attributes included in the data group can take, and after reducing the number of values that the data can take, generates the noise addition query that adds the noise to each of the plurality of data, The information processing apparatus according to claim 1.

6. The first generation unit generates a first update query that is a query for updating the converted first data group by replacing the first data included in the first record included in the converted first data group with the first data included in other first records at a first ratio, The second generation unit generates a second update query that is a query for updating the converted second data group by replacing the second data included in the second record included in the converted second data group with the second data included in other second records at a second ratio, The transmission unit transmits the first update query generated by the first generation unit to the first device and transmits the second update query generated by the second generation unit to the second device, The information processing apparatus according to claim 1.

7. An update unit that updates the integrated data group by replacing the data included in the record included in the integrated data group with the data included in other records at a third ratio is provided, The information processing apparatus according to claim 1.

8. The first generation unit generates the first irreversible conversion query that adds random data to each of the plurality of data identification information included in the first data group and then performs irreversible conversion by the predetermined method. The second generation unit generates the second irreversible conversion query that adds the same random data as the random data added to the data identification information included in the first data group, which corresponds to the data identification information, to each of the plurality of data identification information included in the second data group and then performs irreversible conversion by the predetermined method. The information processing apparatus according to claim 1.

9. The integration unit further processes the integrated data group to generate statistical data, and performs correction to remove the noise using the predetermined probability used for adding the noise to the statistical data. The information processing apparatus according to claim 1.

10. The first data group and the second data group include specific data that can be used to identify newly added records. When regenerating the first irreversible conversion query, the first generation unit generates the first noise addition query that adds the noise to the first record newly added based on the specific data. When regenerating the second irreversible conversion query, the second generation unit generates the second noise addition query that adds the noise to the second record newly added based on the specific data. The information processing apparatus according to claim 1.

11. Based on the data identification information of each of the plurality of first records included in the first data group after conversion and the data identification information of each of the plurality of second records included in the second data group after conversion, the integration unit integrates the first data group after conversion and the second data group after conversion, excludes the data identification information, and generates the integrated data group. The information processing apparatus according to claim 1.

12. Executed by an information processing apparatus A first noise addition query, which is a query for adding noise to each of the plurality of first data included in the first data group such that noise is added to the plurality of data included in the integrated data group with a predetermined probability when the first data group and the second data group are integrated into an integrated data group, and a first irreversible conversion query, which is a query for irreversibly converting the plurality of data identification information included in the first data group by a predetermined method, are generated for each of the plurality of first data included in the first data group among a first data group including a plurality of first records associating data identification information for identifying data with the first data and a second data group including a plurality of second records associating the data identification information with the second data. A second noise addition query, which is a query for adding noise to each of the plurality of second data included in the second data group such that noise is added to the plurality of data included in the integrated data group with a predetermined probability when the first data group and the second data group are integrated into an integrated data group, and a second irreversible conversion query, which is a query for irreversibly converting the plurality of data identification information included in the second data group by a predetermined method, are generated for each of the plurality of second data included in the second data group. The generated first noise addition query and the first irreversible conversion query are transmitted to a first device corresponding to the provider of the first data group, and the generated second noise addition query and the second irreversible conversion query are transmitted to a second device corresponding to the provider of the second data group. A converted first data group including a plurality of first records associating the data identification information converted based on the first irreversible conversion query with the plurality of first data to which noise is added based on the first noise addition query, and a converted second data group including a plurality of second records associating the data identification information converted based on the second irreversible conversion query with the plurality of second data to which noise is added based on the second noise addition query are acquired. Based on the data identification information of each of the plurality of first records included in the acquired converted first data group and the data identification information of each of the plurality of second records included in the acquired converted second data group, the integrated data group obtained by integrating the converted first data group and the converted second data group is generated. An information processing method having the above steps.

Citation Information

Patent Citations

  • Database system, data coupling method, integrating server, data coupling program, database system sharing method and database system sharing program

    JP2018010424A

  • Database management system and database processing method

    JP2021056921A

  • Coordination server program, business operator server program, and data coordinated system

    JP2021117679A

  • Data Analytics Privacy Platform with Quantified Re-identification Risk

    JP2023543716A

  • Privacy-aware query management system

    US20170169253A1