Information processing device and information system
Patent Information
- Application Number
- JP2025527456
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-10-01
- Publication Date
- 2026-01-09
AI Technical Summary
Current methods for data utilization between multiple companies face challenges in selecting appropriate data items for linkage to enhance statistical information output while maintaining privacy, due to limitations in privacy budgets and the complexity of numerous data combinations.
An information processing device calculates the degree of association between data items and indices related to data utilization purposes, using machine learning predictive models and correlation coefficients to select relevant data items for aggregation patterns, and allocates privacy budgets to ensure differential privacy standards are met.
This approach allows for the appropriate selection of data items to improve the usefulness of statistical information output in data linkage, optimizing privacy budget allocation and ensuring the privacy of aggregated results.
Abstract
Description
Information processing device and information system
[0001] The present disclosure relates to an information processing device and an information system.
[0002] By utilizing data held by multiple companies, it is expected that value can be created that cannot be obtained with data held by a single company. One method for realizing data utilization between multiple companies is to output aggregated results without mutually disclosing the data held by multiple companies. In this case, privacy information in the output data can be protected by adding noise based on differential privacy standards to the aggregated results within a predetermined privacy budget. Regarding the above privacy budget, Non-Patent Document 1 below describes a technology related to allocating a privacy budget in a differential privacy use case among multiple companies.
[0003] "Optimal distribution of privacy budget in differential privacy," by Bkakria Anis et al., published October 16, 2018
[0004] However, when utilizing data between multiple companies, there has been no study or proposal as to which data items should be linked to improve the usefulness of the output statistical information.In addition, because the number of data combinations between multiple companies is enormous, it is practically difficult to aggregate all of these combinations due to privacy budget limitations.
[0005] Therefore, an object of the present disclosure is to appropriately select data items to be linked in order to contribute to improving the usefulness of statistical information output by data linkage when utilizing statistical data between multiple companies.
[0006] The information processing device disclosed herein is an information processing device that, when performing aggregation between its own device and a counterpart device that holds user data including a user ID and attribute information about the user, targeting an aggregation pattern that includes multiple combinations of data items included in the attribute information, adds noise based on differential privacy standards to the aggregation results within a privacy budget and performs aggregation targeting the aggregation pattern, and is equipped with a relevance calculation unit that calculates the relevance between an index related to the purpose of data utilization and the data items included in the attribute information, and a selection unit that selects data items to be included in the aggregation pattern based on the relevance calculated by the relevance calculation unit.
[0007] According to the present disclosure, when utilizing statistical data between multiple companies, data items to be linked can be appropriately selected to contribute to improving the usefulness of statistical information output by data linkage.
[0008] 9 is a functional block diagram of an information processing device in the first to fourth embodiments. FIG. 10 is a diagram showing example data in the first to third embodiments. FIG. 11 is a flow diagram showing processing executed in the first to fourth embodiments. FIG. 12 is a diagram showing example data in the fourth embodiment. FIG. 13 is a functional block diagram of an information system in the fifth to sixth embodiments. FIG. 14 is a diagram showing a configuration related to determining allocation of a privacy budget in the fifth embodiment. FIG. 15 is a diagram showing a configuration related to tallying processing in the fifth to sixth embodiments. FIG. 16 is a flow diagram showing processing executed in the fifth to sixth embodiments. FIG. 17 is a flow diagram showing processing related to determining allocation of a privacy budget in the fifth embodiment. FIG. 18 is a diagram for explaining tallying patterns. FIG. 19 is a diagram for explaining the processing of steps S1 to S8 of FIG. 9. FIG. 19 is a diagram for explaining data processing. FIG. 20 is a diagram for explaining irreversible ID conversion. FIG. 21 is a diagram for explaining encryption of an ID. (a) is a diagram for explaining encryption of attribute information, and (b) is a diagram for explaining transmission and reception of encrypted IDs. FIG. 22 is a diagram for explaining re-encryption of an ID. FIG. 23 is a diagram for explaining data matching. FIG. 24 is a diagram for explaining the processing of steps S9 to S12 of FIG. 9. FIG. 25 is a diagram for explaining the processing of steps S13 to S14 of FIG. 10 is a flow diagram showing the tallying process in the fifth and sixth embodiments. FIG. 11 is a diagram for explaining irreversible ID conversion. FIG. 12 is a diagram for explaining encryption of an ID. (a) is a diagram for explaining encryption of attribute information, and (b) is a diagram for explaining transmission and reception of encrypted IDs. FIG. 13 is a diagram for explaining re-encryption of an ID. FIG. 14 is a diagram for explaining data matching. FIG. 15 is a diagram for explaining tallying process. FIG. 16 is a diagram for explaining concealment process. FIG. 17 is a diagram for explaining decryption process. FIG. 18 is a diagram showing a configuration related to the allocation decision of a privacy budget in the sixth embodiment. FIG. 19 is a flow diagram showing the processing related to the allocation decision of a privacy budget in the sixth embodiment. FIG. 19 is a diagram for explaining a first allocation example of a privacy budget. FIG. 19 is a diagram for explaining a second allocation example of a privacy budget. FIG. 19 is a diagram for explaining a first example of output statistical information. FIG. 19 is a diagram for explaining a second example of output statistical information. FIG. 19 is a configuration diagram of an information system according to a modified example. FIG. 19 is a diagram showing an example of the hardware configuration of an information processing device.
[0009] Hereinafter, first to sixth embodiments of the present disclosure will be described in order with reference to the drawings. These first to sixth embodiments are broadly divided into first to fourth embodiments and fifth to sixth embodiments. The former (first to fourth embodiments) provide a detailed description of the process for selecting data items to be included in the aggregation pattern, while the latter (fifth and sixth embodiments) provide a detailed description of the privacy budget allocation and aggregation process for the selected data items.
[0010] To outline each embodiment, a first embodiment calculates the degree of association between an index (hereinafter referred to as "index") related to the purpose of data utilization and a data item based on a predictive model generated by supervised machine learning, and selects data items to be included in an aggregation pattern based on the degree of association. A second embodiment calculates the degree of association between an index and a data item based on a correlation coefficient between the index and the data item, and selects data items to be included in an aggregation pattern based on the degree of association. A third embodiment calculates the degree of association between an index and a category for each of a plurality of categories into which each data item is classified, and selects data items to be included in an aggregation pattern based on the number of categories with high degrees of association. A fourth embodiment calculates the degree of association between an index and a category for each of a plurality of categories into which each data item is classified, and selects categories with high degrees of association as data items to be included in an aggregation pattern. Furthermore, the fifth embodiment is an embodiment in which, in addition to selecting a data item, the allocation of the privacy budget to an aggregation pattern including the selected data item (including calculation of a base value) and aggregation processing is performed, and the sixth embodiment is an embodiment in which, in addition to selecting a data item, the allocation of the privacy budget to an aggregation pattern including the selected data item (including calculation of a base value) and aggregation processing is performed.
[0011] First Embodiment In the first embodiment, the purpose of data utilization is to increase sales of a predetermined target product (target product), and data sharing between companies (Company A and Company B) is assumed to be performed to increase sales of the target product. In this data sharing, when Company A and Company B perform aggregation for an aggregation pattern that includes multiple combinations of data items included in the attribute information of each user, noise based on differential privacy standards is added to the aggregation results within the privacy budget, and aggregation for the aggregation pattern is performed. The above-mentioned objectives and assumptions are also applicable to the second to sixth embodiments. Note that in the first to fourth embodiments, a detailed description is given of the process up to the selection of data items to be included in the aggregation pattern in Company A's information processing device 100 shown in FIG. 1 as a pre-stage process for data sharing.
[0012] 1 , an information processing device 100 according to a first embodiment includes, as characteristic components according to the present disclosure, an association degree calculation unit 101 that calculates an association degree between an index associated with a data utilization objective (increasing sales of a target product) and a data item included in attribute information, and a selection unit 102 that selects data items to be included in a counting pattern based on the calculated association degrees. In the first embodiment, a form will be described in which the association degree calculation unit 101 generates a supervised machine learning prediction model using the index as a target variable and data items included in the attribute information as explanatory variables, and calculates the association degree based on the obtained prediction model.
[0013] As shown in FIG. 2 , Company A's data includes a user ID for identifying a user and attribute information linked to the user ID. The attribute information includes "Product Purchase Status," which indicates whether the user purchased the target product, "Generation," which indicates the user's age group, "Residential Area," which indicates the area where the user lives, and "Number of Cohabitants," which indicates the number of people living in the home. Among these, "Product Purchase Status" is treated as an indicator related to the purpose of data utilization (increasing sales of the target product), and the selection unit 102 selects, from among the other data items ("Generation," "Residential Area," and "Number of Cohabitants"), data items highly related to "Product Purchase Status" as data items to be included in the aggregation pattern. In the data example of FIG. 2 , "Product Purchase Status" includes two categories, "Yes" and "No," and "Generation" includes seven categories, including teens and younger, 20s, 30s, 40s, 50s, 60s, and 70s and older. "Residence area" includes eight categories: Hokkaido, Tohoku, Kanto, Chubu, Kinki, Chugoku, Shikoku, and Kyushu, and "number of people living together" includes three categories: 1 person, 2 people, and 3 or more people. For ease of explanation, the above-mentioned "product purchase status," "age group," "residence area," and "number of people living together" may be referred to as "item A," "item B," "item C," and "item D," respectively. As will be explained in the fifth and sixth embodiments below, data on Company B, a business partner of Company A, includes a user ID and "gender" as attribute information linked to the user ID. "Gender" includes two categories: male and female, and may be referred to as "item E."
[0014] Next, the processing executed by the information processing device 100 will be described with reference to the flowchart of FIG. 3 . First, the relevance calculation unit 101 calculates the relevance between an index related to the purpose of data utilization and a data item included in the attribute information (step ST1). Specifically, the relevance calculation unit 101 generates a supervised machine learning prediction model using the index as the objective variable and the data item included in the attribute information as the explanatory variable, and calculates the relevance based on the obtained prediction model. In the data example of FIG. 2 , the relevance calculation unit 101 generates a prediction model that predicts the indicator "product purchase status" based on "age," "residential area," and "number of cohabitants." For example, the prediction model is generated using a decision tree-based model called LightGBM (Light Gradient Boosting Machine), with "age," "residential area," and "number of cohabitants" as explanatory variables (features) and "product purchase status" as the objective variable (teaching data).
[0015] The learning method for the prediction model here may be a known method disclosed in the paper "LightGBM: A Highly Efficient Gradient Boosting Decision Tree" by Guolin Ke et al., Advances in neural information processing systems 30 (2017). As a result of learning, the feature importance of each explanatory variable (feature) can be obtained, and these can be used as the relevance. Here, for example, the relevance w of "age" 年代 As "0.8", "relevance of residential area" w 居住地域 As "0.4", the relevance of "number of people living together" 同居人数 Here, we have shown an example of generating a prediction model using the above decision tree-based model, but it is also possible to generate a prediction model using various supervised machine learning methods that can output feature importance, such as linear regression, support vector machines, and random forests.
[0016] Next, the selection unit 102 selects data items to be included in the aggregation pattern based on the calculated relevance (step ST2 in FIG. 3). As an example, the selection unit 102 may select a predetermined number (k) of data items with the highest relevance. For example, if k=2, the top two data items with the highest relevance, "age group" and "number of people living together," are selected.
[0017] According to the first embodiment described above, a supervised machine learning prediction model is generated using an index related to the purpose of data utilization as the objective variable and data items included in the attribute information as explanatory variables, and data items to be included in the aggregation pattern can be more appropriately selected based on the relevance calculated based on the obtained prediction model.
[0018] [Second Embodiment] Hereinafter, as a second embodiment, an embodiment will be described in which the degree of association between an index and a data item related to the purpose of data utilization is calculated based on a correlation coefficient between the index and the data item, and data items to be included in an aggregation pattern are selected based on the degree of association. In the second embodiment, the configuration ( FIG. 1 ) and processing flow ( FIG. 3 ) of the information processing device 100 are the same as in the first embodiment, and therefore the description will focus on the details of the degree of association calculation process and selection process that are unique to the second embodiment.
[0019] The relevance calculation unit 101 calculates the correlation coefficients between "product purchase or not," which is an index related to the purpose of data utilization, and each of the data items "age group," "residential area," and "number of people living together," as the relevance (step ST1 in FIG. 3). Of the above, "age group" and "number of people living together" are quantitative variables, but "residential area" is a nominal scale. Therefore, here, the following correlation ratio (an example of a correlation coefficient), which calculates the correlation between all three variables on a nominal scale, is calculated as the relevance.
[0020] The correlation ratio for a data item is calculated by dividing the data item into a number of sections and dividing the data in section i (i is an integer between 1 and a) by n i Here, the average of "product purchase" for the entire data item is Let the average of "product purchase" in category i be Let the jth item in the ith division (j is 1 to n i(integer up to x) data (value indicating "product purchase"; 1: product purchase, 0: product not purchase) ij Then, the correlation ratio can be calculated by the following equation (1). In addition, if the data items for which the correlation coefficient is to be calculated are not nominal scales but are only quantitative variables (numeric data), the Pearson correlation coefficient (an example of a correlation coefficient) may be calculated as the degree of association using a known method.
[0021] In step ST1 of FIG. 3, for example, the relevance w 年代 As "0.8", "relevance of residential area" w 居住地域 As "0.4", the relevance of "number of people living together" 同居人数 Assuming that "0.6" is obtained as the correlation coefficient, in the next step ST2, the selection unit 102 may select a predetermined number (k) of data items with the highest relevance, as in the first embodiment. For example, if k=2, the top two data items with the highest relevance, "age group" and "number of people living together," are selected.
[0022] According to the second embodiment described above, it is possible to more appropriately select data items to be included in the aggregation pattern based on the relevance calculated based on the correlation coefficient between the index related to the purpose of data utilization and the data item.
[0023] [Third Embodiment] Hereinafter, as a third embodiment, an embodiment will be described in which the degree of association between an index and a category is calculated for each category into which each data item is classified, and data items to be included in an aggregation pattern are selected based on the number of categories with high degrees of association. In the third embodiment, the configuration ( FIG. 1 ) and processing flow ( FIG. 3 ) of the information processing device 100 are the same as in the first embodiment, and therefore the description will focus on the details of the degree of association calculation process and selection process that are unique to the third embodiment.
[0024] In the third embodiment, the relevance calculation unit 101 generates a supervised machine learning prediction model using an index related to the purpose of data utilization as the objective variable and data items included in the attribute information as explanatory variables, and calculates the relevance based on the obtained prediction model (step ST1 in FIG. 3 ). For example, in the data example of FIG. 2 , the relevance calculation unit 101 generates a linear regression model to predict "product purchase status," an index related to the purpose of data utilization, from the data items "age group," "area of residence," and "number of cohabitants." Specifically, the "age group," "area of residence," and "number of cohabitants" are each represented by a one-hot vector (a vector in which only one corresponding item is "1" and the others are "0"; hereinafter referred to as a "one-hot vector"), and the "product purchase status" is predicted using the following equation (2): Here, y is the product purchase status (1: purchased, 0: not purchased), x1 to x7 are the one-hot features of "age", and x8 to x 15 is the one-hot feature of "residential area", x 16 ~x 18 is the one-hot feature of the number of people living together, and w j (j: 1 to 18) represents the weight of each feature x, and b represents the intercept.
[0025] The learning method is to use past data to determine the weights w of each feature x in the model so that the difference (error) between the prediction and the actual results shown in the following equation (3) is minimized. j Learn (j: 1-18). The weights that minimize the above error can be found using the least squares method. Note that although a linear regression model is used as an example here, a prediction model may also be generated using various supervised machine learning methods that can output feature importance, such as support vector machines, decision trees, and random forests.
[0026] In step ST1 of Figure 3 above, for example, the weight w of "Age: 60s" 6 The weight of "0.8" and "Age: 50s" 5 As "0.6", the weight of "Residential area: Kanto" 10 The weight of "0.4" is "Number of people living together: 1" 16 The weight of "0.4" is "Number of people living together: 2"17 As "0.4", weights other than the above w j The weights w j It can be said that the feature (category) with a larger value is more closely related to "whether or not a product was purchased," which is an indicator related to the purpose of data utilization.
[0027] Therefore, in the next step ST2, the selection unit 102 identifies feature quantities (categories) with high relevance from the above feature quantities (categories) based on a predetermined criterion, and selects data items to be included in the aggregation pattern based on the number of feature quantities (categories) with high relevance for each data item. As an example, the selection unit 102 selects a weight w j The top k features (categories) are identified in descending order, and their occurrence count is counted for each data item (before one-hot vectorization). Here, if k=5, the data item "age" is counted twice, the data item "place of residence" is counted once, and the data item "number of people living together" is counted twice, so "age" and "number of people living together," which are the data items with the highest counts, are selected as the data items to be included in the aggregation pattern.
[0028] According to the third embodiment described above, the relevance of each data item is calculated for each of the multiple categories into which the data items are classified, and the data items to be included in the aggregation pattern are selected based on the number of categories with high relevance counted using a specified method, thereby enabling more appropriate selection of the data items.
[0029] Regarding the above frequency of occurrence, data items with a larger number of vector dimensions (i.e., the number of "categories" included in the data item) when one-hot vectorized are more likely to be selected. Therefore, as a modification of the third embodiment, data items to be included in the aggregation pattern may be selected based on the ratio obtained by dividing the above count for each data item by the number of vector dimensions. In the above example, the data item "age" has a vector dimension of 7 and a count of 2, so the ratio is 2 / 7 = 0.285. Similarly, the ratio for the data item "residential area" is 1 / 8 = 0.125, and the ratio for the data item "number of people living together" is 2 / 3 = 0.666. Then, "number of people living together," which has the highest ratio, is selected as the data item to be included in the aggregation pattern. In this way, data items to be included in the aggregation pattern can be more appropriately selected, taking into account the fact that the number of vector dimensions (number of categories) differs for each data item.
[0030] Furthermore, in the third embodiment, similar to the first embodiment, a supervised machine learning prediction model is generated and the relevance is calculated based on the obtained prediction model. However, as a form of calculating the relevance, a form in which the correlation coefficient between the index and the data item (for example, the aforementioned correlation ratio) is calculated as the relevance, as in the second embodiment, may also be adopted.
[0031] [Fourth Embodiment] Hereinafter, as a fourth embodiment, an embodiment will be described in which the degree of association between an index and a category is calculated for each category into which data items are classified, and a category with a high degree of association is selected as a data item to be included in an aggregation pattern. In the fourth embodiment, the configuration ( FIG. 1 ) and processing flow ( FIG. 3 ) of the information processing device 100 are the same as those in the first embodiment, and therefore the description will focus on the details of the degree of association calculation process and selection process that are unique to the fourth embodiment.
[0032] In the data example of the fourth embodiment shown in Figure 4, "product purchase" includes two categories: "yes" and "no," and "age" includes seven categories: teens and under, 20s, 30s, 40s, 50s, 60s, and 70s and older. "Region of residence" includes eight categories: Hokkaido, Tohoku, Kanto, Chubu, Kinki, Chugoku, Shikoku, and Kyushu. Furthermore, "gender" in Company B's data includes two categories: male and female.
[0033] In the fourth embodiment, the relevance calculation unit 101 generates a supervised machine learning prediction model using an index related to the purpose of data utilization as the objective variable and data items included in the attribute information as the explanatory variables, and calculates the relevance based on the obtained prediction model (step ST1 in Figure 3).
[0034] In the data example of Fig. 4, the relevance calculation unit 101 generates a prediction model using a linear regression model to predict "product purchase or not," an indicator related to the purpose of data utilization, from the data items "age" and "residential area." Specifically, "age" and "residential area" are each represented by a one-hot vector, and product purchase is predicted using the following equation (4). Here, y is the product purchase status (1: purchased, 0: not purchased), x1 to x7 are the one-hot features of "age", and x8 to x 15 is the one-hot feature of "residential area", w j (j: 1 to 15) represents the weight of each feature x, and b represents the intercept.
[0035] The learning method is to use past data to determine the weights w of each feature of the model so that the difference (error) between the prediction and the actual results shown in the following equation (5) is minimized. j Learn (j: 1-15). The weights that minimize the above error can be found using the least squares method. Note that although a linear regression model is used as an example here, a prediction model may also be generated using various supervised machine learning methods that can output feature importance, such as support vector machines, decision trees, and random forests.
[0036] In step ST1 of Figure 3 above, for example, the weight w of "Age: 60s" 6 The weight of "0.8" and "Age: 50s" 5As "0.6", the weight of "Residential area: Kanto" 10 As "0.4", weights other than the above w j The weights w j It can be said that the feature (category) with a larger value is more closely related to "whether or not a product was purchased," which is an indicator related to the purpose of data utilization.
[0037] Therefore, in the next step ST2, the selection unit 102 identifies feature quantities (categories) with high relevance from the above feature quantities (categories) based on a predetermined criterion, and selects the identified feature quantities (categories) with high relevance as targets to be included in the aggregation pattern. For example, the selection unit 102 selects a weight w j The top k feature quantities (classifications) are identified in descending order of magnitude, and the identified top k feature quantities (classifications) are selected as targets to be included in the aggregation pattern. Here, if k=3, the weight w j The top three feature values (categories) in descending order of magnitude, "age: 60s," "age: 50s," and "area of residence: Kanto," are selected as targets to be included in the aggregation patterns.
[0038] According to the fourth embodiment described above, the relevance of each data item is calculated for each of the multiple categories into which the data items are classified, and the data items (here, feature quantities (categories)) to be included in the aggregation pattern are selected based on the number of categories with high relevance counted using a predetermined method, thereby enabling more appropriate selection of the data items.
[0039] Note that, like the first embodiment, the fourth embodiment generates a supervised machine learning prediction model and calculates the relevance based on the obtained prediction model. However, as a form of calculating the relevance, a form in which the correlation coefficient (for example, the aforementioned correlation ratio) between the index and the feature (category) is calculated as the relevance, as in the second embodiment, may also be adopted.
[0040] [Fifth Embodiment] Below, as the fifth embodiment, in addition to selecting data items as in the first to fourth embodiments, an embodiment will be described in which allocation of the privacy budget to aggregation patterns including the selected data items (including calculation of the base value) and aggregation processing are performed.
[0041] 5, the system configuration of the fifth embodiment includes an information processing device 100 of company A and an information processing device 200 of company B, and these information processing devices 100 and 200 cooperate to execute the tallying process. Of these, the information processing device 100 of company A includes the relevance calculation unit 101 and the selection unit 102 described in the first to fourth embodiments.
[0042] Furthermore, the information processing device 100 of Company A further includes, as a characteristic configuration of this embodiment, a determination unit 103 that determines the allocation of the privacy budget to each combination of data items including the data item selected by the selection unit 102, and an aggregation execution unit 104 that adds noise based on differential privacy standards to the aggregation results within the privacy budget allocated based on the determined allocation for each of the above combinations, and performs aggregation targeting the aggregation pattern.
[0043] On the other hand, Company B's information processing device 200, details of which will be described later, comprises an information providing unit 201 that cooperates with Company A's decision making unit 103 to provide information necessary for the decision making unit 103 to decide on the allocation of the privacy budget, and a tally output unit 202 that cooperates with Company A's tally execution unit 104 to provide information necessary for the tally execution unit 104 to perform tallying, and decrypts and outputs the tally results.
[0044] That is, in FIG. 5 , the components indicated by arrow X (determination unit 103 and information providing unit 201) have the function of determining the allocation of the privacy budget, and the components indicated by arrow Y (tallying unit 104 and tally output unit 202) have the function of performing tallying for tallying patterns and outputting the tallying results. As shown in FIG. 5 , the following example shows the relevance calculation unit 101, the components indicated by arrow X, and the components indicated by arrow Y receiving the latest user data to be processed and processing them. However, this is not required. User data once input to information processing devices 100 and 200 may be stored in a storage unit (not shown) and read from the storage unit during subsequent processing. Below, the functions of the components indicated by arrow X and the components indicated by arrow Y will be described in order.
[0045] Fig. 6 shows a functional block configuration of the configuration indicated by arrow X in Fig. 5. As shown in Fig. 6, the determination unit 103 includes a data input unit 11, a pre-execution data generation unit 12, a basic value calculation unit 13, a basic tally information acquisition unit 14A, an estimation unit 15, an evaluation unit 16, and an allocation determination unit 17, while the information provision unit 201 includes a data input unit 11, a pre-execution data generation unit 12, and a basic tally information provision unit 14B. The functions of each unit will be outlined below, with details to be provided later with reference to the flowchart in Fig. 9.
[0046] The data input unit 11 is a functional unit that accepts input data from outside, and is provided in common with the determination unit 103 of Company A and the information providing unit 201 of Company B. However, although the input data for both Company A and Company B is user data including a user ID and attribute information related to the user, the content of the user data differs between the companies. Specific examples of user data will be described later.
[0047] The pre-execution data generation unit 12 is a functional unit that generates data for pre-execution (referred to as "pre-execution data") when a trial tabulation for calculating basic values such as the data matching rate between the two companies, which will be described later, is performed in advance before the actual tabulation, and is provided in common with the determination unit 103 of Company A and the information provision unit 201 of Company B. The pre-execution data generation unit 12 includes a data processing unit 12A that processes input data into small-scale data with a narrowed-down set of data items for calculating the basic values, an anonymization processing unit 12B that performs an anonymization process for protecting the privacy of attribute information in the user data, an ID irreversible conversion unit 12C that performs an irreversible conversion process to a user ID in the user data, an encryption unit 12D that encrypts the user data, and a data transmission / reception unit 12E that transmits and receives the encrypted user data.
[0048] The base value calculation unit 13 is a functional unit that calculates a base value from the pre-execution data A generated by the pre-execution data generation unit 12 in the decision unit 103 of Company A and the pre-execution data B generated by the pre-execution data generation unit 12 in the information provision unit 201 of Company B, and is provided in the decision unit 103 of Company A. The base value calculation unit 13 includes a data matching unit 13A that matches the pre-execution data A with the pre-execution data B, a counting processing unit 13B that performs counting processing based on the matching results, a concealment processing unit 13C that performs concealment processing on the counting processing results, and a calculation unit 13D that calculates a base value from the counting processing results after the concealment processing.
[0049] The basic aggregation information providing unit 14B is a functional unit provided in the information providing unit 201 of Company B, which calculates the sample size and item-specific proportions, which are the basic aggregation information of Company B's user data, and provides them to the determination unit 103 of Company A.
[0050] The basic summary information acquisition unit 14A is a functional unit provided in the determination unit 103 of Company A, which calculates the sample size and item-specific ratio, which are the basic summary information of Company A's user data, and receives the sample size and item-specific ratio, which are the basic summary information of Company B's user data, calculated by Company B's information provision unit 201, thereby acquiring the basic summary information of both Company A and Company B.
[0051] The estimation unit 15 is a functional unit that estimates the sample size of the aggregation results based on the basic aggregation information of both Company A and Company B acquired by the basic aggregation information acquisition unit 14A and the basic value calculated by the basic value calculation unit 13, and is provided in the determination unit 103 of Company A.
[0052] The evaluation unit 16 is a functional unit that quantitatively evaluates the influence of noise on each combination included in the aggregation pattern based on the sample size of the estimated aggregation results, and is provided in the determination unit 103 of Company A.
[0053] The allocation determination unit 17 is a functional unit that determines the allocation of the privacy budget to each combination based on the evaluation results for each combination obtained by the evaluation unit 16, and is provided in the determination unit 103 of Company A.
[0054] Fig. 7 shows a functional block configuration of the configuration indicated by arrow Y in Fig. 5. As shown in Fig. 7, the tally execution unit 104 includes an anonymization processing unit 21, an ID irreversible conversion unit 22, an encryption unit 23, a data transmission / reception unit 24, a data matching unit 25, a tally processing unit 26, and a confidentiality processing unit 27, while the tally output unit 202 includes an anonymization processing unit 21, an ID irreversible conversion unit 22, an encryption unit 23, a data transmission / reception unit 24, and a decryption unit 28. The function of each unit will be described below.
[0055] The anonymization processing unit 21 is a functional unit that executes processing for privacy protection of attribute information on user data held by its own device prior to encryption of the user ID, and is provided in common with the tally execution unit 104 of company A and the tally output unit 202 of company B. Note that, as the privacy protection, for example, one or more of k-anonymization, l-diversity, and t-approximation are adopted, and an example of k-anonymization among the above will be described later.
[0056] The ID irreversible conversion unit 22 is a functional unit that performs irreversible conversion processing to a user ID on user data held by its own device prior to encrypting the user ID, and is provided in common with the tally execution unit 104 of Company A and the tally output unit 202 of Company B. Note that the above-mentioned irreversible conversion processing includes hashing processing, and after executing the hashing processing to the user ID, the ID irreversible conversion unit 22 discards the salt used in the hashing processing.
[0057] The encryption unit 23 provided in Company A's aggregation execution unit 104 is a functional unit that has the function of generating encrypted user data for the user data to be aggregated by encrypting the user ID in the user data to be aggregated based on the self-encryption key and keyed one-way commutative calculation held by its own device, and includes an ID encryption unit 23A that performs this functional operation.
[0058] In contrast, the encryption unit 23 provided in Company B's tally output unit 202 is a functional unit that encrypts the user ID in the user data to be tallied based on a self-encryption key and a keyed one-way commutative operation held by its own device, and encrypts the attribute information in the user data to be tallied using a homomorphic encryption method that enables tallying processing, thereby generating encrypted user data for the user data to be tallied, and includes an ID encryption unit 23A that has the function of encrypting the user ID, and an attribute information encryption unit 23B that has the function of encrypting the attribute information. Note that the order of execution of the encryption of the user ID and the encryption of the attribute information may be arbitrary and may be any order.
[0059] The data transmission / reception unit 24 is a functional unit that transmits and receives encrypted user data, and is provided in common to the tally execution unit 104 of company A and the tally output unit 202 of company B.
[0060] The data matching unit 25 is a functional unit that matches the encrypted user data of Company A generated by the encryption unit 23 of the aggregation execution unit 104 of Company A with the encrypted user data of Company B based on the user ID corresponding portion identified based on predetermined structural information of the user data, and is provided in Company A's aggregation execution unit 104.
[0061] The tallying processing unit 26 is a functional unit that generates encrypted tally data for a target user by counting the number of encrypted user data whose user ID corresponding portions match as a result of matching by the data matching unit 25, and is provided in the tallying execution unit 104 of Company A. The counting method will be described later.
[0062] The confidentiality processing unit 27 is a functional unit that performs confidentiality processing on the encrypted aggregated data generated by the aggregation processing unit 26 and generates encrypted statistical information, and is provided in Company A's aggregation execution unit 104.
[0063] The decryption unit 28 is a functional unit that decrypts the encrypted statistical information transmitted from the aggregation execution unit 104 of Company A based on a decryption method corresponding to the encryption by the attribute information encryption unit 23B, and outputs the obtained statistical information to an external device, and is provided in the aggregation output unit 202 of Company B.
[0064] (5-2: Processing Executed in Fifth Embodiment) (5-2-1: Overview (Steps ST1 to ST4 in FIG. 8) of the Overall Process) The processing executed in the fifth embodiment will be described below with reference to FIGS. 8 to 28. As shown in FIG. 8, in the information processing device 100 of Company A, the relevance calculation unit 101 calculates the relevance between an index related to the purpose of data utilization and a data item (step ST1), and the selection unit 102 selects data items to be included in the aggregation pattern based on the calculated relevance (step ST2). In these steps ST1 to ST2, any of the processing in the first to fourth embodiments described above may be executed.
[0065] After the data items to be included in the tallying pattern are selected, the determination unit 103 in the information processing device 100 of company A and the information providing unit 201 in the information processing device 200 of company B cooperate to determine the allocation of the privacy budget as follows (step ST3): Then, the tallying unit 104 in the information processing device 100 of company A and the tally output unit 202 in the information processing device 200 of company B cooperate to execute the tallying process for the tallying pattern (described later) (step ST4).
[0066] (5-2-2: Details of Step ST3 in FIG. 8 (Determining Allocation of Privacy Budget)) Next, details of Step ST3 in FIG. 8 (Determining Allocation of Privacy Budget) will be described with reference to FIGS.
[0067] First, in each of Company A's determination unit 103 and Company B's information provision unit 201, user data to be tallied is input to data input unit 11 (steps S1 and S2 in FIG. 9). Here, the data for Company A input to determination unit 103 is the data for Company A (user ID and items A to D) shown in FIG. 2, and the data for Company B input to information provision unit 201 is the data for Company B (user ID and item E) shown in FIG. 2. In addition, two tally patterns are assumed in this embodiment: tally (1) for the combination of items A, B, and E (hereinafter referred to as "item A x B x E") shown in FIG. 10, and tally (2) for the combination of items A, D, and E (hereinafter referred to as "item A x D x E").
[0068] Next, in each of the determination unit 103 of Company A and the information providing unit 201 of Company B, the pre-execution data generation unit 12 generates the pre-execution data described above (steps S3 and S4 in FIG. 9 ). As shown in FIG. 9 , the pre-execution data generation process includes multiple processing steps, which are roughly divided into data processing (steps S3A and S4A) to generate small-scale data for calculating the base value, and pre-processing (steps S3B to S3G, S4B to S4H) to perform data matching (step S5) in a de-identified (anonymized or hashed) and encrypted state.
[0069] For the former, in the determination unit 103 of Company A and the information provision unit 201 of Company B, the data processing unit 12A processes the input data into small-scale data for calculating the base value (steps S3A and S4A). Here, for example, "item E," which is common to the totals for the two combinations of the totalization patterns (item A x B x E and item A x D x E) in this embodiment, is considered to be the main item, and the data is processed into small-scale data with attribute information limited to item E only. As a result, as shown in Figures 11 and 12, Company A's data is processed into only the user ID, and Company B's data is processed into the user ID and item E, respectively.
[0070] The latter "pre-processing" includes anonymization processing, irreversible ID conversion, ID encryption, etc. In the anonymization processing, the anonymization processing unit 12B in each of the determination unit 103 of company A and the information providing unit 201 of company B performs anonymization processing on the user data held by the respective devices as a process for protecting the privacy of attribute information (steps S3B and S4B in FIG. 9 ). Here, k-anonymization is exemplified, which converts user data so that k or more pieces of user data with the same attribute information exist within the target user data (satisfying k-anonymity), thereby reducing the probability of identifying an individual to 1 / k or less. If there are fewer than k pieces of user data with the same attribute information in the attribute information of company B's user data, the attribute information of user data with similar attribute information is converted so that there are k or more pieces of user data.
[0071] 13, in the determination unit 103 of company A and the information providing unit 201 of company B, the ID irreversible conversion unit 12C performs an irreversible conversion process to a user ID on the user data held by each device (steps S3C and S4C). Specifically, the ID irreversible conversion unit 12C performs a hashing process on the user ID, and then discards the salt used in the hashing process.
[0072] 14, in the determination unit 103 of Company A, the encryption unit 12D encrypts the non-identifiable hash (the portion of the user data corresponding to the user ID) with a secret key a prepared in advance to obtain ID-encrypted data for Company A (step S3D). Similarly, in the information providing unit 201 of Company B, the encryption unit 12D encrypts the non-identifiable hash (the portion of the user data corresponding to the user ID) with a secret key b prepared in advance to obtain ID-encrypted data for Company B (step S4D).
[0073] 15(a), in the information providing unit 201 of company B, the encryption unit 12D encrypts the attribute information in the ID encrypted data of company B obtained in step S4D with a secret key B prepared in advance by a homomorphic encryption method that enables aggregation processing, thereby generating encrypted user data of company B (step S4E). Note that each piece of attribute information in the generated encrypted user data is configured by binary values in a predetermined format so that aggregation processing can be performed in the encrypted state.
[0074] Next, as shown in FIG. 15(b), in Company A's determination unit 103, the data transmission / reception unit 12E transmits Company A's encrypted ID contained in Company A's ID encrypted data obtained in step S3D to the data transmission / reception unit 12E of Company B's information provision unit 201 (step S3E), and in Company B's information provision unit 201, the data transmission / reception unit 12E transmits Company B's encrypted user data obtained in step S4E (i.e., data including Company B's encrypted ID and encrypted attribute information) to the data transmission / reception unit 12E of Company A's determination unit 103 (step S4F).
[0075] 16, in the determination unit 103 of Company A, the encryption unit 12D re-encrypts the encrypted ID of Company B, which is included in the encrypted user data of Company B transmitted from the information providing unit 201 of Company B, with the secret key a, to obtain the encrypted ID of Company B encrypted with both the secret keys a and b (step S3F). Similarly, in the information providing unit 201 of Company B, the encryption unit 12D re-encrypts the encrypted ID of Company A transmitted from the determination unit 103 of Company A with the secret key b, to obtain the encrypted ID of Company A encrypted with both the secret keys a and b (step S4G). Then, in the information providing unit 201 of Company B, the data transmission / reception unit 12E transmits the "encrypted ID of Company A encrypted with both the secret keys a and b" obtained in step S4G to the determination unit 103 of Company A (step S4H), and the data transmission / reception unit 12E of the determination unit 103 of Company A receives the encrypted ID of Company A (step S3G).
[0076] Next, as shown in Fig. 17, in the determination unit 103 of Company A, the data matching unit 13A replaces the "encrypted ID encrypted with private key a" obtained by the ID encryption in step S3D with the "encrypted ID of Company A encrypted with both private keys a and b" received in step S3G, thereby obtaining the encrypted ID of Company A shown in the upper left of Fig. 17. Furthermore, the data matching unit 13A replaces the "encrypted ID of Company B encrypted with private key b" in the "encrypted user data of Company B" received in step S3E with the "encrypted ID of Company B encrypted with both private keys a and b" obtained by re-encryption in step S3F, thereby obtaining the encrypted user data of Company B shown in the upper right of Fig. 17. Then, the data matching unit 13A matches the encrypted user data of Company A with the encrypted user data of Company B using the "encrypted ID of Company A encrypted with both private keys a and b" and the "encrypted ID of Company B encrypted with both private keys a and b" as keys (step S5). Here, if "Company A's encrypted ID encrypted with both private keys a and b" and "Company B's encrypted ID encrypted with both private keys a and b" match, the attribute information in Company A's encrypted user data and the encrypted attribute information in Company B's encrypted user data are combined into a single record in the encrypted matching data. After the combination, "Company A's encrypted ID encrypted with both private keys a and b" and "Company B's encrypted ID encrypted with both private keys a and b" are deleted, resulting in encrypted matching data. However, in this embodiment, the attribute information in Company A's user data is not subject to pre-execution processing. Therefore, as shown in the lower part of Figure 17, the encrypted matching data is composed of Company B's encrypted attribute information (encrypted item E). As described above, the encrypted attribute information is composed of binary values in a predetermined format and can be aggregated in an encrypted state. Of course, because Company B's attribute information is encrypted, Company A cannot know its contents. Therefore, the encrypted matching data is generated without revealing to Company A the contents of the attribute information that should be kept confidential from Company B.
[0077] 9 and 11 , the data matching unit 13A matches the pre-execution data A with the pre-execution data B in the above-described manner, and the tallying unit 13B counts the number of records for "male" and "female" in item E for the matched records to obtain tallying results for "male" and "female" (step S6). As described above, the contents of item E, which is "Company B's encrypted attribute information" in the encrypted matched data, are encrypted, and therefore the determination unit 103 for Company A cannot know the contents. However, as described above, the "Company B's encrypted attribute information (item E)" in the encrypted matched data is composed of predetermined binary values. Therefore, in the encrypted state, it is possible to calculate the total number of records in the binary value format corresponding to "male" and the total number of records in the binary value format corresponding to "female."
[0078] The anonymization processing unit 13C then performs the following anonymization processing on the results of the above aggregation processing (step S7). For example, as shown in FIG. 11, the aggregation processing result for item E, in which "Male" is Num1 and "Female" is Num2, is targeted. Adding noise equivalent to consuming a privacy budget of 0.1 results in the anonymized aggregation processing result, as shown on the right side of FIG. 11, in which "Male" is Num1' and "Female" is Num2'. Furthermore, the calculation unit 13D calculates a base value from the aggregation processing result after the anonymization processing (step S8). Here, for example, a "data matching rate" and a "data bias" are calculated as the base values. The "data matching rate" refers to the percentage of records that were successfully matched out of the total number of records that were the subject of data matching. For example, a data matching rate of 40% is calculated. "Data bias" refers to the difference between the item-specific percentages for a given item derived from a single company's data and the item-specific percentages for that given item derived from data matched between companies. Data bias exists for each item for which an item-specific percentage is calculated. However, due to privacy budget constraints, it is difficult to calculate data bias for all items in advance. Therefore, data bias is calculated for item E, which is the main item for aggregation. For item E, a 45% item-specific percentage (the proportion of "males") is calculated by basic aggregation for Company B alone in step S10 of Figure 9 (described below), which is performed in parallel with steps S3 to S8 above. Meanwhile, as shown in Figure 11, the 45% item-specific percentage (the proportion of "males") derived from data matched between Companies A and B results in a data bias of "0%" being calculated as the base value.
[0079] Returning to FIG. 9 , in the determination unit 103 of Company A, the basic tabulation information acquisition unit 14A calculates the sample size and item-specific percentages, which are basic tabulation information, for Company A's data (step S9). For example, for tabulation (1) of items A x B x E, as shown in FIG. 18 , the sample sizes and percentages for each of "Yes x Teens and Under," "No x Teens and Under," ..., and "No x Age 70 and Over" are calculated for "Item A (Whether or Not a Product Purchased) x Item B (Age Group)" in Company A's data. Furthermore, for tabulation (2) of items A x D x E, the sample sizes and percentages for each of "Yes x 1 person," "No x 1 person," ..., and "No x 3 or more people" are calculated for "Item A (Whether or Not a Product Purchased) x Item D (Number of Cohabitants)" in Company A's data.
[0080] Similarly, in the information provider 201 of Company B, the basic tabulation information provider 14B calculates the sample size and item-specific ratios, which are basic tabulation information, for Company B's data, and provides them to the determination unit 103 of Company A (step S10). For example, for tabulation (1) of items A x B x E and tabulation (2) of items A x D x E, the sample sizes and ratios of males and females for item E (gender) in Company B's data are calculated, as shown in Figure 18.
[0081] Then, the basic summary information acquisition unit 14A of Company A's decision unit 103 acquires the basic summary information of both Company A and Company B by receiving the basic summary information (sample size and item-specific proportion for item E (gender)) provided (sent) by the basic summary information provision unit 14B of Company B's information provision unit 201 (step S11).
[0082] Next, the estimation unit 15 estimates the sample size of the tabulation results based on the basic tabulation information of Company A and Company B acquired by the basic tabulation information acquisition unit 14A and the basic values calculated by the basic value calculation unit 13, for example, as follows (step S12). For "item A×B×E" in the tabulation pattern (item A×B×E, item A×D×E), the sample size of the tabulation results is estimated by applying the following formula (6) as shown in FIG. 18 : Sample size ni (estimated value) = Sample size by age group and whether or not purchase was made in Company A's data × Proportion by gender in Company B's data × Data matching rate (6).
[0083] For example, for "item A x B x E," the sample size estimate n1 for the aggregated results is obtained by multiplying the sample size Sa1 of Company A's data with purchases and those under 10 years of age by the 45% male percentage of Company B's data and the data matching rate of 40%. Similarly, the above estimate n2 is obtained for those without purchases, those under 10 years of age, and men; the above estimate n3 is obtained for those with purchases, those under 10 years of age, and women; ..., the above estimate n28 is obtained for those without purchases, those over 70 years of age, and women. Then, the median of the 28 estimates n1 to n28 is obtained, m1, which is midway between the 14th largest value n14 and the 15th largest value n15. Similarly, for "item A x D x E," the median sample size estimate m2 for the aggregated results is obtained.
[0084] In this embodiment, an example has been shown in which the "median" of the sample size estimate of the aggregation results has been used, but the average, minimum value, etc. may also be used. Also, in this embodiment, as described above, the "data bias" of the basic values calculated in step S8 of Fig. 9 was 0%, but if there is data bias for a major item (item E in this case), it is desirable to fine-tune (correct) the "gender ratio of Company B's data" in equation (6) for calculating the above sample size ni (estimated value) by the amount of the "data bias" so that it approaches the gender ratio in the cross-checked data.
[0085] 9 , in the next step S13, the evaluation unit 16 quantitatively evaluates the influence of noise on each combination included in the aggregation pattern based on the estimated sample size of the aggregation result (here, the median of the estimated sample size values) as follows: In this embodiment, the variance of the amount of change in the sample size n is used as an evaluation index for the influence of noise. An example using the following will be explained. The variance of noise z is The value varies depending on the probability distribution used to generate the noise. When the variance of the noise z is substituted into the above equation (7), the following is obtained. Here, the allocation of the privacy budget for the summaries (1) and (2) is defined as ε a1 , ε a2By using the above formula (9), the following evaluation indices for the noise influence are obtained for the summaries (1) and (2), as shown in the table on the left side of FIG. 19: It is not essential to use the variance of the change in the sample size n as the evaluation index, and different indices for measuring the accuracy of data (for example, sampling error rate, statistical power, etc.) may be used depending on the purpose of aggregation.
[0086] In the next step S14, the allocation determination unit 17 determines the allocation of the privacy budget to each combination based on the evaluation result for each combination obtained by the evaluation unit 16, as follows. Here, for example, in order to equalize the influence of noise, the allocation determination unit 17 weights the remaining privacy budget ε by a function inversely proportional to the magnitude of the influence of noise. total Specifically, as shown on the right side of Figure 12, The remaining privacy budget ε total "0.9" is the budget allocation ε for each aggregation a1 , ε a2 In this case, the influence of noise z in the aggregations (1) and (2) is the same (V a1 =V a2 ).
[0087] 8 (determining the allocation of the privacy budget) quantitatively evaluates the impact of noise on each combination included in the aggregation pattern, and then determines the allocation of the privacy budget to each combination based on the evaluation results for each combination obtained in the evaluation, thereby optimizing the allocation of the privacy budget. Furthermore, by performing a trial aggregation in advance using the actually entered user data of companies A and B, basic values such as the data matching rate can be appropriately calculated based on the actual input data and used in the aggregation process, thereby contributing to optimizing the aggregation process.
[0088] (5-2-3: Details of Step ST4 (Counting Process) in FIG. 8) Next, details of Step ST4 (counting process) in FIG. 8 will be described with reference to FIGS. 7 and 20 to 28.
[0089] First, in each of the tallying execution unit 104 of Company A and the tallying output unit 202 of Company B shown in Fig. 7, user data to be tallied is input to the anonymization processing unit 21 (steps A1 and B1 in Fig. 20). The user data of Company A and Company B input here is the user data shown in Fig. 2 above.
[0090] Next, in each of the tallying unit 104 of Company A and the tallying output unit 202 of Company B, the anonymization processing unit 21 performs an anonymization process on the user data of each organization as a process for protecting the privacy of attribute information (steps A2 and B2). Here, as an example of the anonymization process, k-anonymization is performed, which converts the user data so that k or more pieces of user data with the same attribute information exist within the target user data (satisfying k-anonymity), thereby reducing the probability of identifying an individual to 1 / k or less. However, if k-anonymization is deemed unnecessary, the processes of steps A2 and B2 are not essential.
[0091] 21, in the tally execution unit 104 of Company A and the tally output unit 202 of Company B, the ID irreversible conversion unit 22 performs an irreversible conversion process to a user ID on the user data of each organization (steps A3 and B3). Specifically, the ID irreversible conversion unit 22 performs a hashing process on the user ID, and then discards the salt used in the hashing process.
[0092] 22, in the tally execution unit 104 of Company A, the ID encryption unit 23A encrypts the de-identified hash (the portion of the user data corresponding to the user ID) with a secret key a prepared in advance, thereby obtaining the ID-encrypted data of Company A (step A4). Similarly, in the tally output unit 202 of Company B, the ID encryption unit 23A encrypts the de-identified hash (the portion of the user data corresponding to the user ID) with a secret key b prepared in advance, thereby obtaining the ID-encrypted data of Company B (step B4).
[0093] Next, as shown in Figure 23(a), in Company B's aggregation output unit 202, the attribute information encryption unit 23B encrypts the attribute information in Company B's ID encrypted data obtained in step B4 using a homomorphic encryption method capable of aggregation processing with a pre-prepared secret key B, thereby generating Company B's encrypted user data (step B5).
[0094] 23(b), the data transmitter / receiver 24 in the tally execution unit 104 of Company A transmits the encrypted ID of Company A included in the encrypted ID data of Company A obtained in step A4 to the data transmitter / receiver 24 of the tally output unit 202 of Company B (step A5). Also, the data transmitter / receiver 24 in the tally output unit 202 of Company B transmits the encrypted user data of Company B obtained in step B5 (i.e., data including the encrypted ID and encrypted attribute information of Company B) to the data transmitter / receiver 24 of the tally execution unit 104 of Company A (step B6).
[0095] 24, in Company A's tally execution unit 104, the ID encryption unit 23A re-encrypts Company B's encrypted ID, included in Company B's encrypted user data transmitted in step B6, with private key a, to obtain Company B's encrypted ID encrypted with both private keys a and b (step A6). Similarly, in Company B's tally output unit 202, the ID encryption unit 23A re-encrypts Company A's encrypted ID transmitted in step A5 with private key b, to obtain Company A's encrypted ID encrypted with both private keys a and b (step B7). Then, in Company B's tally output unit 202, the data transmission / reception unit 24 transmits the "Company A's encrypted ID encrypted with both private keys a and b" obtained in step B7 to Company A's tally execution unit 104 (step B8), and the data transmission / reception unit 24 of Company A's tally execution unit 104 receives Company A's encrypted ID (step A7).
[0096] Next, as shown in Figure 25, in the tallying execution unit 104 of Company A, the data matching unit 25 replaces the "encrypted ID encrypted with private key a" in the "encrypted ID data of Company A" obtained by ID encryption in step A4 with the "encrypted ID of Company A encrypted with both private keys a and b" received in step A7, thereby obtaining the encrypted user data of Company A shown in the upper left of Figure 25. Also, the data matching unit 25 replaces the "encrypted ID of Company B encrypted with private key b" in the "encrypted user data of Company B" received in step A5 with the "encrypted ID of Company B encrypted with both private keys a and b" obtained by re-encryption in step A6, thereby obtaining the encrypted user data of Company B shown in the upper right of Figure 25. Then, the data matching unit 25 matches the encrypted user data of Company A with the encrypted user data of Company B using the "encrypted ID of Company A encrypted with both private keys a and b" and the "encrypted ID of Company B encrypted with both private keys a and b" as keys (step A8). Here, if "Company A's encrypted ID encrypted with both private keys a and b" and "Company B's encrypted ID encrypted with both private keys a and b" match, the attribute information in Company A's encrypted user data and the encrypted attribute information in Company B's encrypted user data are combined into a single record in the encrypted matching data, and after the combination, "Company A's encrypted ID encrypted with both private keys a and b" and "Company B's encrypted ID encrypted with both private keys a and b" are deleted. This results in encrypted matching data, shown in the lower part of Figure 25, which includes Company A's unencrypted attribute information and Company B's encrypted attribute information. Here, because Company B's attribute information is encrypted, Company A cannot know its contents, and the encrypted matching data is generated without revealing to Company A the contents of the attribute information that should be kept secret from Company B.
[0097] Next, as shown in FIG. 26 , in Company A's tally execution unit 104, the tally processing unit 26 performs the following tallying process on the encrypted matching data (step A9). As mentioned above, the attribute information in Company B's ID-encrypted data was encrypted in step B5 using a pre-prepared secret key B using a homomorphic encryption method that enables tallying processes. Therefore, in practice, each attribute information included in the encrypted matching data in FIG. 26 is configured with a binary value whose format (bit string arrangement) is predetermined. For example, for "gender," a predetermined value of "10" at the Pth and (P+1)th bits from the beginning of the encrypted matching data indicates "male," while a value of "01" indicates "female." Since Company A does not know secret key B, it cannot grasp the actual content of "gender" as described above, and can only grasp it as a simple bit string. Based on the above, the tally processing unit 26 performs the following tallying process on, for example, "item A (product purchase / non-purchase) x B (age group) x E (gender)" of tally pattern (1).
[0098] First, the aggregation processing unit 26 categorizes the multiple records constituting the encrypted matching data based on Company A's attribute information (unencrypted "product purchase" and "age"). That is, the multiple records constituting the encrypted matching data are divided into multiple categories, such as "product purchase, under 10s," "product purchase, in their 20s," ..., and "product purchase, over 70s." Next, the aggregation processing unit 26 performs aggregation processing for each category by vertically summing the multiple records represented by binary values that belong to that category (i.e., the same bits in the bit string). Note that Company B's attribute information represented by binary values is encrypted using an encryption method with homomorphic encryption properties, so the sum can be calculated even in the encrypted state. This results in the number of records for each category in which the Pth and (P+1)th bits are "10" (i.e., the aggregation result for gender being male) and the number of records in which the Pth and (P+1)th bits are "01" (i.e., the aggregation result for gender being female). In this embodiment, the input data is divided by category in the aggregation processing unit 26, but when the data to be aggregated is input to the aggregation execution unit 104 in Figure 7, the data may be divided in advance by aggregation pattern and aggregation processing may be performed.
[0099] Similarly, for the "item A (whether or not a product was purchased) x D (number of people living together) x E (gender)" in aggregation pattern (2), the aggregation processing unit 26 divides the multiple records that make up the encrypted matching data into categories of "product purchased, 1 person," "product purchased, 2 people," and "product purchased, 3 or more people" based on the attribute information of Company A (unencrypted "whether or not a product was purchased" and "number of people living together"), and performs aggregation processing (aggregating by gender for each of the above categories) by vertically adding up (the same bits in the bit string) the multiple records that belong to that category and are expressed as binary values.
[0100] Furthermore, the tabulation processing unit 26 obtains encrypted aggregated data by converting the encrypted aggregated data, which is the categorized encrypted matching data plus the aggregated data for each category, into the organized data format shown in the lower left and lower right of Fig. 26. In the tabulation processing of step A9 as described above, Company A does not know private key B, and therefore cannot grasp the actual content of "gender," which is Company B's attribute information, and it is possible to prevent the actual content of Company B's attribute information from being leaked to Company A.
[0101] 27, in the tally execution unit 104 of company A, the confidentiality processing unit 27 performs the following confidentiality processing on the encrypted tally data generated by the tally processing unit 26 to generate encrypted statistical information (step A10). For example, the confidentiality processing unit 27 generates noise using company B's computational key B' (a type of public key) and adds the generated noise to the encrypted tally data to generate encrypted statistical information. Note that the above computational key B' corresponds to secret key B and is assumed to have been shared in advance by company B with company A.
[0102] Then, the data transmission / reception unit 24 in Company A's aggregation execution unit 104 transmits the encrypted statistical information generated in step A10 to Company B's aggregation output unit 202 (step A11), and the data transmission / reception unit 24 of Company B's aggregation output unit 202 receives the encrypted statistical information (step B9).
[0103] Furthermore, the tally output unit 202 of Company B sends the received encrypted statistical information to the decryption unit 28. The decryption unit 28 of Company B, which knows private key B, decrypts the encrypted statistical information and outputs the resulting statistical information as appropriate, as shown in FIG. 28 (step B10). For example, the statistical information may be displayed or printed out by a predetermined operation performed by an operator of Company B's information processing device 200. FIG. 33 shows an example output in which the number of people with "Product Purchase: No" status is omitted from the above tally and the number of people with "Product Purchase: Yes" status is tallied. As shown in FIG. 33, for tally pattern (1), the number of people by age and gender who purchased the product is output, and for tally pattern (2), the number of people by number of people living with the same household and gender who purchased the product is output. Furthermore, as in the fourth embodiment described above, if a more detailed classification of data items (for example, "Age: 60s," "Age: 50s," and "Area of residence: Kanto") is selected as the target to be included in the aggregation pattern, the number of men and women in their 60s who purchased the product (a), the number of men and women in their 50s who purchased the product (b), and the number of men and women living in the Kanto region who purchased the product (c) will be output, as shown in Figure 34.
[0104] By performing the above-described step ST4 (aggregation process) in Figure 8, cooperation between the information processing device 100 of Company A and the information processing device 200 of Company B can be achieved, thereby generating statistical information that excludes correspondence with individuals, without disclosing the attribute information or private key that should be kept confidential to other devices.
[0105] The fifth embodiment described above makes it possible to appropriately select data items based on the degree of association between the data items and indicators related to the purpose of data utilization, and then determine the allocation of privacy budgets and perform aggregation processing for those data items. This makes it possible to improve the usefulness of statistical information output by data linkage when utilizing statistical data between multiple companies.
[0106] Sixth Embodiment In the following, as the sixth embodiment, in addition to selecting data items as in the first to fourth embodiments, an embodiment is described in which allocation of a privacy budget to aggregation patterns including the selected data items (without calculating a base value) and aggregation processing are performed. The sixth embodiment differs from the fifth embodiment only in the processing of step ST3 (determining allocation of a privacy budget) in Figure 8, and therefore the following description focuses on the configuration and processing (Figures 29 to 32) related to step ST3 that is unique to the sixth embodiment.
[0107] Because pre-processing, such as calculating the basic value, is not performed, the configuration indicated by arrow X in Fig. 5 does not require functional units corresponding to the basic value calculation unit 13 and the pre-execution data generation unit 12 in Fig. 6 described in the fifth embodiment, as shown in Fig. 29. Therefore, the decision unit 103 of Company A includes a data input unit 11, a basic tally information acquisition unit 14A, an estimation unit 15, an evaluation unit 16, and an allocation decision unit 17, while the information provision unit 201 of Company B includes a data input unit 11 and a basic tally information provision unit 14B. Since each functional unit has been described in the fifth embodiment, a repeated description will be omitted here.
[0108] As shown in Fig. 30, the process in the sixth embodiment does not include the generation of pre-execution data (steps S3 and S4 in Fig. 9) and the process related to calculation of the base value (steps S5 to S8 in Fig. 9). Furthermore, the processes in steps S1, S2, and S9 to S11 shown in Fig. 30 are the same as those in the fifth embodiment, and therefore, redundant explanations will be omitted.
[0109] In step S12, the estimation unit 15 of Company A's determination unit 103 estimates the sample size of the tabulated results using the same method as in the fifth embodiment, based on the "basic tabulated information" of Company A and Company B acquired by the basic tabulated information acquisition unit 14A in step S11 and "basic values" using known or estimated values. Here, m1 and m2 are obtained as the sample sizes of the tabulated results (here, medians of sample size estimates) for each item in the tabulation pattern (item A x B x E, item A x D x E). Note that although this embodiment shows an example in which the "median" of the sample size estimates is used, the average, minimum, etc. may also be used.
[0110] In the next step S13, the evaluation unit 16 of the determination unit 103 of company A quantitatively evaluates the influence of noise on each combination included in the aggregation pattern based on the estimated sample size of the aggregation result, using the same method as in the fifth embodiment. As an evaluation index for the influence of noise, the variance of the amount of change in the sample size n is used. By using the above formula (10), the following can be obtained for the summaries (1) and (2) as evaluation indices for the influence of noise, as shown in the table on the left side of FIG. In addition to the above, the evaluation index may be an index that measures the accuracy of different data (for example, sampling error rate, power of test, etc.) depending on the purpose of aggregation.
[0111] In the next step S14, the allocation determination unit 17 of the determination unit 103 of company A determines the allocation of the privacy budget to each combination as follows, based on the evaluation results for each combination obtained by the evaluation unit 16. Here, as a first allocation example, an example will be described in which, as in the fifth embodiment, each combination is weighted by a function that is inversely proportional to the magnitude of the noise influence for each combination, and the allocation of the privacy budget to each combination is determined; and further, as a second allocation example, an example will be described in which the allocation of the privacy budget to each combination is determined so as to satisfy the level of accuracy required to suppress the influence of noise for each combination.
[0112] In the first allocation example, in order to equalize the influence of noise, the allocation determination unit 17 weights the remaining privacy budget ε by a function that is inversely proportional to the magnitude of the influence of noise. total In the sixth embodiment, the "generation of pre-execution data (steps S3 and S4 in FIG. 2)" which consumes the privacy budget is not executed, so the remaining privacy budget ε total is "1.0", unlike in the fifth embodiment. Specifically, as shown on the right side of FIG. The remaining privacy budget ε total "1.0" is the total budget allocation ε for (1) and (2). b1 , ε b2 In this case, the influence of noise z in the aggregations (1) and (2) is the same (V b1=V b2 ).
[0113] In the second allocation example, the allocation determination unit 17 allocates the remaining privacy budget ε to each combination so as to satisfy the level of accuracy required to suppress the influence of noise for each combination. total An allocation of "1.0" is determined. In this case, if the required level of accuracy cannot be achieved within the privacy budget, it is possible to not perform a particular aggregation among aggregations (1) and (2), or to relax the required level of accuracy. One example of a method for quantitatively determining the required level of accuracy is to set the required level of accuracy as the level at which the items to be detected (e.g., purchase rates by attribute (age or gender), differences in the number of purchasers by attribute (age or gender)), etc.) are expected to be detectable from the aggregation results. For example, the right side of Figure 32 shows an example of determining the allocation of the privacy budget so that the noise influence for each combination satisfies the required level of accuracy (here, in accordance with the level (i.e., so that it is the same as the level)). In this case, the noise influence for each combination satisfies the required level of accuracy, and the remaining privacy budget is allocated to the budget allocation ε c1 , ε c2 Here, the privacy budget (ε c1 +ε c2 ) is the remaining privacy budget ε total It will be smaller than "1.0", resulting in a surplus privacy budget.
[0114] Note that the method of allocating the privacy budget is not limited to the first and second allocation examples described above, and various allocation methods other than the first and second allocation examples described above may be adopted depending on the purpose of the aggregation.
[0115] In the sixth embodiment described above, as in the fifth embodiment, data items can be appropriately selected based on the degree of association between the data items and indicators related to the purpose of data utilization, and then the allocation of privacy budgets and aggregation processing can be performed for those data items. This makes it possible to improve the usefulness of statistical information output by data linkage when utilizing statistical data between multiple companies.
[0116] Regarding the method of allocating the privacy budget, various allocation methods such as the first and second allocation examples have been explained in the sixth embodiment above, but these various allocation methods can also be applied to the fifth embodiment.
[0117] In the fifth and sixth embodiments, the information processing device 100 of Company A includes functional units such as the relevance calculation unit 101 and the selection unit 102, while the information processing device 200 of Company B does not include these functional units, as shown in FIG. 5 . However, the present invention is not limited to this configuration. As a modified example, both the information processing devices 100 and 200 of the respective companies may have a similar configuration including the functional units described above, as shown in FIG. 35 . In this case, the relevance calculation and selection processes can be performed by either device. However, the determination unit 103 also has a function to operate as the information provision unit 201 of FIG. 5 and operates as the information provision unit 201 when the determination unit 103 of the other device takes the initiative in determining the privacy budget allocation. The tallying unit 104 also has a function to operate as the tally output unit 202 of FIG. 5 and operates as the tally output unit 202 when the tallying unit 104 of the other device takes the initiative in executing the tallying process. Even with this modified example, the same effects as those of the fifth and sixth embodiments can be obtained.
[0118] The gist of the present disclosure lies in the following [1] to [9]. [1] An information processing device that, when performing aggregation between a local device and a remote device that stores user data including a user ID and attribute information about a user, for aggregation patterns including multiple combinations of data items included in the attribute information, adds noise based on a differential privacy standard to aggregation results within a privacy budget and performs aggregation for the aggregation patterns, the information processing device comprising: an association degree calculation unit that calculates an association degree between an index related to a purpose of data utilization and the data items included in the attribute information, and a selection unit that selects data items to be included in the aggregation pattern based on the association degree calculated by the association degree calculation unit. [2] The information processing device described in [1], in which the association degree calculation unit generates a machine learning prediction model using the index as a response variable and the data items included in the attribute information as explanatory variables, and calculates the association degree based on the obtained prediction model. [3] The information processing device according to [1], wherein the relevance calculation unit calculates a correlation coefficient between the index and the data item included in the attribute information, and defines the obtained correlation coefficient as the relevance. [4] The information processing device according to any one of [1] to [3], wherein the relevance calculation unit calculates the relevance between each of the data items included in the attribute information and the index for each category obtained by classifying the data item into a plurality of categories, and the selection unit selects data items to be included in the tally pattern based on the relevance for each category for each of the data items. [5] The information processing device according to [4], wherein the selection unit identifies the categories with high relevance based on a predetermined criterion, and selects data items to be included in the tally pattern based on the number of categories with high relevance for each data item. [6] The information processing device according to [4], wherein the selection unit identifies the categories with high relevance based on a predetermined criterion, and selects the categories with high relevance as data items to be included in the tally pattern.[7] The information processing device according to any one of [1] to [6], further comprising: a determination unit that determines an allocation of a privacy budget to each of the combinations including the data items selected by the selection unit; and an aggregation execution unit that adds noise based on a differential privacy standard to the aggregation results within the privacy budget allocated based on the allocation determined by the determination unit for each of the combinations, and executes aggregation targeting the aggregation pattern. [8] An information system including a first device and a second device that retain user data including a user ID and attribute information related to the user, wherein the first device is a device that, when performing aggregation between the first device and the second device for an aggregation pattern that includes a plurality of combinations of data items included in the attribute information, performs aggregation for the aggregation pattern by adding noise based on a differential privacy standard to the aggregation result within a privacy budget, wherein the information system includes: a relevance calculation unit that calculates a relevance between an index related to a purpose of data utilization and the data items included in the attribute information; a selection unit that selects data items to be included in the aggregation pattern based on the relevance calculated by the relevance calculation unit; a determination unit that determines an allocation of a privacy budget to each combination that includes the data items selected by the selection unit; and a aggregation execution unit that performs aggregation for the aggregation pattern by adding noise based on a differential privacy standard to the aggregation result within the privacy budget allocated based on the determined allocation for each of the combinations; and An information system comprising: an information providing unit that cooperates with the decision unit to provide information necessary for the decision unit to determine an allocation of a privacy budget; and a tally output unit that cooperates with the tally execution unit to provide information necessary for the tally execution unit to tally, and decrypts and outputs the tally results.[9] An information system including a plurality of devices that hold user data including a user ID and attribute information related to the user, wherein each device is a device that, when performing aggregation between its own device and a counterpart device for an aggregation pattern that includes a plurality of combinations of data items included in the attribute information, performs aggregation for the aggregation pattern by adding noise based on differential privacy standards to the aggregation results within a privacy budget, wherein each device comprises: a relevance calculation unit that calculates the relevance between an index related to a purpose of data utilization and the data items included in the attribute information; a selection unit that selects data items to be included in the aggregation pattern based on the relevance calculated by the relevance calculation unit; a determination unit that determines an allocation of the privacy budget to each combination that includes the data items selected by the selection unit; and an aggregation execution unit that, for each of the combinations, adds noise based on differential privacy standards to the aggregation results within the privacy budget allocated based on the determined allocation, and performs aggregation for the aggregation pattern.
[0119] (Explanation of terms, explanation of hardware configuration (Figure 36), etc.) Note that the block diagrams used in the explanation of the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or may be realized using two or more devices that are physically or logically separated and connected directly or indirectly (for example, using wires, wirelessly, etc.) and these multiple devices. A functional block may be realized by combining software with the one device or the multiple devices.
[0120] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0121] For example, an information processing device according to an embodiment of the present disclosure may function as a computer that executes the processes of the present disclosure. Fig. 36 is a diagram illustrating an example of a hardware configuration of an information processing device 100 according to an embodiment of the present disclosure. The information processing device 100 described above may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc., and the information processing device 200 may be configured in a similar manner.
[0122] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the information processing apparatus 100 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.
[0123] Each function in the information processing device 100 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.
[0124] The processor 1001 controls the entire computer by running, for example, an operating system, and may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc.
[0125] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. While the various processes have been described as being executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may be transmitted from a network via a telecommunications line.
[0126] The memory 1002 is a computer-readable recording medium and may be configured by, for example, at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing a wireless communication method according to an embodiment of the present disclosure.
[0127] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.
[0128] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD).
[0129] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0130] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0131] The information processing device 100 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0132] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0133] Each aspect / embodiment described in the present disclosure may be implemented using any of the following standards: LTE (Long Term Evolution), LTE-Advanced (LTE-A), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), 6th generation mobile communication system (6G), xth generation mobile communication system (xG) (xG (x is, for example, an integer or a decimal number)), FRA (Future Radio Access), NR (new Radio), New radio access (NX), Future generation radio access (FX), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.17 (WiMAX (registered trademark)), IEEE 802.19 (WiMAX (registered trademark)), IEEE 802.20 (WiMAX (registered trademark)), IEEE 802.21 (Wi-Fi (registered trademark)), IEEE 802.22 (WiMAX (registered trademark)), IEEE 802.23 (WiMAX (registered trademark)), IEEE 802.24 (WiMAX (registered trademark)), IEEE 802.25 (WiMAX (registered trademark)), IEEE 802.26 (WiMAX (registered trademark)), IEEE 802.27 (WiMAX (registered trademark)), IEEE 802.28 (WiMAX (registered trademark)), IEEE 802.29 (WiMAX (registered trademark)), IEEE 802.30 (WiMAX (registered trademark)), IEEE 802.31 (Wi-Fi (registered trademark)), IEEE 802.32 (WiMAX (registered trademark)), IEEE 802.33 (WiMAX (registered trademark)), IEEE 802.34 ( The present invention may be applied to at least one of systems using 802.20, UWB (Ultra-Wide Band), Bluetooth (registered trademark), or other suitable systems, and next-generation systems that are extended, modified, created, or defined based on these systems. The present invention may also be applied to a combination of multiple systems (e.g., a combination of LTE and / or LTE-A with 5G).
[0134] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0135] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be transmitted to another device.
[0136] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0137] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0138] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0139] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0140] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0141] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0142] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.
[0143] As used in this disclosure, the terms "system" and "network" are used interchangeably.
[0144] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, may be expressed using relative values from a predetermined value, or may be expressed using other corresponding information. For example, a radio resource may be indicated by an index.
[0145] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0146] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0147] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0148] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0149] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.
[0150] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0151] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0152] 11...data input unit, 12...pre-execution data generation unit, 12A...data processing unit, 12B...anonymization processing unit, 12C...ID irreversible conversion unit, 12D...encryption unit, 12E...data transmission / reception unit, 13...basic value calculation unit, 13A...data matching unit, 13B...tallying processing unit, 13C...confidentiality processing unit, 13D...calculation unit, 14A...basic summary information acquisition unit, 14B...basic summary information provision unit, 15...estimation unit, 16...evaluation unit, 17...allocation determination unit, 21...anonymization processing unit, 22...ID irreversible conversion unit, 23...encryption unit, 23A ...ID encryption unit, 23B...attribute information encryption unit, 24...data transmission / reception unit, 25...data matching unit, 26...counting processing unit, 27...secrecy processing unit, 28...decryption unit, 100...information processing device, 101...association calculation unit, 102...selection unit, 103...determination unit, 104...counting execution unit, 200...information processing device, 201...information provision unit, 202...counting output unit, 1001...processor, 1002...memory, 1003...storage, 1004...communication device, 1005...input device, 1006...output device, 1007...bus.
Claims
1. a relevance calculation unit that calculates a relevance between an index related to a purpose of data utilization and a data item included in attribute information related to a user, the attribute information being used for aggregation between a user's own device and a counterpart device, each of the user's attribute information being held; a selection unit that selects, based on the relevance, data items to be included in a counting pattern that is a target of the counting and includes noise added based on a differential privacy criterion within a privacy budget, the counting pattern including a plurality of combinations of the data items; and An information processing device comprising:
2. the relevance calculation unit generates a machine learning prediction model using the index as a response variable and the data item included in the attribute information as an explanatory variable, and calculates the relevance based on the obtained prediction model. The information processing device according to claim 1 .
3. the relevance calculation unit calculates a correlation coefficient between the index and the data item included in the attribute information, and defines the obtained correlation coefficient as the relevance. The information processing device according to claim 1 .
4. the relevance calculation unit calculates the relevance between each of the data items included in the attribute information and the index for each category obtained by classifying the data item into a plurality of categories; the selection unit selects data items to be included in the aggregation pattern based on the relevance for each of the categories for each of the data items. The information processing device according to claim 1 .
5. the selection unit identifies the categories with high relevance based on predetermined criteria, and selects data items to be included in the aggregation pattern based on the number of categories with high relevance for each data item. The information processing device according to claim 4 .
6. the selection unit identifies the category having a high degree of association based on a predetermined criterion, and selects the category having a high degree of association as a data item to be included in the aggregation pattern. The information processing device according to claim 4 .
7. a determination unit that determines an allocation of a privacy budget to each of the combinations that includes the data item selected by the selection unit; a counting execution unit that adds noise based on a differential privacy criterion to a counting result within the privacy budget allocated based on the allocation determined by the determination unit for each of the combinations and executes counting for the counting patterns; The information processing device according to claim 1 , further comprising:
8. a relevance calculation unit that calculates a relevance between an index related to a purpose of data utilization and a data item included in attribute information used for aggregation between a first device, which is a user's own device, and a second device, which is a counterpart device, each of which holds attribute information related to a user; a selection unit that selects, based on the relevance, data items to be included in a counting pattern that is a target of the counting and includes noise added based on a differential privacy criterion within a privacy budget, the counting pattern including a plurality of combinations of the data items; and a determination unit that determines an allocation of a privacy budget to each combination that includes the data item selected by the selection unit; a counting execution unit that adds noise based on a differential privacy criterion to a counting result within the privacy budget allocated based on the determined allocation for each of the combinations and executes counting for the counting patterns; a first device comprising: an information providing unit that cooperates with the decision unit to provide information necessary for the decision unit to determine the allocation of the privacy budget; a tally output unit that cooperates with the tally execution unit to provide information necessary for the tally execution unit to perform tallying, and decodes and outputs the tally results; a second device comprising: Information systems, including:
9. a relevance calculation unit that calculates a relevance between an index related to a purpose of data utilization and the data item included in the attribute information used for aggregation between a user's own device and a counterpart device, each of which holds attribute information about the user; a selection unit that selects, based on the relevance, data items to be included in a counting pattern that is a target of the counting and includes noise added based on a differential privacy criterion within a privacy budget, the counting pattern including a plurality of combinations of the data items; and a determination unit that determines an allocation of a privacy budget to each combination that includes the data item selected by the selection unit; a counting execution unit that adds noise based on a differential privacy criterion to a counting result within the privacy budget allocated based on the determined allocation for each of the combinations and executes counting for the counting patterns; a plurality of devices comprising: Information systems, including:
10. An information processing device comprising: a step of calculating a degree of association between an index related to a purpose of data utilization and a data item included in attribute information used for aggregation between a device and a counterpart device, each of which holds attribute information related to a user; a step in which the information processing device selects, based on the relevance, data items to be included in a counting pattern to be subjected to the counting, the counting pattern including a plurality of combinations of the data items, the counting pattern including addition of noise based on a differential privacy criterion within a privacy budget; An information processing method comprising: