Information processing device, information processing method, and information processing program

The information processing device enhances data utilization by classifying and noise-adding personal information based on user-defined conditions, addressing the limitations of uniform anonymization and enabling more extensive data use while preserving privacy.

JP7778645B2Active Publication Date: 2025-12-02LY CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022098770
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-12-02
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

Conventional methods of using personal information for data analysis do not effectively utilize the full potential of available data due to uniform anonymization, limiting the amount of personal information that can be utilized.

Method used

An information processing device that classifies user data based on usage conditions, adds noise to the data according to specific groups, and generates use target data using noise-added data groups to enhance the utilization of personal information.

Benefits of technology

Enables the use of a larger amount of personal information by allowing for more nuanced data utilization while maintaining privacy through differential privacy and k-anonymization techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778645000001
    Figure 0007778645000001
  • Figure 0007778645000002
    Figure 0007778645000002
  • Figure 0007778645000003
    Figure 0007778645000003
Patent Text Reader

Abstract

To provide an information processing apparatus, an information processing method, and an information processing program that can use more pieces of individual information.SOLUTION: An information processing apparatus according to the present application comprises a classification unit, a noise addition processing unit, and a data to be used generation unit. The classification unit classifies pieces of data the use of which is permitted by a user into groups according to use conditions set by the user. The noise addition processing unit generates, for each of the groups, noise addition data in which noise according to the group is added to data of an arithmetic result based on a user data group including the plurality of pieces of data classified into the same group by the classification unit, or a noise addition data group in which the noise according to the group is added to the user data group. The data to be used generation unit generates data to be used by using the noise addition data or the noise addition data group for each of the groups generated by the noise addition processing unit.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] In recent years, there has been a study into the use of anonymized personal information, such as data on users of online services, which has been processed to prevent the identification of specific individuals, for data analysis by third parties, etc. For example, Patent Document 1 proposes a technology for protecting personal information using differential privacy. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2018 / 116366 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the conventional method of using personal information, each user is asked to choose whether to allow or deny the use of their personal information, and the personal information that the user allows is uniformly anonymized and used, so there is room for improvement in terms of making it possible to use more personal information.

[0005] The present application has been made in view of the above, and aims to provide an information processing device, an information processing method, and an information processing program that can make it possible to utilize a larger amount of personal information. [Means for solving the problem]

[0006] The information processing device according to the present application includes a classification unit, a noise addition processing unit, and a use target data generation unit. The classification unit classifies data whose use is permitted by a user into groups according to usage conditions set by the user. The noise addition processing unit generates, for each group, noise-added data or a noise-added data group by adding noise appropriate to the group to data resulting from an operation based on a user data group including a plurality of data classified into the same group by the classification unit. The use target data generation unit generates the use target data using the noise-added data or noise-added data group for each group generated by the noise addition processing unit. [Effects of the Invention]

[0007] According to one aspect of the embodiment, more personal information can be made available. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of information processing according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of the information providing system according to the embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of the information processing device according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of an organization information table stored in the organization information storage unit of the information processing device according to the embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of a service log table stored in the service log storage unit of the information processing device according to the embodiment. [Figure 6] FIG. 6 is a flowchart illustrating an example of information processing executed by a processing unit of an information processing device. [Figure 7] FIG. 7 is a flowchart showing an example of anonymization processing executed by the processing unit of the information processing device. [Figure 8]FIG. 8 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of the information processing device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, modes for implementing an information processing device, an information processing method, and an information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing device, the information processing method, and the information processing program according to the present application are not limited to these embodiments. Furthermore, the respective embodiments can be appropriately combined within the scope of not causing any contradiction in the processing content. Furthermore, the same components in the following embodiments will be assigned the same reference numerals, and redundant explanations will be omitted.

[0010] [1. An example of information processing] First, an example of information processing according to the embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of information processing according to the embodiment.

[0011] The information processing device 1 shown in Figure 1 collects data of multiple users U from each of multiple organizations (e.g., companies, local governments, etc.) and anonymizes the collected data of the multiple users U so that the data of each individual user U cannot be identified.

[0012] The information processing device 1 then generates target data for use from the anonymized data and provides the target data for use. The target data for use is data that is to be used by a data user, and the target data for use can be used to perform, for example, data analysis.

[0013] As shown in FIG. 1, servers 31 to 3 m Each of these provides online services to the user U of the terminal device 2 (steps S11 to S1 m ) m is an integer equal to or greater than 2. Servers 31-3 m are servers that provide online services provided by different organizations, for example.

[0014] If the organization is a business, the online services provided by the organization may be, for example, online services such as a shopping site, news site, auction site, flea market site, weather forecast site, finance (stock price) site, etc. Furthermore, the online services provided by the business may also be, for example, online services such as a search site, map providing site, travel site, restaurant introduction site, weblog site, or SNS site.

[0015] If the organization is a local government, the online services provided by the organization are administrative services, such as online services such as electronic applications, electronic bidding, electronic tax returns, electronic tax payments, reservations for public facilities, reservations for book lending, and provision of various information related to administrative services.

[0016] Server 31~3 m Each of the users inquires about whether or not the personal information can be used, and receives a response from the user U about whether or not the personal information can be used (steps S21 to S2 m ) Personal information is stored, for example, on servers 31-3 m The user information includes the user U's usage history of the online services provided by each of the above, and the user U's attribute information.

[0017] The inquiry regarding whether or not personal information can be used is an inquiry as to whether to deny use of the personal information or to conditionally permit use of the personal information, and the conditional permission includes at least one of the conditions of use for each user U and the conditions of use for each data. Each of the conditions of use for each user U and the conditions of use for each data is indicated by the level of protection of personal information when the data is used, and is determined by each user U.

[0018] For example, each condition of use can be set to multiple protection levels AL1, AL2, AL n The protection levels AL1, AL2, . . ., AL nOf these, protection level AL1 is the lowest level of protection of personal information, protection level AL2 is the next lowest level of protection of personal information after protection level AL1, and protection level AL n The level of protection of personal information is the highest. In the following, the protection levels AL1, AL2, AL n The protection level selected by the user U may be referred to as the selected protection level SAL. n When referring to each of these without distinguishing them individually, they may be referred to as protection level AL.

[0019] If a user U who conditionally permits the use of personal information does not select the above-mentioned conditions of use, the information processing device 1 can also determine the conditions of use based on the information of the user U. For example, the information processing device 1 can determine the conditions of use using a protection level determination model that inputs the information of the user U and outputs a score for each protection level.

[0020] In this case, the information processing device 1 inputs the user U who does not select the usage conditions into the protection level determination model, and determines the protection level with the highest score among the scores for each protection level output from the protection level determination model as the protection level for the user U who does not select the usage conditions.

[0021] The information processing device 1 can generate a protection level determination model by machine learning, for example, using a data set including information about the user U and the protection level selected by the user U as learning data.

[0022] The protection level determination model may be, for example, a learning model generated by a GBDT (Gradient Boosting Decision Tree) or a learning model generated by deep learning using a deep neural network (DNN), but is not limited to these examples and may also be a learning model generated by other machine learning methods.

[0023] Server 31~3 m Each of these transmits information to the information processing device 1 (steps S31 to S3 m ). Server 31~3 m The information transmitted from each of the servers 31 to 3 includes service log data of the online service and information indicating whether the service log data of the online service is available. m The information transmitted from each of the above may include provider organization information indicating the organization that provides the online service, authorized organization information indicating the organization that is authorized to use the data in the service log, and the like.

[0024] The service log data for online services is stored on servers 31-3. m The service log data is usage log data of the online services provided by each of the above services, and includes, for example, information on the usage history of the online services by user U and registration information of user U. Hereinafter, the service log data of the online services may be referred to as service log data or service log. The service log data is an example of provided data.

[0025] The registration information of the user U includes, for example, attribute information and protection level information of the user U. The protection level information includes a plurality of protection levels AL1, AL2, . . . , AL n The protection level SAL is a selected protection level selected by the user U from the above.

[0026] Server 31~3 m The service log data included in the information transmitted from the servers 31 to 30 to the information processing device 1 does not include data indicating that the user U has refused to use personal information. m Each of the above includes data of multiple users U who use the online service for which the user U has conditionally permitted the use of personal information, but does not include data for which the user U has refused the use of personal information. Data for which the user U has refused the use of personal information is data that has been refused on a user U basis or on a data basis.

[0027] The information indicating whether the service log data is available for use includes either denial information that denies use of the service log data or permission information that permits use of the service log data. The information indicating whether the service log data is available for use is, for example, information set by an organization that provides an online service.

[0028] The information processing device 1 includes servers 31 to 3 m The servers 31 to 3 receive the information transmitted from each of the servers 31 to 3. m Each piece of information is stored in association with the corresponding organization (step S4). m When each of these is not individually indicated, it may be referred to as server 3, and steps S31 to S3 m When each of these steps is not individually indicated, they may be referred to as step S3.

[0029] In the process of step S4, the information processing device 1 stores the service log data transmitted from each server 3 by associating it with an organization among the multiple organizations that satisfies the association condition. The organization that satisfies the association condition is, for example, an organization that provides an online service or an organization that provided the service log data.

[0030] For example, if the server 31 is a server of an organization managed by an organization that provides online services, the information processing device 1 associates the service log data transmitted from the server 31 with the organization that manages the server 31.

[0031] In addition, the information processing device 1 may include, for example, a server 3 m Server 3 m If the server provides online services to an organization other than the one that manages it, Server 3 m The service log data sent from Server 3 m This links the online service to the organization that provides it.

[0032] If the information transmitted from the server 3 in step S3 includes providing organization information indicating an organization that provides online services, the information processing device 1 can determine that the organization indicated by the providing organization information is an organization that satisfies the linking condition.

[0033] Furthermore, an organization that satisfies the linking condition may be an organization for which the score output from the relevance estimation model satisfies a preset condition. The relevance estimation model is, for example, a learning model that receives service log data as input and outputs a score indicating the relevance of each organization with the source of the service log data request.

[0034] The relevance estimation model is generated by using service log data that has been determined in advance to be relevant to each organization as the correct answer data for each organization and having the model learn the features of the correct answer data for each organization. The relevance estimation model is a learning model common to multiple organizations, but it may also be a learning model for each organization.

[0035] The information processing device 1 inputs the service log data into the relevance estimation model, and identifies organizations whose scores satisfy predetermined conditions among the scores for each organization output from the relevance estimation model as organizations whose scores output from the relevance estimation model satisfy the preset conditions. The score that satisfies the predetermined condition is, for example, the highest score, but may also be, for example, a score equal to or greater than a threshold. If there are multiple organizations with scores equal to or greater than the threshold, the information processing device 1 can link the service log data to the multiple organizations.

[0036] The relevance estimation model may be, for example, a learning model generated by GBDT or a learning model generated by deep learning using a deep neural network, but is not limited to such examples and may also be a learning model generated by other machine learning methods.

[0037] For example, the relevance estimation model may be generated using machine learning with other learning algorithms, such as a regression method such as linear regression, multiple regression, or logistic regression, or a learning algorithm such as a support vector machine.

[0038] For example, the information processing device 1 can store service log data that has been disabled without linking it to an organization. The information processing device 1 can determine whether the service log data has been disabled based on information indicating whether the service log data is available or not, which is included in the information transmitted from the server 3.

[0039] Next, the information processing device 1 receives a query from the requester device 4 as a request to provide data (step S5). The query from the requester device 4 includes, for example, data identification information for identifying the data to be used or the data to be processed, and organization identification information such as the organization ID (IDentifier) ​​of the requesting organization. The requesting organization is, for example, an organization to which an operator operating the requester device 4 belongs, or an organization that has requested an analysis from the organization to which the operator belongs.

[0040] The data identification information includes target organization information, processing target information, etc. The target organization information includes the organization ID of the target organization, which is the organization that owns the data to be used. For example, if the data that the requesting organization wants to analyze is service log data of an online service provided by organization B, the organization ID of the target organization is the organization ID of organization B.

[0041] The processing target information includes data type information indicating the type of processing target data, which is data to be processed, processing content information indicating the processing content of the processing target data, etc. The data type information is, for example, information indicating the location of a user U who used an online service, information indicating the type of usage history of the online service by the user, information indicating the type of attributes of the user, and information indicating the type of input information and output information of a learning model.

[0042] The information indicating the location of the user is, for example, information indicating the location detected by a location detection unit (not shown) provided in the terminal device 2, and is information transmitted from the terminal device 2 to the server 3 when the user uses an online service.

[0043] The usage history of the online service includes the date and time of use of the online service, the usage content of the online service, etc. For example, if the online service is a service provided by a shopping site, the usage content of the online service includes the content of browsing by the user of the transaction target, the content of purchases by the user of the transaction target, the content of evaluations by the user of the transaction target, and the content of favorite registrations by the user of the transaction target.

[0044] For example, if the online service is an administrative service, the usage content of the online service may include the content of electronic applications made by the user, the content of electronic bidding made by the user, the content of electronic declarations made by the user, the content of electronic tax payments made by the user, the content of public facility reservations made by the user, the content of book lending reservations made by the user, and the viewing of various information by the user.

[0045] The attributes of user U are at least one of demographic attributes and psychographic attributes. Demographic attributes are demographic attributes, such as age, gender, occupation, place of residence, annual income, and family structure. Psychographic attributes are psychological attributes, such as lifestyle, values, and interests.

[0046] The processing content information includes, for example, information indicating the type of statistical calculation such as average, sum, or variance, the type of calculation for generating parameters of a learning model, etc. The processing content specified by the processing target information is, for example, calculation of the average annual income of men in their twenties in Kyoto Prefecture, the average variance of people who like motorcycles, the total number of users U who have made electronic tax payments, or generation of a learning model that inputs information indicating the location and attributes of user U and outputs a score indicating the possibility of user U using online services.

[0047] Next, the information processing device 1 determines whether the provision candidate data, which is data identified by the query received in step S5, is data linked to a specific related organization (step S6). A specific related organization is an organization that has a predetermined relationship with the requester of the provision request. The requester of the provision request is the organization whose organization ID is indicated by the organization identification information included in the query.

[0048] For example, when the target organization indicated by the organization ID included in the target organization information is an organization that is in a predetermined competitive relationship with the requester of the provision request, the information processing device 1 determines that the provision candidate data is data of a specific related organization. For example, the information processing device 1 has a relationship table that includes, for each organization, information indicating organizations that are in a predetermined competitive relationship, and uses the relationship table to determine whether the provision candidate data is data of a specific related organization.

[0049] If the information sent from the server 3 in step S3 includes authorized organization information indicating an organization that is permitted to use the service log data, the information processing device 1 can also determine that the organization indicated in the authorized organization information is an organization in a predetermined competitive relationship.

[0050] The information processing device 1 can also use information about competing organizations as correct data and estimate whether an organization is in a competitive relationship using a related organization estimation model that has learned the characteristics of the correct data about competing organizations. The related organization estimation model is a learning model for each organization that has requested the provision request. In this case, the information processing device 1 inputs information about each organization into the related organization estimation model of the organization that has requested the provision request, and determines that an organization whose score output from the related organization estimation model is equal to or greater than a threshold is a competing organization.

[0051] In the above example, an organization that is in a predetermined competitive relationship with the requester of the provision request has been described as a specifically related organization, but the example is not limited to this. For example, the information processing device 1 can also determine an organization that is not in a predetermined competitive relationship with the requester of the provision request as a specifically related organization.

[0052] For example, the information processing device 1 has a relationship table that includes, for each organization, information indicating organizations that are not in a predetermined competitive relationship, and uses this relationship table to determine a specific related organization.

[0053] Furthermore, the information processing device 1 can determine specific related organizations by using, for example, information about organizations that are not in a competitive relationship as correct answer data and using a related organization estimation model that has learned the characteristics of the correct answer data of organizations that are not in a competitive relationship.

[0054] Next, if the information processing device 1 determines that the candidate data to be provided is data linked to a specific related organization, it identifies the type of data that is linked to the specific related organization and corresponds to the query received in step S5 (step S7).

[0055] As described above, the query received in step S5 includes the processing target information, and the information processing device 1 identifies the type of target data, which is data corresponding to the query, based on the data type information included in the processing target information. For example, if the data type information is information indicating the annual income of men in their twenties living in Kyoto Prefecture, the information processing device 1 identifies the annual income of men in their twenties living in Kyoto Prefecture as the type of target data.

[0056] Furthermore, when the data type information includes information indicating the types of input information and output information of the learning model, the information processing device 1 identifies the types of input information and output information of the learning model as the types of target data. The input information is, for example, information indicating the location and attributes of the user U, and the output information is, for example, information indicating the usage content of the online service by the user U, but is not limited to such examples.

[0057] In the process of step S7, the information processing device 1 can also accept the designation of data so that the number of target data or the number of users U who are the source of the target data is equal to or greater than a threshold. For example, the information processing device 1 can transmit a list of attributes for which the number of target data or the number of users U who are the source of the target data is equal to or greater than a threshold to the requester device 4, thereby allowing the operator to select an attribute from the list of attributes.

[0058] Next, the information processing device 1 performs anonymization on the data linked to the specific related organization and of the type identified in step S7 (step S8).

[0059] For example, suppose the specific related organization is organization B, and the type identified in step S7 is a specific attribute of user U. In this case, the information processing device 1 performs anonymization on data of the specific attribute from the service log data of organization B. For example, if the type identified in step S7 is the annual income of men in their twenties living in Kyoto Prefecture, the information processing device 1 performs anonymization on the annual incomes of multiple users U who are men in their twenties living in Kyoto Prefecture, as data of the specific attribute. Hereinafter, data linked to the specific related organization and of the type identified in step S7 will be referred to as data to be anonymized. The data to be anonymized includes multiple pieces of data of the type identified in step S7, and can also be considered a group of data to be anonymized.

[0060] In step S8, the information processing device 1 first classifies each piece of data included in the data to be anonymized into groups according to the conditions of use set by the user U. The data included in the data to be anonymized is data whose use has been permitted by the user U.

[0061] For example, the usage conditions set by a user U may be set to multiple protection levels AL1, AL2, . . . , AL n In this case, the information processing device 1 may select a protection level from among a plurality of protection levels AL1, AL2, . . . , AL n corresponding to multiple groups G1, G2, , G nClassify the data included in the data to be anonymized as follows.

[0062] For example, data for which a usage condition indicated by protection level AL1 is set is classified into group G1, data for which a usage condition indicated by protection level AL2 is set is classified into group G2, and data for which a usage condition indicated by protection level AL3 is set is classified into group G3. n Data for which the terms of use indicated by are set is in Group G n In the following, we consider multiple groups G1, G2, . . ., G n When referring to each of these without distinguishing them individually, they may be referred to as Group G.

[0063] The information processing device 1 classifies each piece of data included in the data to be anonymized into a group G according to the usage conditions set by the user U, and then generates, for each group G, noise-added data in which noise according to the group G is added to the data of calculation results (e.g., statistical quantities such as average, sum, and variance) based on a user data group including multiple pieces of data classified into the same group G, or a noise-added data group in which noise according to the group G is added to the user data group.

[0064] The information processing device 1 generates noise-added data or a noise-added data group by anonymization using differential privacy or k-anonymization. First, anonymization using differential privacy will be described, and then anonymization using k-anonymization will be described.

[0065] The information processing device 1 calculates data of the calculation result for each group G from the user data group for each group G based on the processing content indicated by the processing content information included in the query. The processing content indicated by the processing content information is, for example, statistical calculations such as average, sum, and variance for specific data, or generation of parameters for a learning model.

[0066] For example, assume that the processing content indicated by the processing content information is the average annual income of users U. In this case, the information processing device 1 calculates data indicating the average annual income of users U for each group G as data of the calculation result for each group G.

[0067] Furthermore, it is assumed that the processing content indicated by the processing content information is the distribution of ages of users U. In this case, the information processing device 1 calculates data indicating the distribution or variance of ages of users U for each group G as data of the calculation result for each group G.

[0068] Furthermore, suppose that the processing content indicated by the processing content information is generation of parameters for a learning model. In this case, the information processing device 1 calculates, for each group G, data indicating the average value of multiple gradients resulting from training data samples when optimizing the learning model using stochastic gradient descent.

[0069] Furthermore, the information processing device 1 calculates, as the calculation result data, data indicating the number (total) of users U for each group G and data indicating the variance for each group G to be statistically processed, in addition to the processing content indicated by the processing content information. Hereinafter, the number of users U may be referred to as the number of users.

[0070] Then, the information processing device 1 generates noise-added data for each group G by adding noise according to the group G to the data of the calculation result for each group G. For example, n corresponding to multiple groups G1, G2, , G n Assume that the data contained in the data to be anonymized is classified.

[0071] In this case, the information processing device 1 is divided into groups G1, G2, . . . , G n Noise with increasing noise level is added to the data resulting from the calculation in this order. In this way, noise-added data is generated for each group G, with noise having a higher noise level added to the group G according to the usage conditions with a higher data protection level AL.

[0072] The noise is, for example, fixed noise, Gaussian noise, Laplace noise, etc. When the noise added to the data of the calculation result is fixed noise, the information processing device 1 increases the value of the fixed noise as the noise level increases.

[0073] Furthermore, when the noise added to the data resulting from the calculation is Gaussian noise, the information processing device 1 can increase the noise level by increasing the absolute value of the average value of the noise or the standard deviation of the noise. Furthermore, when the noise added to the data resulting from the calculation is Laplace noise, the information processing device 1 can increase the noise level by increasing the absolute value of the average value of the noise or the standard deviation of the noise.

[0074] In this way, the information processing device 1 generates noise-added data for each group G by adding noise according to the group G to the data of the calculation result for each group G in the anonymization process using differential privacy.

[0075] Next, a description will be given of anonymization using k-anonymization in the information processing device 1. K-anonymization is a technology that makes it difficult to identify an individual by converting the probability of identifying an individual to 1 / k or less, and, for example, adds noise to the data to be anonymized so that there are k or more users U with the same combination of quasi-identifiers included in the data of the user U.

[0076] A combination of quasi-identifiers is a combination of attribute items such as age, sex, date of birth, and place of residence (or location). The information processing device 1 performs anonymization by, for example, performing suppression processing to add noise by converting a certain attribute item into an asterisk "*" or the like, or by performing generalization processing to add noise by raising the hierarchy of the attribute item. For example, when the attribute item is place of residence (or location), the generalization processing is processing such as changing "Chiyoda-ku, Tokyo" to "Tokyo."

[0077] The information processing device 1 can also add noise by top coding, which groups values ​​above a certain threshold into one category, or bottom coding, which groups values ​​below a certain threshold into one category. Top coding is a process in which, for example, if the attribute item is age, "80 years old," "81 years old," and "92 years old" are converted into "80 years old or older." Bottom coding is a process in which, for example, if the attribute item is age, "0 years old," "1 year old," and "3 years old" are converted into "under 3 years old."

[0078] The information processing device 1 generates a noise-added data group for each group G to which noise at a higher noise level has been added for a user data group in the group G according to usage conditions with a higher data protection level AL. For example, the information processing device 1 increases the addition amount, which is the amount of data to be added as noise, or increases the change amount, which is the amount of data to be changed as noise, for a group G with a higher data protection level AL.

[0079] In the generalization process, the information processing device 1 can increase the noise level by, for example, increasing the generalization level. The generalization level is higher for the second higher hierarchical level than the first higher hierarchical level, and higher for the third higher hierarchical level than the second higher hierarchical level.

[0080] Furthermore, the information processing device 1 can increase the noise level by, for example, increasing the number of attribute items converted to asterisks "*" in the suppression process. Furthermore, the information processing device 1 can increase the noise level by decreasing the threshold when grouping values ​​equal to or greater than a certain threshold into one category. Furthermore, the information processing device 1 can increase the noise level by increasing the threshold when grouping values ​​equal to or less than a certain threshold into one category.

[0081] For example, protection levels AL1 to AL nIt is assumed that each user U is classified into a group G according to a selected protection level SAL, which is the protection level selected by the user U from among the above. In this case, the information processing device 1 adds noise with a higher noise level to groups G with higher data protection levels AL.

[0082] The information processing device 1 can also perform a dummy addition process to add dummy data of a number or content according to the usage conditions with a high data protection level AL to the user data group for each group G. The dummy data may be, for example, the same as the data included in the user data group for each group G, or may be data obtained by processing the data included in the user data group. Furthermore, the dummy data may be data generated randomly or according to a predetermined rule.

[0083] For example, the higher the data protection level AL, the more dummy data the information processing device 1 can add, or use data that has been processed to a higher degree than data included in the user data group as dummy data.

[0084] The information processing device 1 can also add dummy data to the user data group so as not to change the overall trend of each group G. For example, the information processing device 1 adds dummy data to the user data group so that the calculation results (e.g., statistical quantities such as average, sum, and variance) based on the user data group before and after adding the dummy data are within a threshold range so as not to change the overall trend of each group G.

[0085] The information processing device 1 can also add noise to the user data group using a different noise addition method for each protection level AL. Noise addition methods include, for example, suppression processing, generalization processing, and dummy addition processing, and suppression processing also includes noise addition methods such as conversion to asterisks "*", top coating, and bottom coating. Note that the noise addition method is not limited to the above examples.

[0086] In this way, the information processing device 1 generates a noise-added data group for each group G by adding noise according to the group G to the user data group for each group G in the anonymization processing using k-anonymization.

[0087] Next, the information processing device 1 generates use target data based on the noise-added data for each group G or the noise-added data group for each group G generated by anonymization in step S8 (step S9).

[0088] First, a case will be described where noise-added data for each group G is generated by anonymization in step S8. In this case, the information processing device 1 generates data to be used by weighting and adding the noise-added data for each group G. For weighting, a weight according to a value obtained by adding noise to the number of users included in the user data group for each group G, or a weight according to a value obtained by dividing the value obtained by adding noise to the number of users included in the user data group for each group G by the value obtained by adding noise to the variance for each group G, etc. is used.

[0089] For example, groups G1, G2, . . ., G n The calculation results are Ms1, Ms2, , Ms n Let G1, G2, , G n The noise of Mn1, Mn2, , Mn n Also, let G1, G2, . . ., G n The weights of k1,k2,...,k n Let's say.

[0090] In this case, the information processing device 1 can calculate the value Da expressed by the following formula (1) as the value of the use target data. Da=k1(Ms1+Mn1)+k2(Ms2+Mn2)+... +k n (Ms n +Mn n ) ···(1)

[0091] Here, groups G1, G2, . . ., Gn Let the number of users with noise be Ng1, Ng2, , Ng n Let the number of users with noise be Ng1, Ng2, , Ng n The total number of users with noise is Ng1, Ng2,...,Ng n is the group G1, G2, , G n The number of users is the number of users included in the user data group G to which noise according to group G has been added, and is generated by the same noise addition process as for the noise-added data described above.

[0092] In this case, the weights k1, k2, , k n is the number of users Ng1 / Nt,Ng2 / Nt,···,Ng n / Nt. That is, k1 = Ng1 / Nt, k2 = Ng2 / Nt, and k n =Ng n / Nt.

[0093] In this way, the information processing device 1 can generate the use target data by performing weighted addition of the noise-added data for each group G using a weight based on the number of users included in the user data set for each group G.

[0094] Also, groups G1, G2, . . ., G n The noisy variance of is σ1 2 , σ2 2 , , σ n 2 Let the noisy variance σ1 2 , σ2 2 , , σ n 2 is the noise-added calculation result data in which noise according to the group G is added to the variance of the values ​​indicated by the data included in the user data group for each group G. For example, the noise-added variance σ1 2 , σ2 2 , , σ n 2 When the statistical processing target is the annual income of user U, σ is the variance of the annual income of user U for each group G to which noise according to group G is added.

[0095] Weights k1,k2,···,k n is the number of users with noise Ng1,Ng2,...,Ng n Each of the groups G is assigned a noise variance σ1 2 , σ2 2 , , σ n 2 In this case, the weights k1, k2, . . . , k n For example, k1=Ng1 / Nt·σ1 2 , k2=Ng2 / Nt·σ2 2 ,···,k n =Ng n / Nt·σ n 2 It is expressed as follows.

[0096] In this way, the information processing device 1 can generate data to be used by weighting and adding the noise-added data for each group G using a weight corresponding to the value obtained by dividing the number of noise-added users for each group G by the noise-added variance for each group G.

[0097] In addition, when the average value of the above-mentioned multiple gradients is used as the calculation result for each group G, the information processing device 1 calculates the parameters of the learning model determined by machine learning using the data of the calculation results with noise as the data to be used.

[0098] Furthermore, the information processing device 1 calculates the number of noise-added users Ng1, Ng2, . . . , Ng n Alternatively, the noise-added data for each group G can be weighted and added using the number of users for each group G to which noise is not added. The information processing device 1 can also weight and add the noise-added data for each group G using the number of users for each group G to which noise is not added and the variance. The variance is σ1 2 , σ2 2 , , σ n 2 It may be the square root of

[0099] The information processing device 1 can selectively use a first weighted averaging method using a weight according to the number of users with noise, and a second weighted averaging method using a weight according to the value obtained by dividing the number of users with noise by the noise variance.

[0100] The information processing device 1 compares the variations in data included in the user data groups between groups G, and selects one of the first weighted averaging method and the second weighted averaging method based on the comparison result. Then, the information processing device 1 performs weighted addition on the noise-added data for each group G using the selected weighted averaging method.

[0101] For example, the information processing device 1 calculates the variation of data contained in a user data group for each group G, and selects a first weighted average method when the trend similarity, which is the similarity in the trends of data variation between groups G, is high, and selects a second weighted average method when the trend similarity is low.

[0102] For example, the information processing device 1 determines that the trend similarity is high when the variance of the variances for each group G or the standard deviation of the variances for each group G is less than a threshold, and determines that the trend similarity is low when the variance of the variances for each group G or the standard deviation of the variances for each group G is greater than or equal to the threshold.

[0103] Furthermore, the information processing device 1 can determine that the trend similarity is high when the variance of the standard deviations for each group G or the standard deviation of the standard deviations for each group G is less than a threshold, and can determine that the trend similarity is low when the variance of the standard deviations for each group G or the standard deviation of the standard deviations for each group G is equal to or greater than a threshold. Note that the information processing device 1 only needs to select a weighted average method based on the comparison result of the variations for each group G, and is not limited to the above-mentioned example.

[0104] Next, a method for generating data to be used when a noise-added data group for each group G is generated as an anonymously processed data group in step S8 will be described.

[0105] The information processing device 1 calculates data of the operation result (for example, statistical quantities such as average, sum, variance, etc.) based on the anonymously processed data group based on the processing content indicated by the processing content information. The processing content indicated by the processing content information includes information indicating the statistical operation type such as average, sum, variance, etc., and the operation type for generating parameters of a learning model.

[0106] For example, assume that the processing content specified by the processing target information is the average annual income of user U. In this case, the information processing device 1 calculates, as the data of the calculation result based on the anonymously processed data group, data indicating the average annual income of user U included in the anonymously processed data group as the data to be used.

[0107] Furthermore, it is assumed that the processing content specified by the processing target information is the variance of the ages of the user U. In this case, the information processing device 1 calculates, as the data of the calculation result based on the anonymized processed data group, data indicating the variance of the ages of the user U included in the anonymized processed data group as the data to be used.

[0108] Furthermore, it is assumed that the processing content specified by the processing target information is parameters of a learning model generated using the anonymized data as learning data. In this case, the information processing device 1 calculates, as the target data to be used, parameters of the learning model determined by machine learning using average values ​​of multiple gradients resulting from training data samples obtained from the anonymized data group when optimizing the model using, for example, stochastic gradient descent.

[0109] Next, the information processing device 1 transmits the use target data generated in step S9 to the requester device 4 as provided information, thereby providing the provided information to the operator (step S10).

[0110] In this way, the information processing device 1 classifies data whose use is permitted by the user U into groups G according to the usage conditions set by the user U, and generates, for each group G, noise-added data by adding noise according to the group G to data resulting from an operation based on a user data group including a plurality of data classified into the same group G, or a noise-added data group by adding noise according to the group G to the user data group. Then, the information processing device 1 generates data to be used using the noise-added data or noise-added data group for each group G. In this way, by adding noise according to the usage conditions set by the user U, the information processing device 1 can use the data of the user U who has given permission to use personal information depending on the protection level, and can make more personal information usable.

[0111] Furthermore, the information processing device 1 receives a request to provide data and determines whether the candidate data to be provided, which is data specified in the request, is data linked to a specific related organization that has a predetermined relationship with the requester of the request. If the information processing device 1 determines that the candidate data to be provided is data linked to an organization that has a predetermined relationship with the requester, the information processing device 1 anonymizes the candidate data to be provided to generate anonymized data, and provides data in response to the request based on the generated anonymized data. This allows the information processing device 1 to provide personal information more appropriately.

[0112] The configuration of an information providing system including an information processing device 1, a terminal device 2, a server 3, and a requester device 4 that performs such processing will be described in detail below.

[0113] [2. Information provision system configuration] 2 is a diagram showing an example of the configuration of an information providing system according to an embodiment. As shown in FIG. 2, the information providing system 100 according to an embodiment includes an information processing device 1, a plurality of terminal devices 2, and a plurality of servers 31 to 33. m and a requester device 4. n is, for example, an integer of 3 or greater.

[0114] Information processing device 1, multiple terminal devices 2, multiple servers 31-3 m , and the requester device 4 are connected to each other via a network N so as to be able to communicate with each other via a wired or wireless connection. Note that the information providing system 100 shown in FIG. 2 may include a plurality of information processing devices 1 and a plurality of requester devices 4.

[0115] The terminal device 2 is, for example, a desktop PC (Personal Computer), a notebook PC, a tablet terminal, a smartphone, a mobile phone, or a PDA (Personal Digital Assistant), etc. Note that the terminal device 2 is not limited to the above examples and may be, for example, a smart watch or a wearable device.

[0116] Server 31~3 m Each of the servers 31 to 3 provides online services to the user U of the terminal device 2. m are servers that provide online services provided by different organizations, but may also be servers that provide online services that are partly provided by the same organization.

[0117] If the organization is a business, the online services provided by the organization may be, for example, online services such as a shopping site, news site, auction site, flea market site, weather forecast site, finance (stock price) site, etc. Furthermore, the online services provided by the business may also be, for example, online services such as a search site, map providing site, travel site, restaurant introduction site, weblog site, or SNS site.

[0118] If the organization is a local government, the online services provided by the organization are administrative services, such as online services such as electronic applications, electronic bidding, electronic tax returns, electronic tax payments, reservations for public facilities, reservations for book lending, and provision of various information related to administrative services.

[0119] The requester device 4 is, for example, a desktop PC or a notebook PC, but may also be a tablet terminal, a smartphone, a mobile phone, or a PDA. The operator who operates the requester device 4 operates the requester device 4 to obtain target data for use, which is obtained by calculation from data on service logs of the organization to which the operator belongs or an organization that has requested the operator's organization to analyze.

[0120] 3. Configuration of Information Processing Device 1 3 is a diagram showing an example of the configuration of the information processing device 1 according to the embodiment. As shown in FIG. 3, the information processing device 1 includes a communication unit 10, a storage unit 11, and a processing unit 12.

[0121] [3.1. Communication Unit 10] The communication unit 10 is realized by, for example, a NIC. The communication unit 10 is connected to a network N by wire or wirelessly, and transmits and receives information to and from various other devices. For example, the communication unit 10 transmits and receives information to and from the terminal device 2 via the network N.

[0122] [3.2. Storage section 11] The storage unit 11 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 11 has an organization information storage unit 20 and a service log storage unit 21.

[0123] [3.2.1. Organizational information storage unit 20] The organization information storage unit 20 stores an organization information table including information for each organization such as a company, a local government, etc. Fig. 4 is a diagram showing an example of the organization information table stored in the organization information storage unit 20 of the information processing device 1 according to the embodiment.

[0124] In the example shown in FIG. 4, the organization information table stored in the organization information storage unit 20 includes information items such as "organization ID," "organization name," and "related organization ID." "Organization ID" is an identifier that identifies an organization. "Organization name" is information that indicates the name of the organization. "Related organization ID" is the organization ID of an organization that has a predetermined relationship with the organization indicated by the "organization ID."

[0125] In the example shown in Figure 4, the organization with organization ID "O1" has the organization name "Organization A," and the organization with which it has a predetermined relationship is Organization B with organization ID "O1," the organization with organization ID "O2" has the organization name "Organization B," and the organization with which it has a predetermined relationship is Organization C with organization ID "O3." Also, the organization with organization ID "O1" has the organization name "Organization C," and the organizations with which it has a predetermined relationship are Organizations A and B with organization IDs "O1, O2."

[0126] The organization information table stored in the organization information storage unit 20 may have a "related organization estimation model" instead of the "organization ID." The "related organization estimation model" is information about the related organization estimation model, such as information about the parameters of the related organization estimation model.

[0127] 3.2.2. Service Log Storage Unit 21 The service log storage unit 21 stores a service log information table including information such as service log data of online services provided by an organization. Fig. 5 is a diagram showing an example of the service log table stored in the service log storage unit 21 of the information processing device 1 according to the embodiment.

[0128] In the example shown in Fig. 5, the service log table stored in the service log storage unit 21 includes information items such as "organization ID," "service log," and "usage conditions information." "Organization ID" is an identifier that identifies an organization. "Service log" is information that indicates the service log data of an online service provided by the organization.

[0129] The "usage conditions information" is information indicating the usage conditions selected by each user U who conditionally authorizes the use of the data corresponding to the user among the multiple data included in the service log data. The information indicating the usage conditions is, for example, protection levels AL1, AL2, . . ., AL n The protection level is indicated by a protection level selected by the user U from among the above, and is associated with, for example, a user ID included in the service log data.

[0130] 5, the organization with organization ID "O1" has service log data "DA1" and usage condition information "RA1", the organization with organization ID "O2" has service log data "DA2" and usage condition information "RA2", and the organization with organization ID "O3" has service log data "DA3" and usage condition information "RA3".

[0131] In the example shown in FIG. 5, the service log data is expressed using abstract symbols such as "DA1" to "DA3", but the service log data is, for example, data in a file format that includes data of each user U.

[0132] Furthermore, in the example shown in FIG. 5, the usage condition information is expressed using abstract codes such as "RA1" to "RA3", but the usage condition information is, for example, data in a file format that includes data containing information for each user U, including the user ID and selected protection level SAL of each user U. The usage condition information may also be included in a "service log". For example, information on the protection level AL may be included as part of the data of user U included in the service log data.

[0133] [3.3. Processing Unit 12] The processing unit 12 is a controller, and is realized by a processor such as a CPU or an MPU using RAM as a work area to execute various programs (an example of an information provision program) stored in a storage device (e.g., the storage unit 11) inside the information processing device 1. Also, the processing unit 12 may be partially or entirely realized by an integrated circuit such as an ASIC or an FPGA.

[0134] 3, the processing unit 12 has a receiving unit 30, a linking unit 31, a determining unit 32, an anonymization unit 33, and a providing unit 34, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of the processing unit 12 is not limited to the configuration shown in FIG. 3, and may have any other configuration as long as it performs the information processing described below.

[0135] 3.3.1. Reception unit 30 The receiving unit 30 receives various data and queries. The receiving unit 30 includes a data receiving unit 40 and a provision request receiving unit 41.

[0136] 3.3.1.1. Data Receiving Unit 40 The data receiving unit 40 receives data transmitted from various devices and stores the received data in the storage unit 11 .

[0137] For example, the data accepting unit 40 accepts service log data uploaded from each server 3 to the information processing device 1, and stores the accepted data in the storage unit 11. The data accepting unit 40 also accepts data uploaded from an information providing device (not shown), and stores the accepted data in the storage unit 11.

[0138] [3.3.1.2. Provision request reception unit 41] The provision request receiving unit 41 receives a data provision request. For example, the provision request receiving unit 41 receives a query transmitted from the requester device 4 as a data provision request.

[0139] The query from the requester device 4 includes, for example, data identification information for identifying the data to be used or the data to be processed, and organization identification information such as the organization ID of the requesting organization. The requesting organization is, for example, the organization to which the operator operating the requester device 4 belongs or the organization that has requested the analysis from the organization to which the operator belongs.

[0140] The data identification information includes target organization information, processing target information, etc. The target organization information includes the organization ID of the target organization, which is the organization that owns the data to be used. For example, if the data that the requesting organization wants to analyze is service log data of an online service provided by organization B, the organization ID of the target organization is the organization ID of organization B.

[0141] The processing target information includes data type information indicating the type of processing target data, which is data to be processed, processing content information indicating the processing content of the processing target data, etc. The data type information is, for example, information indicating the location of a user U who used an online service, information indicating the type of usage history of the online service by the user, information indicating the type of attributes of the user, and information indicating the type of input information and output information of a learning model.

[0142] The information indicating the location of the user is, for example, information indicating the location detected by a location detection unit (not shown) provided in the terminal device 2, and is information transmitted from the terminal device 2 to the server 3 when the user uses an online service.

[0143] The usage history of the online service includes the date and time of use of the online service, the usage content of the online service, etc. For example, if the online service is a service provided by a shopping site, the usage content of the online service includes the content of browsing by the user of the transaction target, the content of purchases by the user of the transaction target, the content of evaluations by the user of the transaction target, and the content of favorite registrations by the user of the transaction target.

[0144] For example, if the online service is an administrative service, the usage content of the online service may include the content of electronic applications made by the user, the content of electronic bidding made by the user, the content of electronic declarations made by the user, the content of electronic tax payments made by the user, the content of public facility reservations made by the user, the content of book lending reservations made by the user, and the viewing of various information by the user.

[0145] The attributes of user U are at least one of demographic attributes and psychographic attributes. Demographic attributes are demographic attributes, such as age, gender, occupation, place of residence, annual income, and family structure. Psychographic attributes are psychological attributes, such as lifestyle, values, and interests.

[0146] The processing content information includes, for example, information indicating the type of statistical calculation such as average, sum, or variance, the type of calculation for generating parameters of a learning model, etc. The processing content specified by the processing target information is, for example, calculation of the average annual income of men in their twenties in Kyoto Prefecture, the average variance of people who like motorcycles, the total number of users U who have made electronic tax payments, or generation of a learning model that inputs information indicating the location and attributes of user U and outputs a score indicating the possibility of user U using online services.

[0147] [3.3.2. Linking section 31] The linking unit 31 links the provided data received by the data receiving unit 40 to an organization among a plurality of organizations that satisfies the linking conditions. The provided data includes, for example, service log data, that is, data on the usage log of online services.

[0148] An organization that satisfies the linking condition is, for example, an organization that provides an online service or an organization that provided the service log data. For example, if server 31 is a server of an organization managed by an organization that provides an online service, linking unit 31 links the service log data transmitted from server 31 to the organization that manages server 31.

[0149] Furthermore, the linking unit 31 may, for example, m Server 3 m If the server provides online services to an organization other than the one that manages it, Server 3 m The service log data sent from Server 3 m This links the online service to the organization that provides it.

[0150] If the information sent from the server 3 includes providing organization information indicating an organization that provides online services, the linking unit 31 can determine that the organization indicated by the providing organization information is an organization that satisfies the linking conditions.

[0151] Furthermore, an organization that satisfies the linking condition may be an organization for which the score output from the relevance estimation model satisfies a preset condition. The relevance estimation model is, for example, a learning model that receives service log data as input and outputs a score indicating the relevance of each organization with the source of the service log data request.

[0152] The relevance estimation model is generated by using service log data that has been determined in advance to be relevant to each organization as the correct answer data for each organization and having the model learn the features of the correct answer data for each organization. The relevance estimation model is a learning model common to multiple organizations, but it may also be a learning model for each organization.

[0153] The linking unit 31 inputs the service log data into the relevance estimation model, and identifies organizations whose scores satisfy predetermined conditions among the scores for each organization output from the relevance estimation model as organizations whose scores output from the relevance estimation model satisfy the preset conditions. The score that satisfies the predetermined conditions is, for example, the highest score, but may also be, for example, a score equal to or greater than a threshold. If there are multiple organizations with scores equal to or greater than the threshold, the linking unit 31 can link the service log data to multiple organizations.

[0154] The relevance estimation model may be, for example, a learning model generated by GBDT or a learning model generated by deep learning using a deep neural network, but is not limited to such examples and may also be a learning model generated by other machine learning methods.

[0155] For example, the relevance estimation model may be generated using machine learning with other learning algorithms, such as a regression method such as linear regression, multiple regression, or logistic regression, or a learning algorithm such as a support vector machine.

[0156] For example, the linking unit 31 can store service log data that has been made unavailable without linking it to an organization. The linking unit 31 can determine whether the service log data has been made unavailable based on information indicating whether the service log data is available or unavailable, which is included in the information transmitted from the server 3.

[0157] [3.3.3. Judgment unit 32] The determination unit 32 determines whether the candidate data to be provided, which is data specified by a query that is a provision request received by the provision request receiving unit 41, is data linked to a specific related organization. A specific related organization is an organization that has a predetermined relationship with the requester of the provision request. The requester of the provision request is the organization indicated by the organization ID indicated in the organization identification information included in the query.

[0158] For example, when the target organization indicated by the organization ID included in the target organization information is an organization that is in a predetermined competitive relationship with the requester of the provision request, the determination unit 32 determines that the provision candidate data is data of a specific related organization. For example, the determination unit 32 has a relationship table that includes, for each organization, information indicating organizations that are in a predetermined competitive relationship, and uses the relationship table to determine whether the provision candidate data is data of a specific related organization.

[0159] If the information sent from the server 3 includes authorized organization information indicating an organization that is permitted to use the service log data, the determination unit 32 can also determine that the organization indicated by the authorized organization information is an organization in a predetermined competitive relationship.

[0160] The determination unit 32 can also use information about competing organizations as correct data and estimate whether an organization is in a competitive relationship using a related organization estimation model that has learned the characteristics of the correct data about competing organizations. The related organization estimation model is a learning model for each organization that has requested the provision request. In this case, the determination unit 32 inputs information about each organization into the related organization estimation model of the organization that has requested the provision request, and determines that an organization whose score output from the related organization estimation model is equal to or greater than a threshold is a competing organization.

[0161] In the above example, the specific related organization is an organization that is in a predetermined competitive relationship with the requester of the provision request, but the specific related organization is not limited to this example. For example, the determination unit 32 can also determine an organization that is not in a predetermined competitive relationship with the requester of the provision request as a specific related organization.

[0162] For example, the determination unit 32 has a relationship table that includes, for each organization, information indicating organizations that are not in a predetermined competitive relationship, and uses this relationship table to determine the specific related organizations.

[0163] In addition, the judgment unit 32 can, for example, use information about organizations that are not in a competitive relationship as correct data and determine specific related organizations using a related organization estimation model that has learned the characteristics of the correct data of organizations that are not in a competitive relationship.

[0164] If the determination unit 32 determines that the candidate data to be provided is data linked to a specific related organization, it identifies the type of data that is linked to the specific related organization and corresponds to the query received by the provision request receiving unit 41.

[0165] As described above, the query received by the provision request receiving unit 41 includes the processing target information, and the determination unit 32 identifies the type of target data, which is data corresponding to the query, based on the data type information included in the processing target information. For example, if the data type information is information indicating the annual income of men in their twenties living in Kyoto Prefecture, the determination unit 32 identifies the annual income of men in their twenties living in Kyoto Prefecture as the type of target data.

[0166] Furthermore, when the data type information includes information indicating the types of input information and output information of the learning model, the determination unit 32 identifies the types of input information and output information of the learning model as the types of target data. The input information is, for example, information indicating the location and attributes of the user U, and the output information is, for example, information indicating the usage content of the online service by the user U, but is not limited to such examples.

[0167] The determination unit 32 can also accept designation of data such that the number of target data or the number of users U who are the source of the target data is equal to or greater than a threshold. For example, the determination unit 32 can send a list of attributes for which the number of target data or the number of users U who are the source of the target data is equal to or greater than a threshold to the requester device 4, thereby allowing the operator to select an attribute from the list of attributes.

[0168] [3.3.4. Anonymous processing section 33] The anonymization unit 33 performs anonymization on the candidate data to be provided to generate anonymously processed data when the determination unit 32 determines that the candidate data to be provided is data linked to a specific related organization. The anonymization unit 33 includes a classification unit 50 and a noise addition processing unit 51.

[0169] [3.3.4.1. Classification section 50] The classification unit 50 classifies each piece of data included in the data to be anonymized into a group G according to the conditions of use set by the user U. Each piece of data included in the data to be anonymized is data whose use has been permitted by the user U.

[0170] For example, the usage conditions set by a user U may be set to multiple protection levels AL1, AL2, . . . , AL n In this case, the classification unit 50 classifies the protection levels AL1, AL2, . . . , AL n corresponding to multiple groups G1, G2, , G n Classify the data included in the data to be anonymized as follows.

[0171] For example, data for which a usage condition indicated by protection level AL1 is set is classified into group G1, data for which a usage condition indicated by protection level AL2 is set is classified into group G2, and data for which a usage condition indicated by protection level AL3 is set is classified into group G3. n Data for which the terms of use indicated by are set is in Group G n It is classified as follows.

[0172] [3.3.4.2. Noise Addition Processing Unit 51] The noise addition processing unit 51 generates noise-added data or a noise-added data group for each group G based on a user data group for each group G including a plurality of data classified into the same group G by the classification unit 50.

[0173] The noise-added data is data obtained by adding noise according to the group G to data of a calculation result (for example, a statistical quantity such as an average, sum, or variance) based on a user data group including a plurality of pieces of data classified into the same group G by the classification unit 50. Furthermore, the noise-added data group is a data group obtained by adding noise according to the group G to a user data group including a plurality of pieces of data classified into the same group G by the classification unit 50.

[0174] The noise addition processing unit 51 generates noise-added data or a noise-added data group by anonymization processing using differential privacy or k-anonymization. For example, when noise addition type information indicating the noise addition type is included in the information transmitted from the server 3, the noise addition processing unit 51 can perform the anonymization processing indicated by the noise addition type information, either the anonymization processing using differential privacy or the anonymization processing using k-anonymization.

[0175] Furthermore, for example, when the query received by the provision request receiving unit 41 includes noise addition type information indicating the noise addition type, the noise addition processing unit 51 can perform anonymization indicated by the noise addition type information, either anonymization using differential privacy or anonymization using k-anonymization.

[0176] First, anonymization using differential privacy will be described. The noise addition processing unit 51 calculates data of the calculation result for each group G from the user data group for each group G based on the processing content indicated in the processing content information included in the query accepted by the provision request accepting unit 41. The processing content indicated in the processing content information is, for example, statistical calculations such as average, sum, and variance for specific data, or generation of parameters for a learning model.

[0177] For example, assume that the processing content indicated by the processing content information is the average annual income of user U. In this case, the noise addition processing unit 51 calculates data indicating the average annual income of user U for each group G as data of the calculation result for each group G.

[0178] Furthermore, it is assumed that the processing content indicated by the processing content information is the age distribution of users U. In this case, the noise addition processing unit 51 calculates data indicating the distribution or variance of the ages of users U for each group G as data of the calculation result for each group G.

[0179] Furthermore, suppose that the processing content indicated by the processing content information is generation of parameters for a learning model. In this case, the noise addition processor 51 calculates, for each group G, data indicating the average value of multiple gradients resulting from training data samples when optimizing the learning model using stochastic gradient descent.

[0180] Furthermore, in addition to the processing content indicated by the processing content information, the noise addition processing unit 51 calculates, as calculation result data, data indicating the number (total) of users U for each group G and data indicating the variance for each group G to be statistically processed. Hereinafter, the number of users U may be referred to as the number of users.

[0181] Then, the noise addition processing unit 51 generates noise-added data for each group G by adding noise according to the group G to the data of the calculation result for each group G. For example, n corresponding to multiple groups G1, G2, , Gn Assume that the data contained in the data to be anonymized is classified.

[0182] In this case, the noise addition processing unit 51 adds the noise to the groups G1, G2, . . . , G n Noise with increasing noise level is added to the data resulting from the calculation in this order. In this way, noise-added data is generated for each group G, with noise having a higher noise level added to the group G according to the usage conditions with a higher data protection level AL.

[0183] The noise is, for example, fixed noise, Gaussian noise, Laplace noise, etc. When the noise to be added to the data of the calculation result is fixed noise, the noise addition processing unit 51 increases the value of the fixed noise as the noise level increases.

[0184] Furthermore, when the noise added to the calculation result data is Gaussian noise, the noise addition processing unit 51 can increase the noise level by increasing the absolute value of the average value of the noise or by increasing the standard deviation of the noise. Furthermore, when the noise added to the calculation result data is Laplace noise, the noise addition processing unit 51 can increase the noise level by increasing the absolute value of the average value of the noise or by increasing the standard deviation of the noise.

[0185] In this way, the noise addition processing unit 51 generates noise-added data for each group G by adding noise according to the group G to the data of the calculation result for each group G in the anonymization processing using differential privacy.

[0186] Next, anonymization using k-anonymization will be described. The noise addition processing unit 51 performs anonymization by, for example, performing suppression processing to add noise by converting a certain attribute item into an asterisk "*" or the like, or by performing generalization processing to add noise by raising the hierarchy of the attribute item.

[0187] The noise addition processing unit 51 can also add noise by top coding, which groups values ​​equal to or greater than a certain threshold into one category, or by bottom coding, which groups values ​​equal to or less than a certain threshold into one category.

[0188] The noise addition processing unit 51 generates a noise-added data group for each group G, to which noise at a higher noise level has been added, for user data groups in the group G according to usage conditions with a higher data protection level AL. For example, the noise addition processing unit 51 increases the addition amount, which is the amount of data to be added as noise, or increases the change amount, which is the amount of data to be changed as noise, for groups G with a higher data protection level AL.

[0189] In the generalization process, the noise addition processing unit 51 can increase the noise level by, for example, increasing the generalization level. The generalization level is higher for the layer two levels above than for the layer one level above, and higher for the layer three levels above than for the layer two levels above.

[0190] Furthermore, the noise addition processing unit 51 can increase the noise level by, for example, increasing the number of attribute items converted to asterisks "*" in the suppression processing. Furthermore, the noise addition processing unit 51 can increase the noise level by decreasing the threshold when grouping values ​​equal to or greater than a certain threshold into one category. Furthermore, the noise addition processing unit 51 can increase the noise level by increasing the threshold when grouping values ​​equal to or less than a certain threshold into one category.

[0191] For example, protection levels AL1 to AL n It is assumed that each user U is classified into a group G according to a selected protection level SAL, which is the protection level selected by the user U from among the above. In this case, the noise addition processing unit 51 adds noise at a higher noise level to a group G having a higher data protection level AL.

[0192] The noise addition processing unit 51 can also perform dummy addition processing to add dummy data of a number or content according to the usage conditions with a high data protection level AL to the user data group for each group G. The dummy data may be, for example, the same as the data included in the user data group for each group G, or may be data obtained by processing the data included in the user data group. Furthermore, the dummy data may be data generated randomly or according to a predetermined rule.

[0193] For example, the higher the data protection level AL, the more dummy data the noise addition processing unit 51 can add, or the more highly processed data included in the user data group can be used as the dummy data.

[0194] The noise addition processing unit 51 can also add dummy data to the user data group so as not to change the overall trend of each group G. For example, the noise addition processing unit 51 adds dummy data to the user data group so that the calculation results (e.g., statistical quantities such as average, sum, and variance) based on the user data group before and after adding the dummy data are within a threshold range so as not to change the overall trend of each group G.

[0195] The noise addition processing unit 51 can also add noise to the user data group using a different noise addition method for each protection level AL. Noise addition methods include, for example, suppression processing, generalization processing, and dummy addition processing, and suppression processing also includes noise addition methods such as conversion to asterisks "*", top coating, and bottom coating. Note that the noise addition method is not limited to the examples described above.

[0196] In this way, the noise addition processing unit 51 generates a noise-added data group for each group G by adding noise according to the group G to the user data group for each group G in the anonymization processing using k-anonymization.

[0197] [3.3.5.Providing Department 34] The providing unit 34 provides data in response to the provision request received by the provision request receiving unit 41 as data to be analyzed, based on the anonymized data or a group of anonymized data generated by the anonymization unit 33.

[0198] The providing unit 34 includes a target data generating unit 60 that generates target data for use based on the anonymously processed data or a group of anonymously processed data generated by the anonymizing unit 33, and a providing processing unit 61 that provides the target data for use generated by the target data generating unit 60.

[0199] 3.3.5.1. Use target data generation unit 60 The target data generation unit 60 generates target data to be used using the anonymously processed data or a group of anonymously processed data generated by the anonymization unit 33. The anonymously processed data is, as described above, noise-added data or a group of noise-added data for each group G.

[0200] The utilization target data generation unit 60 includes a comparison unit 70 , a selection unit 71 , a weighted average calculation unit 72 , and a calculation processing unit 73 .

[0201] 3.3.5.1.1. Comparison unit 70 The comparison unit 70 compares the variations in the data included in the user data groups between groups G. For example, the comparison unit 70 calculates the variations in the data included in the user data groups for each group G, and calculates a trend similarity, which is the similarity in the trends of the data variations between groups G.

[0202] The trend similarity is, for example, a value indicated by the variance of the variances for each group G or the standard deviation of the variances for each group G, but is not limited to such examples. For example, the trend similarity may be a value indicated by the variance of the standard deviations for each group G or the standard deviation of the standard deviations for each group G.

[0203] [3.3.5.1.2. Selection unit 71] The selection unit 71 selects one weighted averaging method from among the plurality of weighted averaging methods based on the comparison result by the comparison unit 70. The plurality of weighted averaging methods include a first weighted averaging method and a second weighted averaging method.

[0204] The first weighted average method is a weighted average method using a weight according to the number of noise-bearing users for each group G, and the second weighted average method is a weighted average method using a weight according to the value obtained by dividing the number of noise-bearing users by the noise-bearing variance for each group G.

[0205] For example, the selection unit 71 selects the first weighted averaging method when the trend similarity is high, and selects the second weighted averaging method when the trend similarity is low. For example, the selection unit 71 determines that the trend similarity is high when the variance of the variances for each group G or the standard deviation of the variances for each group G is less than a threshold, and determines that the trend similarity is low when the variance of the variances for each group G or the standard deviation of the variances for each group G is equal to or greater than the threshold.

[0206] Furthermore, the selection unit 71 can determine that the trend similarity is high when the variance of the standard deviations for each group G or the standard deviation of the standard deviations for each group G is less than a threshold, and can determine that the trend similarity is low when the variance of the standard deviations for each group G or the standard deviation of the standard deviations for each group G is equal to or greater than a threshold. Note that the selection unit 71 is not limited to the above-described example as long as it selects a weighted averaging method based on the comparison result of the variations for each group G.

[0207] [3.3.5.1.3. Weighted average calculation unit 72] The weighted average calculation unit 72 generates data to be used based on the anonymously processed data when the anonymously processed data generated by the anonymization unit 33 is noise-added data for each group G. For example, the weighted average calculation unit 72 can generate data to be used by calculating the above formula (1).

[0208] The weighted average calculation unit 72 generates data to be used by weighting and adding the noise-added data for each group G. For the weighting, a weight according to a value obtained by adding noise to the number of users included in the user data group for each group G, or a weight according to a value obtained by dividing the value obtained by adding noise to the number of users included in the user data group for each group G by the value obtained by adding noise to the variance for each group G, or the like is used.

[0209] For example, when the selection unit 71 selects the first weighted average method, the weighted average calculation unit 72 generates the data to be used by weighting and adding the noise-added data for each group G using a weight corresponding to the value obtained by adding noise to the number of users included in the user data group for each group G (number of users with noise).

[0210] In addition, when the second weighted average method is selected by the selection unit 71, the weighted average calculation unit 72 generates data to be used by weighting and adding the noise-added data for each group G using a weight corresponding to the value obtained by dividing the number of noise-added users for each group G by the noise-added variance for each group G.

[0211] In addition, when the weighted average calculation unit 72 uses the average value of the above-mentioned multiple gradients as the calculation result for each group G, it calculates the parameters of the learning model determined by machine learning using the data of the calculation results with noise as the data to be used.

[0212] Moreover, the weighted average calculation unit 72 can also weight and add the noise-added data for each group G using the number of users for each group G to which noise is not added, instead of the number of users with noise. Moreover, the weighted average calculation unit 72 can weight and add the noise-added data for each group G using the number of users for each group G to which noise is not added and the variance. Note that the variance is σ1 2 , σ2 2 , , σ n 2 It may be the square root of

[0213] [3.3.5.1.4. Processing unit 73] When the anonymization unit 33 generates a noise-added data group for each group G, that is, an anonymously processed data group, the calculation processing unit 73 generates data to be used.

[0214] The calculation processing unit 73 calculates data of calculation results (for example, statistical quantities such as average, sum, variance, etc.) based on the anonymously processed data group based on the processing content indicated by the processing content information. The processing content indicated by the processing content information includes information indicating the statistical calculation type such as average, sum, variance, and the calculation type for generating parameters of the learning model.

[0215] For example, suppose the processing content specified by the processing target information is the average annual income of user U. In this case, the calculation processing unit 73 calculates, as the data of the calculation result based on the anonymously processed data group, data indicating the average annual income of user U included in the anonymously processed data group as the data to be used.

[0216] Furthermore, it is assumed that the processing content specified by the processing target information is the variance of the ages of user U. In this case, the calculation processing unit 73 calculates, as the data of the calculation result based on the anonymously processed data group, data indicating the variance of the ages of user U included in the anonymously processed data group as the data to be used.

[0217] Furthermore, it is assumed that the processing content specified by the processing target information is parameters of a learning model generated using the anonymized data as learning data. In this case, the arithmetic processing unit 73 calculates, as the target data to be used, parameters of the learning model determined by machine learning using average values ​​of multiple gradients resulting from training data samples obtained from the anonymized data group when optimizing the model using, for example, stochastic gradient descent.

[0218] [3.3.5.2. Provision Processing Unit 61] The provision processing unit 61 provides, for example, data in response to a provision request accepted by the provision request accepting unit 41.

[0219] For example, the provision processing unit 61 provides the provision information to the operator by transmitting the use target data generated by the use target data generation unit 60 to the requester device 4 as data in response to the provision request. This allows the provision processing unit 61 to provide personal information more appropriately.

[0220] [4. Processing Procedure] Next, the procedure of information processing executed by the processing unit 12 of the information processing device 1 according to the embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of information processing executed by the processing unit 12 of the information processing device 1.

[0221] 6, the processing unit 12 of the information processing device 1 determines whether or not there is upload data (step S20). For example, when service log data is transmitted from the server 3 to the information processing device 1 and received by the communication unit 10, the processing unit 12 determines that there is upload data.

[0222] If the processing unit 12 determines that there is upload data (step S20: Yes), it links the service log data, which is the upload data, to the organization (step S21). Then, the processing unit 12 stores the service log data, which is the upload data linked to the organization, in the storage unit 11 (step S22).

[0223] When the processing of step S22 is completed or when it is determined that there is no upload data (step S20: No), the processing unit 12 determines whether there is a data provision request (step S23). For example, when a query is transmitted from the requester device 4 to the information processing device 1 and received by the communication unit 10, the processing unit 12 determines that there is a data provision request.

[0224] When determining that there is a data provision request (step S23: Yes), the processing unit 12 determines whether or not the provision candidate data, which is data specified by the query, is data linked to a specific related organization (step S24).

[0225] When the processing unit 12 determines that the provision candidate data is data linked to a specific related organization (step S24: Yes), the processing unit 12 performs anonymization processing (step S25). The anonymization processing is the processing of steps S30 and S31 shown in FIG. 7, and will be described in detail later.

[0226] Next, the processing unit 12 generates data to be used based on the anonymously processed data including noise-added data for each group G generated in the anonymization processing in step S25 or an anonymously processed data group including noise-added data groups for each group G (step S26). Then, the processing unit 12 provides the data to be used generated in step S26 to the provider of the provision request (step S27).

[0227] When the processing of step S27 is completed, when it is determined that there is no request to provide data (step S23: No), or when it is determined that the candidate data to be provided is not data linked to a specific related organization (step S24: No), the processing unit 12 determines whether or not it is time to end (step S28). For example, when the power of the information processing device 1 is turned off, or when it is determined that an end operation has been performed by operating an operation unit (not shown) of the information processing device 1, the processing unit 12 determines that it is time to end.

[0228] If the processing unit 12 determines that the end timing has not arrived (step S28: No), it proceeds to step S20, and if it determines that the end timing has arrived (step S28: Yes), it terminates the processing shown in Figure 6.

[0229] Fig. 7 is a flowchart showing an example of anonymization processing executed by the processing unit 12 of the information processing device 1. As shown in Fig. 7, the processing unit 12 classifies each piece of data included in the candidate data to be provided into a group G according to the usage conditions set by the user (step S30).

[0230] Next, the processing unit 12 generates noise-added data or noise-added data groups for each group G by adding noise according to the group G to the data of the calculation result based on the user data group or the user data group (step S31), and ends the processing shown in Figure 7.

[0231] [5. Modifications] In the above example, the data to be anonymized is service log data, but the data to be anonymized is not limited to service log data and may be, for example, data collected offline.

[0232] In the above example, the anonymization unit 33 adds noise according to the protection level indicated in the terms of use set by the user U, but is not limited to this example. For example, when protection level information indicating the protection level of the data in the service log is transmitted from the server 3 to the information processing device 1, the anonymization unit 33 can set the protection level indicated by the protection level information to the lowest protection level.

[0233] For example, the protection levels of the data included in the service log data are AL1, AL2, AL n In this case, if the protection level indicated by the protection level information is protection level AL2, the anonymization unit 33 changes the protection level AL1 set by the user U as a usage condition to the protection level AL2. Also, if the protection level indicated by the protection level information is protection level AL n If so, the protection level AL of the protection level AL set by the user U as a condition of use is n Change to.

[0234] The protection level information may be information indicating a value p for raising the protection level AL. In this case, the anonymization unit 33 sets the protection levels AL1, AL2, . . . , AL as the protection levels of the data included in the service log data. n If it contains, the protection levels AL1, AL2, AL n Protection level AL 1+p ,AL2+p ,···,AL n+p and then anonymize the information.

[0235] [6. Hardware Configuration] The information processing device 1 or terminal device 2 according to the above-described embodiment is realized by, for example, a computer 80 configured as shown in Fig. 8. The following description will be given taking the information processing device 1 as an example. Fig. 8 is a hardware configuration diagram showing an example of a computer 80 that realizes the functions of the information processing device 1 according to the embodiment. The computer 80 has a CPU 81, a RAM 82, a ROM (Read Only Memory) 83, an HDD (Hard Disk Drive) 84, a communication interface (I / F) 85, an input / output interface (I / F) 86, and a media interface (I / F) 87.

[0236] The CPU 81 operates and controls each part based on programs stored in the ROM 83 or the HDD 84. The ROM 83 stores a boot program executed by the CPU 81 when the computer 80 starts up, programs that depend on the hardware of the computer 80, and the like.

[0237] The HDD 84 stores programs executed by the CPU 81, data used by such programs, etc. The communication interface 85 receives data from other devices via the network N (see FIG. 2) and sends it to the CPU 81, and transmits data generated by the CPU 81 to other devices via the network N.

[0238] The CPU 81 controls output devices such as a display and a printer, and input devices such as a keyboard and a mouse, via the input / output interface 86. The CPU 81 acquires data from the input devices via the input / output interface 86. The CPU 81 also outputs generated data to the output devices via the input / output interface 86.

[0239] The media interface 87 reads a program or data stored in a recording medium 88 and provides it to the CPU 81 via the RAM 82. The CPU 81 loads the program or data from the recording medium 88 onto the RAM 82 via the media interface 87 and executes the loaded program. The recording medium 88 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0240] For example, when the computer 80 functions as the information processing device 1 according to the embodiment, the CPU 81 of the computer 80 executes programs loaded onto the RAM 82 to realize the functions of the processing unit 12. In addition, the HDD 84 stores data in the storage unit 11. The CPU 81 of the computer 80 reads and executes these programs from a recording medium 88, but as another example, the CPU 81 may obtain these programs from another device via the network N.

[0241] [7. Other] Furthermore, among the processes described in the above embodiments, some of the processes described as being performed automatically can also be performed manually. Alternatively, all or some of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0242] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0243] For example, the information processing device 1 described above may be realized by a plurality of server computers, and depending on the function, the configuration can be flexibly changed, such as by calling an external platform using an API or network computing.

[0244] 3 may be held in a storage server or the like, rather than being held by each device. In this case, each device obtains various pieces of information by accessing the storage server.

[0245] [8. Effects] As described above, the information processing device 1 according to the embodiment includes the classification unit 50, the noise addition processing unit 51, and the use target data generation unit 60. The classification unit 50 classifies data whose use is permitted by a user U into groups G according to the use conditions set by the user U. The noise addition processing unit 51 adds noise according to the group G to data resulting from a calculation based on a user data group including a plurality of data classified into the same group G by the classification unit 50, thereby generating noise-added data or noise-added data groups for each group G by adding noise according to the group G to the user data group. The use target data generation unit 60 generates use target data using the noise-added data or noise-added data groups for each group G generated by the noise addition processing unit 51. In this way, by adding noise according to the use conditions set by the user U, the information processing device 1 can use the data of the user U who has given permission to use personal information depending on the protection level, thereby making it possible to utilize more personal information.

[0246] Furthermore, the noise addition processing unit 51 generates noise-added data or noise-added data groups to which a higher noise level is added for groups G according to usage conditions with higher data protection levels. This allows the information processing device 1 to more appropriately add noise according to the usage conditions set by the user U.

[0247] Furthermore, the usage target data generation unit 60 generates, as usage target data, data obtained by weighting the noise-added data for each group G generated by the noise addition processing unit 51 using a weight based on a value obtained by adding noise to the number of users U included in the user data group for each group G. This enables the information processing device 1 to more appropriately add noise according to the usage conditions set by the user U.

[0248] Furthermore, the use target data generation unit 60 generates, as the use target data, data obtained by weighting the noise-added data for each group G generated by the noise addition processing unit 51 using a weight according to the value obtained by dividing the number of users U included in the user data set for each group G by the value obtained by adding noise to the variance for each group G of the data used to calculate the calculation result. This enables the information processing device 1 to more appropriately add noise according to the use conditions set by the user U.

[0249] The usage target data generation unit 60 also includes a comparison unit 70, a selection unit 71, and a weighted average calculation unit 72. The comparison unit 70 compares the variations in data included in the user data groups between groups G. Based on the comparison results by the comparison unit 70, the selection unit 71 selects, as a weighted average method, one of a method using a weight corresponding to a value obtained by adding noise to the number of users U included in the user data group for each group G, and a method using a weight corresponding to a value obtained by dividing a value obtained by adding noise to the number of users U included in the user data group for each group G by a value obtained by adding noise to the variance for each group G of the data used to calculate the calculation results. The weighted average calculation unit 72 calculates a weighted average of the calculation result data for each group G calculated by the noise addition processing unit 51 using the weighted average method selected by the selection unit 71. This allows the information processing device 1 to more appropriately add noise according to the usage conditions set by the user U.

[0250] Furthermore, the data used to calculate the calculation result is data on the attributes of the user U. This allows the information processing device 1 to appropriately protect the attributes of the user U.

[0251] Furthermore, the data used to calculate the calculation result is gradient data used for model learning. This allows the information processing device 1 to appropriately protect the personal information of the user U when calculating parameters used in machine learning.

[0252] The information processing device 1 also includes a providing unit 34 and a receiving unit 30. The providing unit 34 provides the use target data generated by the use target data generating unit 60. The receiving unit 30 receives a query specifying the use content. The use target data generating unit 60 calculates, for each group G, data of the calculation result corresponding to the use content specified in the query. This allows the information processing device 1 to provide personal information more appropriately.

[0253] Furthermore, the usage conditions include at least one of usage conditions for each user U and usage conditions for each data item. This allows the information processing device 1 to make more personal information usable.

[0254] The above describes the embodiments of the present application in detail based on the drawings, but this is merely an example, and the present invention can be implemented in other forms that include the embodiments described in the Disclosure of the Invention section and that have been modified and improved in various ways based on the knowledge of those skilled in the art.

[0255] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, a providing section can be read as providing means or providing circuit. [Explanation of symbols]

[0256] 1. Information processing equipment 2. Terminal Device 3,31~3 m server 4 Requester device 10. Communications Department 11 Storage section 12 Processing section 20 Organization information storage section 21 Service log storage unit 30 Reception 31 Stringing section 32 Judgment section 33 Anonymous processing department 34 Providing Department 40 Data Reception Department 41 Provision Request Reception Department 50 Classification Department 51 Noise addition processing section 60 Target data generation unit 61 Provision Processing Unit 70 Comparison Section 71 Selection section 72 Weighted average calculation section 73 Processing unit 100 Information Provision System

Claims

1. a classification unit that classifies data whose use is permitted by a user into groups according to the use conditions set by the user; a noise addition processing unit that generates noise-added data for each group by adding noise according to the group to data resulting from an operation based on a user data group including a plurality of data classified into the same group by the classification unit; a use target data generation unit that generates use target data using the noise-added data for each group generated by the noise addition processing unit, The usage target data generation unit The noise-added data for each of the groups generated by the noise addition processing unit is weighted-averaged using a weight based on a value obtained by adding noise to the number of users included in the user data group for each of the groups, and the resulting data is generated as the data to be used.

1. An information processing device comprising:

2. The usage target data generation unit The noise-added data for each of the groups generated by the noise addition processing unit is weighted and averaged using a weight corresponding to a value obtained by dividing a value obtained by adding noise to the number of users included in the user data group for each of the groups by a value obtained by adding noise to the variance for each of the groups of data used to calculate the calculation result, and the data obtained is generated as the data to be used.

2. The information processing apparatus according to claim 1, wherein:

3. The usage target data generation unit a comparison unit that compares the variations in data included in the user data groups between the groups; a selection unit that selects, based on a comparison result by the comparison unit, one of a method using a weight according to a value obtained by adding noise to the number of users included in the user data group for each group, and a method using a weight according to a value obtained by dividing the value obtained by adding noise to the number of users included in the user data group for each group by a value obtained by adding noise to the variance for each group of data used to calculate the calculation result; and a weighted average calculation unit that calculates a weighted average of the data of the calculation results for each group calculated by the noise addition processing unit using the weighted average method selected by the selection unit.

3. The information processing apparatus according to claim 2, wherein:

4. The data used to calculate the calculation result is: The data is attribute data of the user.

4. The information processing device according to claim 1, wherein the information processing device is a computer.

5. The data used to calculate the calculation result is: The gradient data used to train the model 4. The information processing device according to claim 1, wherein the information processing device is a computer.

6. a providing unit that provides the use target data generated by the use target data generating unit; a reception unit that receives a query specifying the content of use, The usage target data generation unit Calculating the data of the calculation result corresponding to the usage content specified in the query for each of the groups.

4. The information processing device according to claim 1, wherein the information processing device is a computer.

7. The terms of use are: The user-specific usage conditions include at least one of the user-specific usage conditions and the data-specific usage conditions.

4. The information processing device according to claim 1, wherein the information processing device is a computer.

8. 1. A computer-implemented information processing method, comprising: a classification step of classifying the data licensed for use by a user into groups according to the conditions of use set by the user; a noise addition processing step of generating noise-added data for each group by adding noise according to the group to data resulting from an operation based on a user data group including a plurality of data classified into the same group by the classification step; a use target data generating step of generating use target data using the noise-added data for each group generated by the noise addition processing step, The use target data generation step includes: The noise-added data for each of the groups generated by the noise-adding processing step is weighted-averaged using a weight based on a value obtained by adding noise to the number of users included in the user data group for each of the groups, and the resulting data is generated as the data to be used.

1. An information processing method comprising:

9. a classification step of classifying data whose use is permitted by a user into groups according to the use conditions set by the user; a noise addition processing procedure for generating noise-added data for each group by adding noise according to the group to data resulting from an operation based on a user data group including a plurality of data classified into the same group by the classification procedure; a use target data generation procedure for generating use target data using the noise-added data for each group generated by the noise addition processing procedure; The use target data generation procedure includes: The noise-added data for each of the groups generated by the noise-adding processing procedure is weighted-averaged using a weight based on a value obtained by adding noise to the number of users included in the user data group for each of the groups, and data obtained is generated as the data to be used. An information processing program characterized by:

Citation Information

Patent Citations

  • Privacy protection type data provision system

    JP2014229039A

  • Server device and data transfer method

    JP2020112922A

  • Generation method, generation program, and generation device

    JP2021163014A

  • Privacy protection-type data providing system

    US20140351946A1

  • Training User-Level Differentially Private Machine-Learned Models

    US20190227980A1