Anonymization device, anonymization method, and anonymization program
The anonymization device addresses the challenge of anonymizing data with inconsistent formats by using a learned model to identify and process sensitive data, ensuring secure and compliant data provision across diverse data sources.
Patent Information
- Application Number
- JP2021102290
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-06-21
AI Technical Summary
Existing technologies struggle to anonymize specific parts of data with inconsistent formats from multiple data sources, failing to identify and protect sensitive information when providing data to users.
An anonymization device that extracts requested data, applies a learned model to determine anonymization targets, and performs anonymization processes on identified sensitive data using techniques like deletion, censoring, or encryption, ensuring consistent anonymization across diverse data formats.
Effectively anonymizes sensitive data within inconsistent data formats, ensuring secure and compliant data provision to users with varying access levels.
Smart Images

Figure 0007706271000001 
Figure 0007706271000002 
Figure 0007706271000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an anonymization device, an anonymization method, and an anonymization program for anonymizing a portion of data to be anonymized among the provided data.
Background Art
[0002] In recent years, information and communication technologies have attracted attention, and the opportunities to collect and utilize data from multiple data sources using information and communication technologies have been increasing. In the future, it is expected that the utilization of data collected from multiple data sources will further increase. On the other hand, some of the diverse data may contain parts that need to be anonymized. In this case, it is desirable to provide the data to the user who utilizes the data after anonymizing the parts to be anonymized. Patent Document 1 discloses a technology in a plant system including nodes such as sensors and a data collection device that collects data from each node, where each node encrypts the data and transmits it to the data collection device. In the technology described in Patent Document 1, each node determines whether to encrypt the data by comparing the value of the data with a threshold value, and the threshold value is determined according to the processing capacity of the data collection device. Thereby, the plant system described in Patent Document 1 can encrypt the data as much as possible even when the data collection device does not have high processing capacity.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In smart city development and the like, it is expected to utilize data collected from a variety of multiple data sources. When providing data collected from these data sources to users, for example, when there is a part indicating the content of items to be anonymized, such as names and addresses, that is, when the part to be anonymized is included in the data, the part to be anonymized is anonymized and the data is provided. However, in the data collected from a variety of multiple data sources, the item names, formats, etc. included in the data are also diverse, and it is difficult to identify where the part to be anonymized is. In the technology described in Patent Document 1, each node only compares the data with a known format and a threshold value that it has acquired, and then encrypts and transmits the entire data or transmits it as it is without encryption. When providing data with an inconsistent format to users, it is not possible to anonymize the part to be anonymized.
[0005] The present disclosure has been made in view of the above, and an object thereof is to obtain an anonymization device capable of anonymizing a part to be anonymized when providing data extracted from multiple types of data with inconsistent formats to users.
Means for Solving the Problems
[0006] In order to solve the above-described problems and achieve the object, the anonymization device according to the present disclosure includes a data extraction unit that extracts requested data, which is data requested by a user, from data collected from a plurality of data source devices, and for each piece of divided data obtained by dividing the requested data, by inputting the divided data into a learned model for inferring whether the input data is an object to be anonymized, an anonymization determination unit that determines whether the divided data is an object to be anonymized. The anonymization device further includes a data anonymization unit that performs an anonymization process on the divided data determined to be an object to be anonymized among the requested data using the determination result by the anonymization determination unit, and a request processing unit that provides the requested data after the anonymization process to the user. When the content of the first item corresponding to the corresponding content is a numerical value and is defined as an object to be anonymized, if the content corresponding to the second item associated with the first item is detected in the request data, and a numerical value is detected at a position satisfying the condition that the relative position with the content corresponding to the second item in the request data is determined, the detected numerical value is determined as an object to be anonymized. is provided.
Advantages of the Invention
[0007] According to the present disclosure, when providing data extracted from a plurality of types of data with inconsistent formats to a user, it is possible to anonymize the parts to be anonymized.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0009] Hereinafter, an anonymization device, an anonymization method, and an anonymization program according to an embodiment will be described in detail with reference to the drawings.
[0010] FIG. 1 is a diagram showing a configuration example of a data providing system including an anonymization device according to an embodiment. The data providing system of this embodiment includes data source devices 2-1 to 2-n, an anonymization device 1, and user devices 3-1 to 3-m. n is an integer of 2 or more, and m is an integer of 1 or more.
[0011] The data source devices 2-1 to 2-n transmit the data acquired in various systems to the anonymization device 1. The data transmitted by each of the data source devices 2-1 to 2-n may be of any type. For example, it includes at least one of various contracts, installation and update information of facilities, services, application forms for training, various measurement information, etc. The contract may be, for example, a contract related to electricity, a lease contract related to a parking lot, etc. The facilities may be, for example, street lights, traffic signal facilities, etc. The various measurement information may be, for example, measurement data of power consumption, measurement data of gas usage, etc. The data acquired by the data source devices 2-1 to 2-n is, for example, data acquired in a smart city, but is not limited thereto. The data transmitted by each of the data source devices 2-1 to 2-n does not have a fixed format. The data transmitted by each of the data source devices 2-1 to 2-n may include data with an unknown format to the anonymization device 1. Hereinafter, when showing the data source devices 2-1 to 2-n without distinguishing them individually, they are described as the data source device 2.
[0012] The user devices 3-1 to 3-m send a data provision request including information indicating the data to be provided, out of the data collected from the data source devices 2-1 to 2-n, to the anonymization device 1. The data provision request includes user identification information which is identification information indicating the requesting user. Note that the user identification information may not be included in the data provision request, and the user identification information used at the time of logging in to the anonymization device 1 etc. may be associated with the data provision request. Note that the user devices 3-1 to 3-m may correspond one-to-one with the users, or a plurality of user devices 3-1 to 3-m may correspond to one user, or one of the user devices 3-1 to 3-m may correspond to a plurality of users. When one of the user devices 3-1 to 3-m corresponds to a plurality of users, it is assumed that the viewing and use of the data within the device are managed for each user within the user device 3-1 to 3-m. When the user devices 3-1 to 3-m correspond one-to-one with the users, the user identification information may be the identification information of the user devices 3-1 to 3-m. Hereinafter, when not individually distinguishing the user devices 3-1 to 3-m, they are described as user device 3.
[0013] The anonymization device 1 accumulates the data received from the data source devices 2-1 to 2-n, extracts the data corresponding to the data provision request from the accumulated data, anonymizes the anonymization target part which is the part to be anonymized among the extracted data, and provides the data after anonymizing the anonymization target part to the user devices 3-1 to 3-m. Also, the anonymization target part is determined according to the authority of the user corresponding to the user devices 3-1 to 3-m, and the anonymization device 1 anonymizes the anonymization target part according to the authority of the user who is the request source of the data provision request.
[0014] As shown in FIG. 1, the anonymization device 1 includes a data collection unit 11, a data lake unit 12, a model generation unit 13, a model storage unit 14, a data extraction unit 15, an anonymization determination unit 16, a data anonymization unit 17, a request processing unit 18, and an authority information storage unit 19.
[0015] The data collection unit 11 receives data from the data source devices 2-1 to 2-n and stores the received data in the data lake unit 12. The data collection unit 11 may request the data source devices 2-1 to 2-n to transmit data, and each of the data source devices 2-1 to 2-n may transmit data, or each of the data source devices 2-1 to 2-n may spontaneously transmit data to the anonymization device 1 periodically or when there is data update, etc.
[0016] The data lake unit 12 is a data storage unit that accumulates the data collected from the data source devices 2-1 to 2-n. Here, it is assumed that the collected data is stored in the data lake unit 12 in its raw or minimally processed form.
[0017] The model generation unit 13 generates a trained model for determining whether each part in the data is an anonymization target by machine learning using the data stored in the data lake unit 12, and stores the generated trained model in the model storage unit 14.
[0018] The request processing unit 18 receives a data provision request from the user devices 3-1 to 3-m and passes the received data provision request to the data extraction unit 15. Also, the data provision request is provided with request identification information for identifying the data provision request. The request identification information may be added to the data provision request in the user devices 3-1 to 3-m, or may be provided by the request processing unit 18. The request processing unit 18 maintains the correspondence between the request identification information and the source user devices 3-1 to 3-m. Further, the request processing unit 18 receives the provided data generated in response to the data provision request from the data anonymization unit 17, and transmits the provided data to the source user devices 3-1 to 3-m of the data provision request corresponding to the provided data.
[0019] The data extraction unit 15 extracts requested data, which is the data requested by the user, from the data collected from a plurality of data source devices 2-1 to 2-n. Specifically, based on the data provision request received from the request processing unit 18, the data extraction unit 15 extracts data from the data lake unit 12 and outputs the extracted data to the anonymization determination unit 16 together with the user identification information and the request identification information corresponding to the data provision request. The data extraction unit 15 can extract by combining the data collected from different data source devices 2-1 to 2-n stored in the data lake unit 12 according to the data provision request.
[0020] The authority information storage unit 19 stores authority information indicating the correspondence with the level of the browsing authority for data corresponding to each of a plurality of users. The authority information is, for example, information indicating the correspondence between user identification information and the browsing authority. The browsing authority is set at two levels, for example, an authority without browsing restrictions and an authority without the right to view personal information. FIG. 2 is a diagram showing an example of the browsing information of the present embodiment. In FIG. 2, the authority is defined at two levels, but it is not limited to this, and it may be set at three or more levels. Also, in the example shown in FIG. 2, there are no browsing restrictions at level 1, but there are browsing restrictions at all levels, and the items that cannot be browsed may be set to differ depending on the level. Further, the authority information storage unit 19 also stores non-browsable information indicating non-browsable items, which are items determined not to be browsable for each level of authority.
[0021] FIG. 3 is a diagram showing an example of information that cannot be viewed in the present embodiment. In the example shown in FIG. 3, there is no viewing restriction for the authority level 1, and for the authority level 2, viewing of items such as names, addresses, company names, phone numbers, and people's faces is prohibited. FIG. 3 is an example, and the specific items that cannot be viewed are not limited to the example shown in FIG. 3. For example, the items that cannot be viewed may include an image indicating the position of a surveillance camera, a serial number of a product, an image related to a specific facility within a company, etc. Similarly, when there are three or more levels of viewing authority, the items that cannot be viewed are determined by the information that cannot be viewed for each level. Hereinafter, an example in which the authority information and the information that cannot be viewed are determined as shown in FIGS. 2 and 3 will be described. However, as described above, these information are not limited to this example.
[0022] Returning to the description of FIG. 1. For each piece of divided data obtained by dividing the request data, the anonymization determination unit 16 determines whether the divided data is an anonymization target by inputting the divided data into a learned model for inferring whether the input data is an anonymization target. Specifically, the anonymization determination unit 16 uses the user identification information received from the data extraction unit 15 and the authority information stored in the authority information storage unit 19 to determine the level of the viewing authority of the user corresponding to the user identification information, and uses the information that cannot be viewed stored in the authority information storage unit 19 to determine the items to be anonymized. The anonymization determination unit 16 further reads out the learned model stored in the model storage unit 14, and by inputting the data received from the data extraction unit 15 into the read learned model, determines whether each item in the data received from the data extraction unit 15 is an anonymization target. The anonymization determination unit 16 outputs the determination result for each item to the data anonymization unit 17 together with the data received from the data extraction unit 15 and the request identification information. When there are no items to be anonymized, that is, when the data corresponds to a user with no viewing restriction, the anonymization determination unit 16 outputs a determination result indicating that there are no items to be anonymized to the data anonymization unit 17 together with the data received from the data extraction unit 15 and the request identification information.
[0023] The data anonymization unit 17 performs anonymization processing on the divided data determined to be the anonymization target among the request data using the determination result by the anonymization determination unit 16. That is, the data anonymization unit 17 anonymizes the data received from the anonymization determination unit 16 based on the determination result received from the anonymization determination unit 16. That is, the data anonymization unit 17 performs anonymization processing on the portions determined to be the anonymization target by the anonymization determination unit 16 among the data received from the anonymization determination unit 16. The anonymization processing may be any processing as long as it can be anonymized. For example, when the corresponding portion is text, examples of the processing include deleting the corresponding portion, censoring the corresponding portion, encrypting the corresponding portion, etc. Also, when the corresponding portion is an image, examples of the processing include mosaic processing, encryption processing, etc. The data anonymization unit 17 outputs the data after anonymizing the anonymization target portion to the request processing unit 18 together with the request identification information. Note that, except for the anonymization target portion, the data received from the anonymization determination unit 16 is output as it is to the request processing unit 18. Also, when the determination result received from the anonymization determination unit 16 indicates that there are no items to be anonymized, the data anonymization unit 17 outputs the data received from the anonymization determination unit 16 to the request processing unit 18 together with the request identification information. The request data after anonymization processing is provided to the user by the request processing unit 18.
[0024] Note that, in the example shown in FIG. 1, the anonymization device 1 includes the data lake unit 12, but it is not limited thereto, and the data lake device having the data lake and the anonymization device performing anonymization processing may be separated. When separating into a data lake device and an anonymization device, the data lake device includes the data collection unit 11 and the data lake unit 12, and the anonymization device includes the model generation unit 13, the model storage unit 14, the data extraction unit 15, the anonymization determination unit 16, the data anonymization unit 17, the request processing unit 18, and the authority information storage unit 19. Then, the data extraction unit 15 and the model generation unit 13 acquire data from the data lake device.
[0025] Further, the model generation unit 13 may be provided in a learning device different from the anonymization device 1. In this case, the learning device including the model generation unit 13 generates a learned model, and the learned model generated by the learning device is stored in the model storage unit 14 of the anonymization device 1.
[0026] Next, the operation of this embodiment will be described. The data source devices 2-1 to 2-n are data obtained in various systems, and may include data with an unknown format, and even for the same item, the item names may be different. For example, regarding the item of person's name, multiple names such as "Full Name", "Name", "Name" etc. may be used as the item name. Also, for example, it is conceivable that a person's name is included in a place where the item name is not "Full Name", such as in the "Remarks" item.
[0027] Generally, when all the formats are known or all the item names are unified, it is easy to extract the items to be anonymized. However, as described above, there may be data with an unknown format, and when the item names may also be diverse, it is difficult to identify where in the data the content corresponding to the item to be anonymized is located.
[0028] Figs. 4 to 6 are diagrams showing an example of the data obtained from the data source device 2 of this embodiment. In the example shown in Fig. 4, the data obtained from the data source device 2 is data indicating a power contract. The contract data shown in Fig. 4 includes person name information 201 representing a person's name, address information 202 representing an address, and telephone number information 203 representing a telephone number. Therefore, as illustrated in Figs. 2 and 3, if the authority information and the non-viewable information are defined, the anonymization device 1 anonymizes the person name information 201, the address information 202, and the telephone number information 203 and provides the data shown in Fig. 4 to a user with level 2 authority.
[0029] In the example shown in FIG. 5, the data acquired from the data source device 2 is data indicating a contract regarding a lease contract of a parking lot. The data of the contract shown in FIG. 5 includes personal name information 201 and address information 202, and further includes company name information 204 representing a company name. Therefore, as illustrated in FIGS. 2 and 3, if authority information and non-viewable information are defined, the anonymization device 1 will anonymize the personal name information 201, address information 202, and company name information 204 and provide the data shown in FIG. 5 to a user with level 2 authority.
[0030] In the example shown in FIG. 6, the data acquired from the data source device 2 is data indicating an application form for the use of a certain service. The data of the contract shown in FIG. 6 includes personal name information 201, address information 202, and telephone number information 203, and further includes face image information 205 representing a person's face. Therefore, as illustrated in FIGS. 2 and 3, if authority information and non-viewable information are defined, the anonymization device 1 will anonymize the personal name information 201, address information 202, telephone number information 203, and face image information 205 and provide the data shown in FIG. 6 to a user with level 2 authority.
[0031] As can be seen by referring to FIGS. 4 to 6, the positions of the personal name information 201 representing a personal name are different in each piece of data. Also, in the example shown in FIG. 4, the item name corresponding to the personal name information 201 is "Name", while in the example shown in FIG. 5, there is no item name corresponding to the personal name information 201, and in the example shown in FIG. 6, the item name corresponding to the personal name information 201 is "First Name". Furthermore, in the example shown in FIG. 6, the personal name information 201 and the telephone number information 203 are included in the column with the item name "Remarks". Therefore, it is difficult to simply identify the personal name information 201, which is the content corresponding to the personal name item, from the item name, arrangement position, etc.
[0032] In this embodiment, the model generation unit 13 learns the content of each item to be anonymized by machine learning for each item to be anonymized. Then, the anonymization device 1 uses the learned model generated by learning to identify the part corresponding to the content of the item to be anonymized among the data provided to the user, anonymizes the identified part, and provides the data to the user. As a result, when providing data extracted from a plurality of types of data with inconsistent formats to the user, the part to be anonymized can be anonymized. Here, an example of providing data with an inconsistent format to the user is described. However, the anonymization device 1 of this embodiment may be applied when the formats of the data source devices 2-1 to 2-n are consistent. Thereby, while having the effect of being able to anonymize the part to be anonymized of data with an inconsistent format, when the formats of the data source devices 2-1 to 2-n are consistent, even when the content to be anonymized is described in a place where it cannot be determined from the item name, such as when a person's name is described in the remarks, the corresponding part can be anonymized.
[0033] For example, the model generation unit 13 uses, as a data set, the divided data obtained by dividing the data stored in the data lake unit 12 item by item and the correct answer data corresponding to the divided data, generates a learned model by machine learning using a plurality of data sets as learning data, and stores it in the model storage unit 14. The correct answer data is given, for example, by an operator of the anonymization device 1 or the like checking the data. The correct answer data may be input from an operator or the like using an input means not shown in FIG. 1, or may be transmitted from another device not shown in FIG. 1. The learned model may be a model for determining which item the input data corresponds to among the items to be anonymized, or may be a learned model for determining for each item to be anonymized whether the content corresponding to the item is included, or may be a learned model for determining for each level of authority whether the content of the item to be anonymized corresponding to the level is included. Further, the learned model may be created for each of image use and character string use. Further, learned models may be generated separately for image use and character string use for each level of authority.
[0034] When a trained model for determining which item of the input data is a target for anonymization is generated, for example, as learning data, data indicating a person's name and correct answer data of "person's name" are generated as a data set, and data indicating an address and correct answer data of "address" are generated as a data set. In this case, after learning is performed, when data is input to the trained model, which item of data among items such as "person's name" and "address" the input data is will be output as an inference result. By learning using data corresponding to the items set as anonymization targets, when data is input to the learned data, it is possible to determine which of the items set as anonymization targets the data is. Also, as described above, the trained model may be generated separately for images and character strings.
[0035] When a trained model is generated for each item, for example, in the generation of a trained model related to a person's name, data is divided for each item such as a person's name, address, and phone number, and a plurality of data sets are created that are composed of the content of the item in the divided data and correct answer data indicating whether it is a person's name or not. Whether it is a person's name may be given as binary data of, for example, 1 and 0, or may be given as other values. Note that, without using the data stored in the data lake unit 12 as learning data, for example, data representing a person's name and data representing other than a person's name may be generated for learning, and the data representing the person's name and the correct answer data may be used as learning data. For the data representing a person's name, a value indicating that it is a person's name is given as the correct answer data, and for the data representing other than a person's name, a value indicating that it is not a person's name is given as the correct answer data. The model generation unit 13 generates a trained model using a plurality of such data sets. Similarly, trained models are generated for other items such as addresses.
[0036] When a learned model is generated for each level of authority, a dataset is generated that consists of divided data divided for each item and correct answer data indicating whether or not it is a target for anonymization. Whether or not it is a target for anonymization may be given, for example, as binary data of 1 and 0, or may be given as other values. That is, in this case, for example, in the generation of the learned model corresponding to level 2 illustrated in FIG. 3, when the data illustrated in FIG. 4 is used as learning data, each of the person name information 201, address information 202, and telephone number information 203 is treated equally as a target for anonymization without distinction. Therefore, if the output value of the learned model in the case of being a target for anonymization is set to 1, the correct answer data corresponding to the divided data including the person name information 201, the correct answer data corresponding to the divided data including the address information 202, and the correct answer data corresponding to the divided data including the telephone number information 203 are all given as 1. Note that, similar to the example described above, data such as data representing a person name generated for learning may be used instead of using the data stored in the data lake unit 12 as learning data. Also, as described above, learned models may be generated separately for images and character strings for each level. When the data shown in FIG. 6 is used as learning data, for example, the face image information 205 portion shown in FIG. 6 is used for the generation of the learned model for images, and the character string portion is used for the generation of the learned model for character strings.
[0037] All of the above-mentioned learned models are learned models for inferring whether the input data is subject to anonymization. The specific generation method may be any of the above three methods, or other methods. When using a model for determining which item of the input data is subject to anonymization, if it is at the two-level of whether to restrict viewing as shown in FIGS. 2 and 3, learning may be performed using the learning data corresponding to each item with the non-viewable items corresponding to the restricted viewing level as the items to be anonymized. When the non-viewable items differ depending on the level, a learned model learned for all items included as non-viewable items at any level is generated. That is, a learned model learned for the items obtained by taking the logical OR of the non-viewable items at each level is generated. At the time of inference, the anonymization device 1 can determine whether it is subject to anonymization according to the level using the non-viewable information from the classification result obtained by inputting the data into the learned model.
[0038] In this way, machine learning in the model generation unit 13 can use, for example, supervised learning. As the algorithm of supervised learning, any algorithm may be used. For example, a neural network model can also be used. A neural network is composed of an input layer composed of a plurality of neurons, an intermediate layer (hidden layer) composed of a plurality of neurons, and an output layer composed of a plurality of neurons. The intermediate layer may be one layer or two or more layers.
[0039] FIG. 7 is a schematic diagram showing an example of a neural network. For example, in the case of a three-layer neural network as shown in FIG. 7, when a plurality of inputs are input to the input layer (X1-X3), the values are multiplied by the weights W1 (w11-w16) and input to the intermediate layer (Y1-Y2), and the result is further multiplied by the weights W2 (w21-w26) and output from the output layer (Z1-Z3). This output result varies depending on the values of the weights W1 and the weights W2.
[0040] In this embodiment, the weights W1 and W2 are adjusted so that the output from the output layer when the feature amount, which is the divided data described above, is input to the input layer approaches the correct answer data, thereby learning the relationship between the feature amount and the correct answer data. Note that the machine learning algorithm is not limited to neural networks. Also, after data is provided to the user, if the anonymization is incomplete, the result may be fed back and re-learning may be performed. The machine learning in the model generation unit 13 is not limited to this, and reinforcement learning may be used.
[0041] Next, the operation of data provision in response to the data provision request of this embodiment will be described. FIG. 8 is a flowchart showing an example of the data provision processing procedure in the anonymization device 1 of this embodiment. As shown in FIG. 8, the anonymization device 1 receives a data provision request from the user device 3 (step S1). Specifically, the request processing unit 18 receives a data provision request from the user device 3 and passes the received data provision request to the data extraction unit 15.
[0042] Next, the anonymization device 1 extracts the requested data (step S2). Specifically, the data extraction unit 15 extracts the data requested in the data provision request from the data lake unit 12 and outputs the extracted data to the anonymization determination unit 16 together with the user identification information and the request identification information corresponding to the data provision request. For example, when a user determines whether they are at home by combining delivery information and power consumption, the data provision request requests the provision of delivery information and power consumption. There is no particular restriction on the method of specifying the data requested in the data provision request. For example, the type of data such as a power contract may be specified, or a broad specification method such as all data related to power may be used.
[0043] Next, the anonymization device 1 determines whether or not the requester is a user who can view all items (step S3). Specifically, the anonymization determination unit 16 uses the user identification information received by the data extraction unit 15 and the authority information stored in the authority information storage unit 19 to determine the level of viewing authority, and refers to the non-viewable information stored in the authority information storage unit 19 to obtain non-viewable items corresponding to the determined level. If the anonymization determination unit 16 determines that there are no non-viewable items corresponding to the determined level, it determines that the requester is a user who can view all items. In the examples shown in FIGS. 2 and 3, the user with the user identification information corresponding to level 1 is a user who can view all items.
[0044] If the requester is not a user who can view all items (step S3 No), the anonymization device 1 determines whether or not each part of the extracted data is an anonymization target (step S4). Specifically, the anonymization determination unit 16 divides the data extracted by the data extraction unit 15 into each part, and for each divided data (divided data), uses the learned model stored in the model storage unit 14 to determine whether or not it is an anonymization target, and outputs the data received from the data extraction unit 15 together with the determination result and the request identification information to the data anonymization unit 17. As described above, although the learned model is divided in units of items, the items in the data to be provided are not always known. Therefore, it is divided by a delimiter unit or a defined data size, etc., and the divided parts are input to the learned model. As a result, if the corresponding content of the learned anonymization target item is included in the part, it is determined that it is an anonymization target. Details of step S4 will be described later.
[0045] In addition, when the data is a character string, for example, the units separated by spaces or the like are divided into parts. Also, names that are often used as item names to be anonymized, such as "name", "first name", "address", "phone number", "TEL", etc., are retained, these character strings are extracted, and the data may be divided using the item names as delimiters. Also, the learned model may be input with the item names removed from the divided data. One part may correspond to one item. Or, it may be divided by a predetermined data size. Also, for items with unclear item delimiters, they may be divided into words in combination with morphological analysis. Also, when an image is included, the entire image may be treated as one part, or the image may be divided into a defined size and the divided image may be treated as one part. Also, in the case of data in which a character string is captured as an image, the character string part is recognized as characters by OCR (Optical Character Recognition) processing.
[0046] Next, the anonymization device 1 anonymizes the anonymization target part using the determination result (step S5). Specifically, the data anonymization unit 17 anonymizes the anonymization target part using the determination result as to whether each piece of divided data is an anonymization target, and outputs the data after the anonymization process to the request processing unit 18 together with the request identification information. As described above, the anonymization process includes deletion, censoring, mosaic processing, encryption, and the like. Among the anonymization target items, for those for which leakage to a third party is not preferable separately from the user's authority, for example, even for data corresponding to a data providing request from a user at level 1 who has no browsing restriction, the anonymization target item may be encrypted with the public key corresponding to the user's secret key so that only the corresponding user can decrypt it. In this way, when it is determined by the anonymization determination unit 16 that anonymization is not to be performed, the data anonymization unit 17 may encrypt at least a part of the request data extracted by the data extraction unit 15 with the public key corresponding to the secret key possessed by the user corresponding to the request data, and output the encrypted data to the request processing unit 18. Alternatively, the data anonymization unit 17 may encrypt the data before outputting it to the request processing unit 18 regardless of whether the anonymization process is performed or not.
[0047] Next, the anonymization device 1 provides the data after the anonymization process to the user (step S6), and the anonymization device 1 ends the process. In step S6, specifically, the request processing unit 18 transmits the data received from the data anonymization unit 17 to the user device 3 that is the transmission source of the data providing request corresponding to the request identification information.
[0048] If the requesting element is a user who can view all items (Yes in step S3), the extracted data is provided to the user (step S7), and the process ends. In step S7, specifically, the anonymization determination unit 16 outputs the data received from the data extraction unit 15 and the request identification information to the data anonymization unit 17 together with a determination result indicating that anonymization is not necessary, and the data anonymization unit 17 outputs the data and the request identification information received from the anonymization determination unit 16 to the request processing unit 18. Then, the request processing unit 18 transmits the data received from the data anonymization unit 17 to the user device 3 that is the source of the data providing request corresponding to the request identification information.
[0049] In this way, the anonymization determination unit 16 uses the authority information stored in the authority information storage unit 19 to determine that anonymization is not performed when the user who is the source of the request data corresponds to a level without viewing restrictions, and when the user who is the source of the request data corresponds to a level with viewing restrictions, for each divided data, the divided data is input to a learned model for inferring whether the input data is an anonymization target, thereby determining whether the divided data is an anonymization target. The data anonymization unit 17, when the anonymization determination unit 16 determines whether each divided data is an anonymization target, uses the determination result by the anonymization determination unit 16 to perform anonymization processing on the divided data determined to be an anonymization target among the request data, and outputs the request data after the anonymization target has been anonymized to the request processing unit 18. Also, the data anonymization unit 17 outputs the request data extracted by the data extraction unit 15 to the request processing unit 18 when the anonymization determination unit 16 determines that anonymization is not performed. The request processing unit 18 provides the request data extracted by the data extraction unit 15 to the user when the anonymization determination unit 16 determines that anonymization is not performed.
[0050] In the above example, a case where users with different levels of viewing authority are mixed was described. However, when the same items are to be anonymized for all users, the anonymization device 1 of the present embodiment may be applied. In this case, the anonymization device 1 does not perform the determination in step S3 described above, but performs the processing from step S4 onwards after step S2. In this case, the anonymization device 1 does not need to hold the authority information.
[0051] Next, the details of the processing in step S4 described above will be explained. FIG. 9 is a flowchart showing an example of the anonymization determination processing procedure in step S4. Here, it is assumed that the learned models for outputting the classification results are generated separately for images and character strings. As shown in FIG. 9, the anonymization determination unit 16 divides the extracted data, that is, the data received from the data extraction unit 15 (step S11). As described above, when the data is a character string, the anonymization determination unit 16 divides it, for example, in units such as spaces, and when it is image data, it divides it, for example, into data sizes of a determined size. At this time, when images and character strings are mixed, the data is divided so that the images and character strings become separate divided data. The anonymization determination unit 16 determines whether the extracted data includes image data (step S12).
[0052] When the extracted data does not include image data (step S12 No), the anonymization determination unit 16 determines whether each divided data is an anonymization target using the learned model for inputting character strings (step S13). Specifically, the anonymization determination unit 16 reads out the learned model for character strings stored in the model storage unit 14, inputs each divided data into the read learned model, obtains which classification, that is, which item, the divided data corresponds to, and uses the obtained result and the non-viewable items stored in the authority information storage unit 19 to determine whether the divided data is an anonymization target. The anonymization determination unit 16 outputs the determination result to the data anonymization unit 17 together with the request identification information and data received from the data extraction unit 15, and ends the processing in step S4.
[0053] When the extracted data includes image data (step S12 Yes), the anonymization determination unit 16 determines whether the image contains only characters (step S14). Specifically, it determines whether the image indicated by the image data included in the extracted data is an image that contains only characters. For example, the anonymization determination unit 16 detects characters from the image data, and if no characters are detected, it determines that the image data is an image that contains only characters. The anonymization determination unit 16 may also determine whether the image contains characters by performing OCR processing on the image.
[0054] When the image contains only characters (step S14 Yes), the anonymization determination unit 16 performs OCR processing (step S15). Specifically, the anonymization determination unit 16 converts the image into a character string by performing OCR processing on the image, and proceeds to the processing of step S13.
[0055] When the image does not contain only characters (step S14 No), the anonymization determination unit 16 determines whether each piece of split data is an anonymization target using a learned model with the image as input (step S16). Specifically, the anonymization determination unit 16 reads out the learned model for images stored in the model storage unit 14, and inputs the partial data of the image, which is the area other than the characters in the image data, into the read learned model. As a result, it is determined which classification, that is, which item, the partial data of the image corresponds to, and using the obtained result and the non-viewable items stored in the authority information storage unit 19, it is determined whether the split data is an anonymization target, and the determination result is retained. Note that when an image data, which is one piece of split data, contains an anonymization target part and a non-anonymization target part, the split data may be determined as an anonymization target, or the determination results may be obtained separately for the anonymization target part and the non-anonymization target part.
[0056] Next, the anonymization determination unit 16 determines whether the image indicated by the image data is an image that does not contain characters (step S17). If it is an image that does not contain characters (step S17 Yes), the determination result of step S16 is output to the data anonymization unit 17 together with the request identification information and data received from the data extraction unit 15, and the process of step S4 ends.
[0057] In the case of an image containing characters (step S17 No), the process proceeds to step S13. When step S13 is performed via step S17, the anonymization determination unit 16 outputs both the determination result of the image and the determination result of the character string to the data anonymization unit 17 together with the request identification information and data received from the data extraction unit 15.
[0058] The learned models are generated separately for character strings and images. When the requested data does not include image data, the anonymization determination unit 16 inputs the divided data into the learned model for character strings. When the requested data includes image data and there is a region other than characters in the image indicated by the image data, the anonymization determination unit 16 inputs the data corresponding to the region into the learned model for images.
[0059] When learned models are generated for each item, the divided data is input into the corresponding learned model for each anonymization target item. For example, when the anonymization target items are five types, namely, name, address, company name, telephone number, and human face as shown in FIG. 3, four types of learned models for character strings and one type of learned model for images corresponding to the human face are generated. Therefore, in step S13, the anonymization determination unit 16 sequentially inputs the divided data into the four types of learned models, and determines the divided data as an anonymization target when a determination result that the divided data is the content of the corresponding item is obtained.
[0060] In the examples described above, items such as names, addresses, company names, phone numbers, and people's faces have been described, where it is possible to determine from the content itself that these are values indicating them. On the other hand, when only numerical values are described, such as some measurement data, it may be difficult to learn only the numerical values to determine whether they are items to be anonymized. For example, it is difficult to extract features only from the numerical values of the measurement results of a power meter or the measurement results of a gas meter. Also, even if an item name is described for these items, if it is a general name such as "measured value", it may not be possible to determine what kind of measured value it is even if it is learned in association with the item name. In such cases, by associating with the values of other items within the same data, it may be possible to determine whether the numerical value of the item to be anonymized is present. For example, if the number of the power meter is included near the measured value, it is presumed to be the measured value of the power meter. Therefore, when the measured value of the power consumption is to be anonymized, for example, a learned model is generated for the number of the power meter, and when the anonymization determination unit 16 finds a part corresponding to the number of the power meter using the learned model, if a numerical value is included in the part adjacent to that part, it is determined that it is the measured value of the power meter and is targeted for anonymization.
[0061] In this way, when the content of the first item whose corresponding content is a numerical value is determined to be an object to be anonymized, the anonymization determination unit 16 detects the content corresponding to the second item associated with the first item in the request data, and when a numerical value is detected at a position that satisfies the condition that the relative position with the content corresponding to the second item in the request data is determined, the detected numerical value is determined to be an object to be anonymized. The detection of the second item may use a learned model, or the content corresponding to the second item may be detected by searching using the item name of the second item as a search key. An example of the condition for determining the relative position is the condition of being adjacent, that is, adjacent with a delimiter such as a space as a boundary, but it is not limited to this, and the same row, the same column, etc. may be used as conditions.
[0062] Next, the hardware configuration of the anonymization device 1 of the present embodiment will be described. The anonymization device 1 of the present embodiment functions as an anonymization device 1 when a computer program, which is an anonymization program in which the processing in the anonymization device 1 is described, is executed on a computer system. FIG. 10 is a diagram showing a configuration example of a computer system that realizes the anonymization device 1 of the present embodiment. As shown in FIG. 10, this computer system includes a control unit 101, an input unit 102, a storage unit 103, a display unit 104, a communication unit 105, and an output unit 106, which are connected via a system bus 107.
[0063] In FIG. 10, the control unit 101 is a processor such as a CPU (Central Processing Unit), and executes an anonymization program in which the processing in the anonymization device 1 of the present embodiment is described. The input unit 102 is composed of, for example, a keyboard, a mouse, etc., and is used by the user of the computer system to input various information. The storage unit 103 includes various memories such as a RAM (Random Access Memory) and a ROM (Read Only Memory), and a storage device such as a hard disk, and stores programs that the control unit 101 should execute, necessary data obtained during the processing, and the like. The storage unit 103 is also used as a temporary storage area for programs. The display unit 104 is composed of a display, an LCD (Liquid Crystal Display Panel), etc., and displays various screens for the user of the computer system. The communication unit 105 is a receiver and a transmitter that perform communication processing. The output unit 106 is a printer, a speaker, etc. Note that FIG. 10 is an example, and the configuration of the computer system is not limited to the example of FIG. 10.
[0064] Here, an operation example of the computer system until the anonymization program of the present embodiment becomes executable will be described. In the computer system having the above-described configuration, for example, an anonymization program is installed in the storage unit 103 from a CD-ROM or a DVD-ROM set in a CD (Compact Disc)-ROM drive or a DVD (Digital Versatile Disc)-ROM drive (not shown). Then, when the anonymization program is executed, the anonymization program read from the storage unit 103 is stored in the main storage area of the storage unit 103. In this state, the control unit 101 executes the processing as the anonymization device 1 of the present embodiment according to the program stored in the storage unit 103.
[0065] In the above description, a program describing the processing in the anonymization device 1 is provided using a CD-ROM or a DVD-ROM as a recording medium. However, the present invention is not limited to this, and depending on the configuration of the computer system, the capacity of the provided program, etc., for example, a program provided by a transmission medium such as the Internet via the communication unit 105 may be used.
[0066] The model generation unit 13, the data extraction unit 15, the anonymization determination unit 16, and the data anonymization unit 17 shown in FIG. 1 are realized by the anonymization program stored in the storage unit 103 shown in FIG. 10 being executed by the control unit 101 shown in FIG. 10. The data lake unit 12, the model storage unit 14, and the authority information storage unit 19 shown in FIG. 1 are part of the storage unit 103 shown in FIG. 10. The data collection unit 11 and the request processing unit 18 shown in FIG. 1 are realized by the communication unit 105 shown in FIG. 10. The anonymization device 1 may be realized by a plurality of computer systems. For example, the anonymization device 1 may be realized by a cloud computer system.
[0067] For example, the anonymization program of the present embodiment causes a computer system to perform steps of extracting requested data, which is data requested by a user, from data collected from a plurality of data source devices; for each piece of split data obtained by splitting the requested data, inputting the split data into a learned model for inferring whether the input data is an anonymization target to determine whether the split data is an anonymization target; performing an anonymization process on the split data determined to be an anonymization target among the requested data using the determination result of whether the split data is an anonymization target; and providing the requested data after the anonymization process to the user.
[0068] As described above, in the present embodiment, data to be provided to a user in the data collected from the plurality of data source devices 2-1 to 2-n is extracted, the extracted data is split, and each piece of split data is input into a learned model that has learned the content of items to be anonymized, so as to determine whether each piece of split data is an anonymization target. Therefore, when providing data extracted from multiple types of data with inconsistent formats to a user, the parts to be anonymized can be anonymized.
[0069] The configuration shown in the above embodiment is an example, and it is possible to combine it with another known technology, combine the embodiments with each other, or omit or change a part of the configuration without departing from the gist.
Explanation of Reference Numerals
[0070] 1 Anonymization device, 2-1 to 2-n Data source devices, 3-1 to 3-m User devices, 11 Data collection unit, 12 Data lake unit, 13 Model generation unit, 14 Model storage unit, 15 Data extraction unit, 16 Anonymization determination unit, 17 Data anonymization unit, 18 Request processing unit, 19 Authority information storage unit.
Claims
1. A data extraction unit that extracts requested data, which is data requested by a user, from data collected from a plurality of data source devices; For each piece of divided data obtained by dividing the requested data, by inputting the divided data into a learned model for inferring whether the input data is a target for anonymization, an anonymization determination unit that determines whether the divided data is a target for anonymization; A data anonymization unit that performs anonymization processing on the divided data determined to be a target for anonymization among the requested data, using the determination result by the anonymization determination unit; A request processing unit that provides the requested data after the anonymization processing to the user; Comprising: When the content of the first item whose corresponding content is a numerical value is defined as a target for anonymization, the anonymization determination unit detects the content corresponding to the second item associated with the first item in the requested data, and When a numerical value is detected at a position that satisfies the condition that the relative position with the content corresponding to the second item in the requested data is determined, the detected numerical value is determined to be a target for anonymization. An anonymization device characterized by this.
2. A data extraction unit that extracts requested data, which is data requested by a user, from data collected from a plurality of data source devices; For each piece of divided data obtained by dividing the requested data, by inputting the divided data into a learned model for inferring whether the input data is a target for anonymization, an anonymization determination unit that determines whether the divided data is a target for anonymization; A data anonymization unit that performs anonymization processing on the divided data determined to be a target for anonymization among the requested data, using the determination result by the anonymization determination unit; A request processing unit that provides the requested data after the anonymization processing to the user; Comprising: The learned model is generated separately for character strings and images, When the requested data does not include image data, the anonymization determination unit inputs the divided data into the learned model for character strings, and when the data includes image data and there is a region in the image indicated by the image data that does not include characters as images, the anonymization determination unit inputs the data corresponding to the region into the learned model for images. An anonymization device characterized by this.
3. An authority information storage unit that stores authority information indicating the correspondence with the level of browsing authority regarding the data corresponding to each of the plurality of users; comprising The anonymization determination unit uses the authority information stored in the authority information storage unit to determine that anonymization is not performed when the user who is the requester of the request data corresponds to the level without viewing restrictions, and when the user who is the requester of the request data corresponds to the level with viewing restrictions, for each of the divided data, by inputting the divided data into a learned model for inferring whether the input data is an anonymization target, it determines whether the divided data is an anonymization target. When the data anonymization unit determines whether each of the divided data is an anonymization target by the anonymization determination unit, it uses the determination result by the anonymization determination unit to perform an anonymization process on the divided data determined to be an anonymization target among the request data, outputs the request data after the anonymization target is anonymized to the request processing unit, and when it is determined by the anonymization determination unit that anonymization is not performed, outputs the request data extracted by the data extraction unit to the request processing unit. The anonymization device according to claim 1 or 2, wherein when it is determined by the anonymization determination unit that anonymization is not performed, the request processing unit provides the request data extracted by the data extraction unit to the user.
4. The anonymization device according to claim 3, wherein the data anonymization unit encrypts at least a part of the request data extracted by the data extraction unit with a public key corresponding to a secret key possessed by the user corresponding to the request data, and outputs the encrypted data to the request processing unit.
5. a model generation unit that generates the learned model The anonymization device according to any one of claims 1 to 4, characterized by comprising
6. An anonymization method in an anonymization device, comprising: an extraction step of extracting request data, which is data requested to be provided by a user, from data collected from a plurality of data source devices; a determination step of determining whether each of the divided data obtained by dividing the request data is an anonymization target by inputting the divided data into a learned model for inferring whether the input data is an anonymization target; an anonymization step of performing an anonymization process on the divided data determined to be an anonymization target among the request data by using the determination result of whether the divided data is an anonymization target. An providing step of providing the requested data after the anonymization process to the user; including; In the determination step, when the content of the first item corresponding to a numerical value is defined as an anonymization target, the content corresponding to the second item associated with the first item is detected in the requested data, and a numerical value is detected at a position satisfying the condition that the relative position with the content corresponding to the second item in the requested data is defined, the detected numerical value is determined to be an anonymization target. An anonymization method characterized by this.
7. An anonymization method in an anonymization device, An extraction step of extracting requested data, which is data requested to be provided by a user, from data collected from a plurality of data source devices; For each piece of divided data obtained by dividing the requested data, by inputting the divided data into a learned model for inferring whether the input data is an anonymization target, a determination step of determining whether the divided data is an anonymization target; An anonymization step of performing an anonymization process on the divided data determined to be an anonymization target among the requested data using the determination result of whether the divided data is an anonymization target; An providing step of providing the requested data after the anonymization process to the user; including; The learned model is generated separately for character strings and images, In the determination step, when the requested data does not include image data, the divided data is input into the learned model for character strings, and when the requested data includes image data and there is an area in the image indicated by the image data that does not include characters as images, data corresponding to the area is input into the learned model for images. An anonymization method characterized by this.
8. In a computer system, An extraction step of extracting requested data, which is data requested to be provided by a user, from data collected from a plurality of data source devices; For each piece of divided data obtained by dividing the requested data, by inputting the divided data into a learned model for inferring whether the input data is an anonymization target, a determination step of determining whether the divided data is an anonymization target; An anonymization step of performing an anonymization process on the divided data determined to be an anonymization target among the requested data using the determination result of whether the divided data is an anonymization target; A providing step of providing the requested data after the anonymization process to the user; An anonymization program for causing execution, In the determination step, when the content of the first item whose corresponding content is a numerical value is determined to be an anonymization target, the content corresponding to the second item associated with the first item is detected in the requested data, and when a numerical value is detected at a position satisfying the condition that the relative position with the content corresponding to the second item in the requested data is determined, the detected numerical value is determined to be an anonymization target. An anonymization program characterized by this.
9. In a computer system, An extraction step of extracting requested data, which is data requested to be provided by a user, from data collected from a plurality of data source devices; For each piece of divided data obtained by dividing the requested data, by inputting the divided data into a learned model for inferring whether the input data is an anonymization target, a determination step of determining whether the divided data is an anonymization target; An anonymization step of performing an anonymization process on the divided data determined to be an anonymization target among the requested data, using the determination result of whether the divided data is an anonymization target; A providing step of providing the requested data after the anonymization process to the user; An anonymization program for causing execution, The learned model is generated separately for strings and images, In the determination step, when the requested data does not include image data, the divided data is input into the learned model for strings, and when the data includes image data and there is an area in the image indicated by the image data that does not include characters as images, data corresponding to the area is input into the learned model for images. An anonymization program characterized by this.
Citation Information
Patent Citations
Data management server, data management method and program
JP2007079984A
Data acquisition system, terminal equipment, data acquisition apparatus, and data acquisition method and program
JP2018073182A
Program, information processing unit, and system
JP2018186411A
Information processing apparatus, program and information processing method
JP2019200669A
Character recognition device, imaging device, character recognition method, and character recognition program
JP2021005164A