Anonymization system and anonymization method

JP7905276B2Active Publication Date: 2026-08-14HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

【0011】 本発明によれば、匿名化において様々な情報加工(仮名化加工、一般化加工、削除加工)を行う場合でも取得するデータの正当性を検証可能であり、より利便性のある匿名化システム及び匿名化方法が提供される。前述した以外の課題、構成および効果は、以下の実施形態により明らかにされる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007905276000001
    Figure 0007905276000001
  • Figure 0007905276000002
    Figure 0007905276000002
  • Figure 0007905276000003
    Figure 0007905276000003
Patent Text Reader

Abstract

To provide a more convenient anonymization system and anonymization method.SOLUTION: An anonymization system comprises an anonymization data providing device and an anonymization data user device. The anonymization data providing device performs, on each of elements in data, first processing of generating pseudonymization data and random numbers for pseudonymization data, second processing of generating generalization data and random numbers for generalization data, and third processing of generating deletion data and random numbers for deletion data. In pseudonymizing the data, the anonymization system performs, of the first processing, second processing, and third processing, processing necessary for a pseudonymization request on each of the elements in the data, and generates random numbers for pseudonymization data that is random numbers based on the processing. In verifying the data, the pseudonymization data user device performs processing that is not performed by the anonymization data providing device, generates random numbers for deletion data, and verifies the correctness of the pseudonymization data on the basis of the respective random numbers for deletion data.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , ,

[0006] , , , , ,

[0005] , , , ,

[0001] The present invention relates to an anonymization system and an anonymization method. ​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​ [Patent Document 2] Japanese Patent Publication No. 2020-77256 [Overview of the project] [Problems that the invention aims to solve]

[0007] However, the signature method described in Patent Document 1 focuses solely on the deletion of constituent elements and does not consider data replacement. Furthermore, while the signature method described in Patent Document 2 considers generalization processing in addition to deletion processing, it does not consider anonymization of text data. Therefore, even when various information processing steps are performed during anonymization, it is possible to verify the legitimacy of the acquired data, and there is a challenge in providing a more convenient anonymization system and method. [Means for solving the problem]

[0008] An anonymization system, which is one aspect of the invention disclosed in this application, is as follows. That is, this anonymization system comprises an anonymized data providing device and an anonymized data user device capable of communicating with the anonymized data providing device. When generating values ​​to be used for data verification, the anonymized data providing device performs the following processes for each element in the data: a first process which generates pseudonymized data obtained by replacing the original data in the element with an identifiable value, and a random number for pseudonymized data based on the original data and a random number for the original data; a second process which generates generalized data obtained by replacing the original data in the element with a generalized value, and a random number for generalized data based on the pseudonymized data and a random number for pseudonymized data; and a third process which generates deleted data obtained by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and a random number for generalized data. Furthermore, when anonymizing data in response to an anonymization request from an anonymized data user device, the anonymized data user device performs the necessary processing from the first, second, and third processing steps for each element of the data, generating anonymized data and random numbers for anonymized data based on the processing. The anonymized data user device obtains the anonymized data and random numbers for anonymized data via communication. Then, during data verification, it performs the processing from the first, second, and third processing steps that the anonymized data provider device has not performed on each element of the anonymized data, generating random numbers for deletion data. Based on the random numbers for deletion data generated by the anonymized data provider device and the random numbers for deletion data that it has generated, it verifies the validity of the obtained anonymized data.

[0009] An anonymization system, which is one aspect of the invention disclosed in this application, is as follows. That is, this anonymization system comprises an anonymized data providing device and an anonymized data user device capable of communicating with the anonymized data providing device. When generating values ​​to be used for data verification, the anonymized data providing device performs the following processes for each named entity in the text data: a first process which generates pseudonymized data obtained by replacing the original data with an identifiable value, and a random number for pseudonymized data based on the original data and a random number for the original data; a second process which generates generalized data obtained by replacing the original data with a generalized value, and a random number for generalized data based on the pseudonymized data and a random number for the pseudonymized data; a third process which generates deleted data obtained by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and a random number for generalized data; and a fourth process which generates a hash value of the original data for each non-named entity in the text data. Furthermore, when anonymizing data in response to an anonymization request from an anonymized data user device, the anonymized data user device performs the necessary processing from the first, second, and third processing steps for each named entity in the data, generating anonymized data and random numbers for anonymized data based on the processing. The anonymized data user device obtains the anonymized data and random numbers for anonymized data via communication. Then, during data verification, it performs the processing from the first, second, and third processing steps that the anonymized data provider device has not performed on each named entity in the anonymized data, generating random numbers for deletion data. It also generates a hash value of the original data for each non-named entity in the anonymized data. The authenticity of the obtained anonymized data is verified based on the random numbers and hash values ​​for deletion data generated by the anonymized data provider device and the random numbers and hash values ​​for deletion data that the user device has generated.

[0010] An anonymization method, which is one aspect of the invention disclosed in this application, is as follows. That is, this anonymization method is performed using an anonymized data providing device and an anonymized data user device that can communicate with the anonymized data providing device. When the anonymized data providing device generates values ​​to be used for data verification, it performs the following processes for each element in the data: a first process which generates pseudonymized data obtained by replacing the original data in the element with an identifiable value, and a random number for pseudonymized data based on the original data and a random number for the original data; a second process which generates generalized data obtained by replacing the original data in the element with a generalized value, and a random number for generalized data based on the pseudonymized data and a random number for pseudonymized data; and a third process which generates deleted data obtained by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and a random number for generalized data. Furthermore, when anonymizing data in response to an anonymization request from an anonymized data user device, the anonymized data user device performs the necessary processing from the first, second, and third processing steps for each element of the data, generating anonymized data and random numbers for anonymized data based on the processing. The anonymized data user device obtains the anonymized data and random numbers for anonymized data via communication. Then, during data verification, it performs the processing from the first, second, and third processing steps that the anonymized data provider device has not performed on each element of the anonymized data, generating random numbers for deletion data. Based on the random numbers for deletion data generated by the anonymized data provider device and the random numbers for deletion data that it has generated, it verifies the validity of the obtained anonymized data. [Effects of the Invention]

[0011] According to the present invention, even when various information processing steps (pseudonymization, generalization, deletion) are performed during anonymization, the validity of the acquired data can be verified, and a more convenient anonymization system and anonymization method are provided. Problems, configurations, and effects other than those mentioned above will be clarified by the following embodiments. [Brief explanation of the drawing]

[0012] [Figure 1] Figure 1 is an explanatory diagram showing an example of the system configuration of an anonymization system. [Figure 2] Figure 2 is a block diagram showing an example of the hardware configuration of a computer. [Figure 3] Figure 3 is a block diagram showing an example of the functional configuration of a signature generator terminal. [Figure 4] Figure 4 is an explanatory diagram showing an example of a confidential data table. [Figure 5] Figure 5 is an explanatory diagram showing an example of an index data table. [Figure 6] Figure 6 is an explanatory diagram showing an example of a random number data table. [Figure 7] Figure 7 is an explanatory diagram showing an example of a generalized rule table. [Figure 8] Figure 8 is a block diagram showing an example of the functional configuration of an anonymized data providing server. [Figure 9] Figure 9 is an explanatory diagram showing an example of an anonymized data table. [Figure 10] Figure 10 is a block diagram showing an example of the functional configuration of an anonymized data user terminal. [Figure 11] Figure 11 is a sequence diagram showing an example of the processing of an anonymization system. [Figure 12] Figure 12 is an explanatory diagram showing an example of an anonymization request setting screen. [Figure 13] Figure 13 is a flowchart showing a detailed processing procedure example of the signature generation process (step S1101) shown in FIG. 11. [Figure 14] Figure 14 is a flowchart showing a detailed processing procedure example of the generation of an element hash value within the element hash value list generation process (step S1301) shown in FIG. 13. [Figure 15] Figure 15 is a flowchart showing a detailed processing procedure example of the anonymization process (step S1108) shown in FIG. 11. [Figure 16] Figure 16 is a flowchart showing a detailed processing procedure example of the verification data generation process (step S1506) shown in FIG. 15. [Figure 17]FIG. 17 is a flowchart showing a detailed processing procedure example of the signature verification process (step S1110) shown in FIG. 11. [Figure 18] FIG. 18 is a flowchart showing a detailed processing procedure example of the element hash value generation process within the element hash value list generation process (step S1701) shown in FIG. 17. [Figure 19] FIG. 19 is an explanatory diagram showing a system configuration example of the anonymization system in Embodiment 2. [Figure 20] FIG. 20 is a block diagram showing a functional configuration example of the signature generator terminal in Embodiment 2. [Figure 21] FIG. 21 is a block diagram showing a functional configuration example of the anonymized data providing server in Embodiment 2. [Figure 22] FIG. 22 is a block diagram showing a functional configuration example of the anonymized data user terminal in Embodiment 2. [Figure 23] FIG. 23 is a sequence diagram showing an example of the processing of the anonymization system in Embodiment 2. [Figure 24] FIG. 24 is a flowchart showing a detailed processing procedure example of the signature generation process (step S2301) shown in FIG. 23. [Figure 25] FIG. 25 is a flowchart showing a detailed processing procedure example of the auxiliary data list generation process (step S2402) shown in FIG. 24. [Figure 26] FIG. 26 is a flowchart showing a detailed processing procedure example of the index data generation process (step S2501) shown in FIG. 25. [Figure 27] FIG. 27 is a flowchart showing a detailed processing procedure example of the element hash value generation process within the element hash value list generation process (step S2403) shown in FIG. 24. [Figure 28] FIG. 28 is a flowchart showing a detailed processing procedure example of the anonymization process (step S2307) shown in FIG. 23. [Figure 29] FIG. 29 is a flowchart showing a detailed processing procedure example of the verification data generation process (step S2309) shown in FIG. 23. [Modes for carrying out the invention]

[0013] Embodiments of the present invention will be described below with reference to the drawings. The embodiments are illustrative examples for explaining the present invention, and have been omitted and simplified as appropriate for clarity of explanation. The present invention can also be implemented in various other forms. Unless otherwise specified, each component may be singular or plural. The positions, sizes, shapes, and ranges of the components shown in the drawings may not represent their actual positions, sizes, shapes, and ranges in order to facilitate understanding of the invention. Therefore, the present invention is not necessarily limited to the positions, sizes, shapes, and ranges disclosed in the drawings. Examples of various types of information may be described using terms such as "table," "list," and "queue," but these types of information may also be represented by other data structures. For example, various types of information such as "XX table," "XX list," and "XX queue" may be referred to as "XX information." When describing identification information, terms such as "identification information," "identifier," "name," "ID," and "number" are used, and these terms are interchangeable. When there are multiple components with the same or similar function, they may be described using the same symbol but with different subscripts. Furthermore, when it is not necessary to distinguish between these multiple components, the subscripts may be omitted in the description. In embodiments, processing performed by executing a program may be described. Here, the computer executes the program using a processor (e.g., CPU, GPU) and performs processing defined by the program using memory resources (e.g., memory) and interface devices (e.g., communication ports). Therefore, the main entity performing the processing by executing the program may be the processor. Similarly, the main entity performing the processing by executing the program may be a controller, device, system, computer, or node having a processor. The main entity performing the processing by executing the program may be an arithmetic unit, and may include dedicated circuits that perform specific processing. Here, dedicated circuits include, for example, FPGAs (Field Programmable Gate Arrays), ASICs (Application Specific Integrated Circuits), CPLDs (Complex Programmable Logic Devices), etc. The program may be installed on the computer from the program source. The program source may be, for example, a program distribution server or a storage medium readable by the computer. If the program source is a program distribution server, the program distribution server includes a processor and storage resources for storing the program to be distributed, and the processor of the program distribution server may distribute the program to other computers. In addition, in the embodiment, two or more programs may be implemented as one program, or one program may be implemented as two or more programs.

[0014] This embodiment describes a more convenient anonymization system that allows for verification of data authenticity. By using this anonymization system, data authenticity is guaranteed. Based on this guarantee, the use of tampered data in various services is suppressed. Therefore, the anonymization system can contribute from a social perspective.

[0015] First, the first embodiment will be described with reference to Figures 1-18.

[0016] <Anonymization System> Figure 1 is an explanatory diagram showing an example of the system configuration of an anonymization system. The anonymization system 100 verifies that the provided information has not been improperly altered when providing information anonymized to prevent leakage of personal information and sensitive information. The anonymization system 100 is a system for data owners who possess tabular data containing sensitive information (hereinafter referred to as sensitive data) to provide the information to data users after anonymizing it. The anonymization system 100 includes a signature generator terminal 101, an anonymized data provision server 102, and an anonymized data user terminal 103.

[0017] The signature generator terminal 101, the anonymized data provider server 102, and the anonymized data user terminal 103 are connected to each other via a network 104 to send and receive information. The network 104 can be the Internet, a WAN (Wide Area Network), or a LAN (Local Area Network), among others. The network 104 can be connected via wired or wireless connection.

[0018] The signature generator terminal 101 is a terminal used by the signature generator (for example, a government agency), which is the data holder, and generates a signature on sensitive data so that the legitimacy of the anonymized data, which has been anonymized from the sensitive data, can be verified. The anonymized data provision server 102 provides the signature value of the sensitive data and generates and provides anonymized data from the sensitive data entrusted by the signature generator in accordance with the request of the anonymized data user. The anonymized data user terminal 103 is a terminal used by the anonymized data user and performs verification of the legitimacy of the anonymized data.

[0019] <Example of hardware configuration for a computer (signature generator terminal 101, anonymized data provision server 102, and anonymized data user terminal 103)> Figure 2 is a block diagram showing an example of a computer hardware configuration. Computer 200 includes a processor 201, a storage device 202, an input device 203, an output device 204, and a communication interface (communication IF) 205. The processor 201, storage device 202, input device 203, output device 204, and communication IF 205 are connected by a bus 206. The processor 201 controls computer 200. The storage device 202 serves as the work area for the processor 201. The storage device 202 is a non-temporary or temporary recording medium that stores various programs and data. Examples of storage devices 202 include ROM (Read Only Memory), RAM (Random Access Memory), HDD (Hard Disk Drive), and flash memory. The input device 203 takes data in. Examples of input devices 203 include a keyboard, mouse, touch panel, numeric keypad, scanner, microphone, and sensor. The output device 204 outputs data. Output devices 204 include, for example, displays, printers, and speakers. The communication IF 205 connects to the network 104 and sends and receives data.

[0020] <Example of functional configuration of signature generator terminal 101> Figure 3 is a block diagram showing an example of the functional configuration of the signature generator terminal 101. The signature generator terminal 101 includes a signature generation unit 301, a hash value generation unit 302, a sensitive data table 303, an index data table 304, a random number data table 305, a generalization rule table 306, a signature value table 307, and a signing key 308.

[0021] The signature generation unit 301 and the hash value generation unit 302 are specifically implemented, for example, by having the processor 201 execute a program stored in the storage device 202 shown in Figure 2. The sensitive data table 303, index data table 304, random number data table 305, generalization rule table 306, signature value table 307, and signature key 308 are stored in the storage device 202.

[0022] The signature generation unit 301 generates signature values ​​for sensitive data. The hash value generation unit 302 generates hash values ​​using a one-way function or the like. The sensitive data table 303 stores sensitive data such as personal information. The index data table 304 stores index numbers for identifying pseudonymized data. Details of the index data table 304 will be described later in Figure 5. The random number data table 305 stores random numbers assigned to enhance the security of hash values. Details of the random number data table 305 will be described later in Figure 6. The generalization rule table 306 stores generalization rules for generalizing sensitive data. Details of the generalization rule table 306 will be described later in Figure 7. The signature value table 307 stores pairs of identification information for the sensitive data table 303 and signature values ​​for all sensitive data within the sensitive data table 303. The signature key 308 is key information for encrypting the hash values ​​for the sensitive data table 303. The signature key 308 is, for example, a private key in a public-key cryptography scheme.

[0023] <Sensitive Data Table 303> Figure 4 is an explanatory diagram showing an example of a sensitive data table 303. The sensitive data table 303 has attribute fields: ID 401, name 402, address 403, age 404, and gender 405. The combination of attribute values ​​for each field in the same row constitutes an entry that defines an individual's sensitive data.

[0024] ID401 is an identifier that uniquely identifies an individual's sensitive data. Name402 is the individual's full name in the sensitive data identified by ID401. Address403 is the place of residence of the individual in the sensitive data identified by ID401. Age404 is the number of years since the individual's birth in the sensitive data identified by ID401. Age404 is updated over time. Gender405 is information that distinguishes the individual's gender in the sensitive data identified by ID401.

[0025] <Index Data Table 304> Figure 5 is an explanatory diagram showing an example of the index data table 304. The index data table 304 is a group of tables (541-545) that define index numbers 501-505 for each of the attributes defined in the sensitive data table 303: ID 401, name 402, address 403, age 404, and gender 405. Index numbers 501-505 are values ​​assigned to each type of attribute value.

[0026] <Random number data table 305> Figure 6 is an explanatory diagram showing an example of the random number data table 305. The random number data table 305 has the same attribute ID 601, name 602, address 603, age 604, and gender 605 as the sensitive data table 303. Also, the random number data table 305 has the same number of records as the sensitive data table 303. In other words, the sensitive data table 303 and the random number data table 305 are tables of the same size, and each cell in the random number data table 305 stores a random number.

[0027] <Generalized Rule Table 306> Figure 7 is an explanatory diagram showing an example of a generalization rule table 306. The generalization rule table 306 has two fields: pre-generalization 701 and post-generalization 702. Pre-generalization 701 is the place name before generalization, and post-generalization 702 is a place name that includes the place name defined in pre-generalization 701 and covers a broader area than that place name. For example, if the place name in pre-generalization 701 is a prefecture name, then post-generalization 702 will be a regional name. Also, if the place name in pre-generalization 701 is a city / ward / town / village name, then post-generalization 702 will be a prefecture name. For example, the conversion from Tokyo to the Kanto region is a one-stage generalization. Also, the conversion from Tokyo to Japan is a two-stage generalization because it involves a conversion from Tokyo to the Kanto region and a conversion from the Kanto region to Japan.

[0028] Furthermore, while Figure 7 uses place names as an example to define the generalization rules, other information can also be used. For example, in the case of age, if the age of pre-generalization 701 is 0-4 years old, then the post-generalization 702 will be "young," if the age of pre-generalization 701 is 5-14 years old, then the post-generalization 701 will be "boy," if the age of pre-generalization 701 is 15-24 years old, then the post-generalization 702 will be "young adult," if the age of pre-generalization 701 is 25-44 years old, then the post-generalization 702 will be "adult," if the age of pre-generalization 701 is 45-64 years old, then the post-generalization 702 will be "middle-aged," and if the age of pre-generalization 701 is 65 years or older, then the post-generalization 702 will be "elderly."

[0029] <Example of functional configuration of anonymized data provision server 102> Figure 8 is a block diagram showing an example of the functional configuration of the anonymized data provision server 102. The anonymized data provision server 802 includes a web server function 801, an anonymization processing unit 802, a hash value generation unit 302, a sensitive data table 303, an index data table 304, a random number data table 305, a generalization rule table 306, a signature value table 307, an anonymized data table 803, a verification index data table 804, a verification random number data table 805, and a verification generalization rule table 806.

[0030] The web server function 801 and the anonymization processing unit 802 are specifically implemented, for example, by having the processor 201 execute a program stored in the storage device 202 shown in Figure 2. In addition, the anonymized data table 803, the verification index data table 804, the verification random number data table 805, and the verification generalized rule table 806 are stored in the storage device 202.

[0031] The web server function 801 maintains a web page accessible from the anonymized data user terminal 103, which contains the identification information of the sensitive data table 303 and the signature values ​​for all the sensitive data within the sensitive data table 303. The anonymization processing unit 802 performs verifiable anonymization processing on the sensitive data. The anonymized data table 803 stores the anonymized data, which is the result of the anonymization processing of the sensitive data by the anonymization processing unit 802. The verification index data table 804 stores index data for attributes that were not anonymized from the index data. The verification random number data table 805 is a table of the same size as the random number data table 305 and stores random number data for signature verification. The verification generalization rule table 806 stores the generalization rules necessary for signature verification from the generalization rule table 306.

[0032] <Anonymized Data Table 803> Figure 9 is an explanatory diagram showing an example of an anonymized data table 803. The anonymized data table 803 has the same fields as the sensitive data table 303: ID 901, name 902, address 903, age 904, and gender 905. The combination of values ​​for each field in the same row constitutes an entry that defines the anonymized data of a single individual. ID 901 is identification information that uniquely identifies the anonymized data of an individual. Name 902 is the name 402 of the individual in the anonymized data identified by ID 901, and is blank due to deletion processing. Address 903 is the place name 602 after generalization of address 403. Age 904 is the age of the individual in the anonymized data identified by ID 901. Gender 905 is information that distinguishes the gender of the individual in the anonymized data identified by ID 901.

[0033] <Example of functional configuration of anonymized data user terminal 103> Figure 10 is a block diagram showing an example of the functional configuration of an anonymized data user terminal 103. The anonymized data user terminal 103 includes a web browser function 1001, a signature verification unit 1002, a hash value generation unit 302, an anonymized data table 803, a verification index data table 804, a verification random number data table 805, a verification generalization rule table 806, a signature value 1003, and a verification key 1004.

[0034] The web browser function 1001 and the signature verification unit 1002 are specifically implemented, for example, by having the processor 201 execute a program stored in the storage device 202 shown in Figure 2. The signature value 1003 and the verification key 1004 are also stored in the storage device 202.

[0035] The web browser function 1001 receives the web page published by the anonymized data provision server 102 and displays it on the anonymized data user terminal 103. The signature verification unit 1002 verifies the legitimacy of the anonymization using the anonymized data table 803, the verification index data table 804, the verification random number data table 805, the verification generalization rule table 806, the signature values ​​1003 for all sensitive data in the sensitive data table 303, and the verification key 1004 as input.

[0036] The anonymized data table 803, the validation index data table 804, the validation random number data table 805, and the validation generalization rule table 806 store the anonymized data, validation index data, validation random number data, and validation generalization rules obtained from the anonymized data provision server 102.

[0037] The signature value 1003 is the signature value for all sensitive data in the sensitive data table 303 obtained from the anonymized data provision server 102. The verification key 1004 is key information used to decrypt the signature value 1003, and is, for example, the public key of the public key scheme for the signing key 308.

[0038] Some or all of the programs and data stored in the signature generator terminal 101, the anonymized data providing server 102, and the anonymized data user terminal 103 may be stored in advance in the non-temporary area of ​​the storage device 202 provided by the computer 200, or, if necessary, may be stored in the non-temporary area of ​​the storage device 202 from the non-temporary storage device of other devices connected to the network 104, or from a non-temporary storage medium connected to an interface not shown.

[0039] <Sequence of anonymization system 100> Next, an example of the anonymization system's processing will be described with reference to Figure 11. Figure 11 is a sequence diagram of the anonymization system 100.

[0040] The signature generator terminal 101, using its signature generation unit 302, takes the sensitive data table 303, index data table 304, random number data table 305, generalized rule table 306, and signature key 308 as input and performs a signature generation process to generate a signature value 1003 for all the data in the sensitive data table 303 (step S1101). Details of the signature generation process (step S1101) will be described later with reference to Figures 13 and 14.

[0041] Next, the signature generator terminal 101 transmits the sensitive data, index data, random number data, generalization rules, and the signature value generated by the signature generation process (step S1101) to the anonymized data providing server 102 (step S1102). The signature generator terminal 101 also stores the identification information of the sensitive data table 303 and the signature values ​​for all the sensitive data in the sensitive data table 303 in the signature value table 307.

[0042] The anonymized data provision server 102 obtains the sensitive data, index data, random number data, generalization rule, and signature value transmitted from the signature generator terminal 101 and stores them in the storage device 202 (step S1103).

[0043] Next, the anonymized data provision server 102 generates a web page accessible via the network 104 using its web server function 801, and notifies the web browser function 1001 of the anonymized data user terminal 103 of the web page's URL (step S1104).

[0044] Next, the anonymized data user terminal 103 uses the web browser function 1001 to obtain a web page containing the identification information of the sensitive data table 303, the attributes of the sensitive data table 303, and the signature values ​​for all sensitive data within the sensitive data table 303 (step S1105).

[0045] Next, the anonymized data user terminal 103 downloads the signature values ​​1003 for all sensitive data in the sensitive data table 303 from a web page by operating the data user's input device 203, and stores them in the anonymized data user terminal 103's storage device 202 (step S1106).

[0046] Next, the anonymized data user terminal 103 sets an anonymization request by operating the data user's input device 203 and sends a request to acquire anonymized data, including the anonymization request, to the Web server function 801 of the anonymized data provision server 102 (step S1107). In this embodiment, the request to acquire anonymized data is assumed to include the following three anonymization requests. Anonymization request 1: Rename ID401 to a pseudonym Anonymization request 2: Delete name 402 Anonymization request 3: Generalize address 403 by one level.

[0047] Here, an example of an anonymization request setting screen will be described with reference to Figure 12. Figure 12 is an explanatory diagram showing an example of an anonymization request setting screen. The anonymization request setting screen 1200 is displayed on the output device 204 of the anonymization data user terminal 103 and has a processing rule specification area 1201, a generalization method specification area 1202, and an execution button 1203.

[0048] The processing rule specification area 1201 has radio buttons for each attribute with the options "No processing," "Pseudonymization," "Deletion," or "Generalization," and only one processing rule can be selected for each attribute. For example, suppose the processing rule for ID 401 is set to "Pseudonymization," the processing rule for name 402 is set to "Deletion," the processing rule for address 403 is set to "Generalization," and the processing rules for age 404 and gender 405 are set to "No processing."

[0049] The generalization method specification area 1202 is an area where it is possible to specify how many levels of generalization to apply to each attribute. For example, if "1 level" is specified for the generalization of address 403, it is required to generalize address 403 by one level.

[0050] When the execute button 1203 is pressed, a request to obtain anonymized data, including anonymization requests for each entered attribute, is sent to the anonymized data providing server 102 (step S1107).

[0051] Returning to Figure 11, the explanation continues. The anonymized data provision server 102, via the Web server function 801, obtains the identification information of the sensitive data table 303 and the anonymization request notified from the anonymized data user terminal 103, and passes them to the anonymization processing unit 802. The anonymized data provision server 102 then uses the anonymization processing unit 802 as input to perform anonymization processing, generating anonymized data, verification index data, verification random number data (random numbers for anonymized data), and verification generalization rules, and passes these to the Web server function 801 (step S1108). Details of the anonymization processing (step S1108) will be described later in Figures 15 and 16.

[0052] The web server function 801 registers the anonymized data, verification index data, verification random number data, and verification generalization rules on a downloadable web page for data users, and notifies the web browser function 1001 of the anonymized data user terminal 103 of that URL.

[0053] Next, the anonymized data user terminal 103 uses its web browser function 1001 to access a download web page for data users by inputting the notified URL, downloads and obtains the anonymized data, verification index data, verification random number data, and verification generalization rules, and stores them in the storage device 202 (step S1109). As a result, the anonymized data table 803, the verification index data table 804, the verification random number data table 805, and the verification generalization rule table 806 are created in the storage device 202 of the anonymized data user terminal 103.

[0054] Finally, the anonymized data user terminal 103, using the signature verification processing unit 1102, performs a signature verification process to verify the legitimacy of the anonymization process (step S1108) performed by the anonymized data provision server 102, taking the anonymized data, verification index data, verification random number data, verification generalization rule, signature value, and verification key as input (step S1110). The anonymized data user terminal 103 displays "Verification successful" on the output device 204 if the legitimacy of the anonymization process (step S1108) is verified, or "Verification failed" if the legitimacy of the anonymization process (step S1008) is not verified (step S1110).

[0055] If the anonymized data user terminal 103 successfully verifies the data, it displays the anonymized data on the output device 204. If the verification fails, it does not display the anonymized data on the output device 204. Details of the signature verification process (step S1110) will be described later in Figures 17 and 18.

[0056] As shown in Figure 11, users of anonymized data can verify whether the anonymized data they obtained has been properly anonymized through legitimate anonymization processing (step S1008). This prevents the provision of services that utilize analysis results obtained from fraudulently anonymized data.

[0057] <Signature generation process (step S1101)> Figure 13 is a flowchart showing a detailed example of the signature generation process (step S1101) shown in Figure 11.

[0058] The signature generation unit 302 generates index data for each attribute and stores it in the index data table 304 (step S1301). Specifically, for example, the signature generation unit 302 performs the following processing for all attributes of the sensitive data. First, the signature generation unit 302 obtains attribute values ​​for each attribute from the sensitive data. For example, in the case of the sensitive data table 303 in Figure 4, the signature generation unit 302 obtains the attribute values ​​"247", "500", "294", and "80" for ID 401, which are the attribute values ​​of the ID 401 column. Next, the signature generation unit 302 generates index numbers "1", "2", "3", and "4" for each of the obtained attribute values ​​"247", "500", "294", and "80" of ID 401, and stores the pairs of attribute values ​​and index numbers in the index data table 306. The same applies to other attributes (name 402, address 403, age 404, gender 405).

[0059] The signature generation unit 302 generates random numbers equal to the number of cells in the sensitive data 303, i.e., the number of attributes M × the number of records N (in Figure 4, M=5 and N=4, so M × N = 20), and stores them in the random number data table 306 (S1302).

[0060] The signature generation unit 302 calculates a hash value (hereinafter referred to as element hash value) EH for each cell of the sensitive data 303 and stores it in the element hash value list (S1303). The element hash value list stores the number of element hash values ​​(EH_11, EH_12, ..., EH_1N, EH21, ..., EH_MN) for the number of cells in the sensitive data table 303, M × N (20 in Figure 4). The detailed processing of the element hash value generation process will be described later in Figure 14.

[0061] Next, the signature generation unit 302 generates the hash value TH of the entire sensitive data table 303 from the element hash value list (S1304). Specifically, the hash value generation unit 303 is input with a value obtained by concatenating the M × N element hash values ​​included in the element hash value list, and the overall hash value TH is generated. In other words, the overall hash value TH is calculated as follows. TH=Hash(EH_11+EH_12+···+EH_MN) Here, Hash() is a one-way function executed in the hash value generation section, and + represents string concatenation.

[0062] Finally, the signature generation unit 302 generates a signature value 1003 using the overall hash value TH and the signing key 308 (step S1305).

[0063] <Element hash value generation process> Figure 14 is a flowchart showing a detailed example of the processing procedure for generating element hash values ​​in the element hash value list generation process (S1303) shown in Figure 13.

[0064] The signature generation unit 302 reads the cell value Data_O from the sensitive data 303 to generate an element hash value, the index number paired with Data_O from the index data, and the random number R_O corresponding to Data_O from the random number data, and performs pseudonymization processing (S1401). Specifically, it generates pseudonymized data Data_P and the random number R_P corresponding to the pseudonymized data Data_P. Data_P is a value obtained by concatenating the name of the attribute to which Data_O belongs and the index number paired with Data_O. For example, the pseudonymized data for the attribute value "Tokyo" in address 403 in Figure 4 is "Address 1". R_P is the hash value of the concatenated value of Data_O and R_O. Data_P = attribute name + index number R_P = Hash(Data_O + R_O)

[0065] Next, the signature generation unit 302 determines whether a generalized rule table exists for the attributes of Data_O (S1402).

[0066] If a generalization rule table exists (S1402: Yes), the signature generation unit 302 performs a generalization process in which the generalized data Data_G is the generalized value of Data_O obtained by referring to the generalization rule table, and the random number R_G corresponding to the generalized data Data_G is the hash value of the concatenated value of Data_P and R_P (S1403). The generalized value of Data_G = Data_O R_G = Hash(Data_P + R_P)

[0067] On the other hand, if no generalization rule exists (S1402: No), the signature generation unit 302 performs a generalization process in which the generalized data Data_G is the name of the attribute to which Data_O belongs, and the random number R_G is the hash value of the concatenated value of Data_P and R_P, similar to S1403 (S1404). Data_G=attribute name R_G = Hash(Data_P + R_P)

[0068] The signature generation unit 302 determines whether Data_G can be further generalized (S1405). Specifically, it determines whether the value of Data_G is included in the pre-generalization 701 of the generalization rule table 306.

[0069] If generalization is not possible (S1405: No), the process proceeds to S1408. On the other hand, if generalization is possible (S1405: Yes), the signature generation unit 302 performs generalization processing such that the generalized data Data_G' is the generalized value of Data_G obtained by referring to the generalization rule table, and the random number R_G' is the hash value of the concatenated value of Data_G and R_G (S1406). Data_G' = Generalized value of Data_G R_G'=Hash(Data_G+R_G)

[0070] The signature generation unit 302 assigns Data_G' to Data_G and R_G' to R_G (S1407), and then the process proceeds to S1405.

[0071] The signature generation unit 302 performs a deletion process of blanking out the post-deletion data Data_R and setting the random number R_R corresponding to the post-deletion data Data_R as the hash value of the concatenated value of the generalized data Data_G and the random number R_G (S1408). The random number R_R is set as the element hash value EH, and the element hash value generation process ends. Data_R = blank R_R = Hash(Data_G + R_G) = EH

[0072] <Anonymization process (step S1108)> FIG. 15 is a flowchart showing a detailed processing procedure example of the anonymization process (step S1108) shown in FIG. 11. The anonymization process (step S1108) is a process in which the anonymization data providing server 102 anonymizes sensitive data that can be verified using sensitive data, index data, random number data, generalization rule data, and anonymization processing requests by the anonymization processing unit 802.

[0073] The anonymization processing unit 802 initializes a variable i representing which attribute the processing is for to 0 (step S1501).

[0074] Next, the anonymization processing unit 802 determines whether i < M (step S1502). Here, M is the number of all attributes in the sensitive data. If i < M is not satisfied (step S1502: No), the anonymization process S1108 ends. On the other hand, if i < M (step S1502: Yes), the anonymization processing unit 802 determines whether processing of attribute i is requested by step S1107 (step S1503).

[0075] If processing of attribute i is not requested (step S1503: No), the anonymization processing unit 802 copies the N cells of attribute i in the sensitive data table 303 to attribute i in the anonymization data table 803, and copies the N cells of attribute i in the random number data table to the verification random number data table 805 (step S1505), and the process proceeds to step S1506.

[0076] On the other hand, if processing of attribute i is requested (step S1503: Yes), anonymization processing is performed on each cell value of attribute i (step S1504). Specifically, the following processing is performed on each of the N cells of attribute i, and the processed data is stored in the anonymized data table 803, and the random numbers corresponding to the processed data are stored in the verification random number data table 805. The processed data and random numbers for the cell value in the jth row of attribute i are stored in the jth row cell of attribute i in the anonymized data table 803 and the verification random number data table 805, respectively.

[0077] First, the anonymization processing unit 802 determines which of the following processing methods is requested for attribute i: "pseudonymization," "generalization," or "deletion."

[0078] If "pseudonymization" is requested, the anonymization processing unit 802 performs the pseudonymization process S1401 of the element hash value generation process shown in Figure 14, and stores the pseudonymized data Data_P in the anonymization data table 803 and the random number R_P corresponding to the pseudonymized data in the verification random number data table 805.

[0079] If "generalization" is requested, the anonymization processing unit 802 performs the S1401-S1407 processes in which the determination in S1405 of the element hash value generation process shown in Figure 14 is changed from "whether further generalization is possible" to "whether the generalization level of the anonymization request has been reached". The generalized data Data_G obtained as a result of the processing, and the random number R_G corresponding to the generalized data Data_G are stored in the anonymization data table 803 and the verification random number data table 805, respectively.

[0080] If "deletion" is requested, the anonymization processing unit 802 performs all of S1401 to S1408 of the element hash value generation process shown in Figure 14. The deleted data Data_R obtained as a result of S1408, and the random number R_R corresponding to the deleted data Data_R, are stored in the anonymization data table 803 and the verification random number data table 805, respectively.

[0081] The anonymization processing unit 802 generates verification index data and verification generalization rules for attribute i using the index data, generalization rules, and anonymization processing request (step S1506). Details of the verification data generation process (step S1506) will be described later in Figure 16.

[0082] The anonymization processing unit 802 increments the variable i (step S1507), and the process returns to step S1502.

[0083] <Verification data generation process S1506> Figure 16 is a flowchart showing a detailed example of the processing procedure for the verification data generation process (step S1506) shown in Figure 15.

[0084] The anonymization processing unit 802 refers to the anonymization request for attribute i and determines whether the anonymization request for attribute i is "deletion" or not (S1601). If "deletion" is requested (S1601: Yes), the verification data generation process (S1506) is terminated.

[0085] On the other hand, if the request is for anything other than "deletion," that is, a processing request for "pseudonymization" or "generalization," or if there is no anonymization processing request (S1601: No), the anonymization processing unit 802 determines whether or not a generalization rule exists for attribute i (S1602). If no generalization rule exists (S1602: No), the process proceeds to S1604. On the other hand, if a generalization rule exists (S1602: Yes), the generalization rule for attribute i is copied to the verification generalization rule table 806 (S1603).

[0086] The anonymization processing unit 802 refers to the anonymization request for attribute i and determines whether the anonymization request for attribute i is "generalization" or not (S1604). If "generalization" is requested (S1604: Yes), the verification data generation process (S1506) is terminated.

[0087] On the other hand, if the request is for something other than "generalization," that is, a request for "pseudonymization," or if there is no request for anonymization (S1604: No), the anonymization processing unit 802 sequentially refers to the attribute value and index number pairs stored in the index data and adds the generalization rule for generalizing the pseudonymized data to the verification generalization rule table 806 (S1605).

[0088] Specifically, it is determined whether the attribute value in question exists in the pre-generalization attribute of the generalization rule for attribute i. If it does not exist, a generalization rule is added to the verification generalization rule table 806 with the pre-generalization value set to "attribute name + index number" and the generalized value set to "attribute name". On the other hand, if it exists in the generalization rule, a generalization rule is added to the verification generalization rule table 806 with the pre-generalization value set to "attribute name + index number" and the generalized value set to "the generalized value of the attribute in question".

[0089] The anonymization processing unit 802 refers to the anonymization request for attribute i and determines whether the anonymization request for attribute i is "pseudonymization" or not (S1606). If "pseudonymization" is requested (S1606: Yes), the verification data generation process (S1506) is terminated. On the other hand, if the request is for something other than "pseudonymization," that is, if there is no anonymization request (S1606: No), the anonymization processing unit 802 copies the index data for attribute i to the verification index data table 804 (S1607) and terminates the verification data generation process (S1506).

[0090] <Signature verification process (step S1012)> Figure 17 is a flowchart showing a detailed example of the signature verification process (step S1110) shown in Figure 11.

[0091] The anonymized data user terminal 103 calculates a hash value (hereinafter referred to as element hash value) EH for each cell of the anonymized data table 803 obtained in step S1109 and stores it in the element hash value list (S1701). The element hash value list contains the same number of element hash values ​​(EH_11, EH_12, ..., EH_1N, EH21, ..., EH_MN) as the number of cells in the anonymized data table 803, M × N (M is the number of attributes, N is the number of records). The detailed processing of the element hash value generation process will be described later in Figure 18.

[0092] The signature verification unit 1002 generates an anonymized overall hash value TH' representing the entire anonymized data from the element hash value list (step S1702). The anonymized overall hash value generation process is the same as the process of generating the hash value TH of the entire sensitive data table 303 from the element hash value list (S1304). The sum of the element hash values ​​included in the element hash value list is input to the hash value generation unit 302 to generate the overall hash value TH'.

[0093] The signature verification unit 1002 decrypts the signature value 1003 for all sensitive data in the sensitive data table 303 using the verification key 1004 and obtains the overall hash value TH (step S1703).

[0094] The signature verification unit 1002 determines whether the decrypted overall hash value TH matches the anonymized overall hash value TH' (step S1704). If they match (step S1704: Yes), the signature verification unit 1002 outputs "Verification successful" to the output device 204 (step S1705). On the other hand, if they do not match (step S1704: No), the signature verification unit 1002 outputs "Verification failed" to the output device 204 (step S1706), and the signature verification process (step S1110) ends.

[0095] <Element hash value generation process S1701> Figure 18 is a flowchart showing a detailed example of the processing procedure for the element hash value generation process (step S1701) shown in Figure 17.

[0096] The signature verification unit 1002 reads the cell value Data_O from the anonymized data table 803 to generate an element hash value, and the random number R_O corresponding to Data_O from the verification random number data table 805, and determines whether Data_O is blank or not (S1801). If Data_O is blank (S1801: Yes), the element hash value generation process is terminated with R_O as the element hash value.

[0097] On the other hand, if Data_O is not blank (S1801: No), the signature verification unit 1002 determines whether or not index data for the attributes of Data_O exists (S1802). If index data exists (S1802: Yes), it refers to the verification index data table 804 and performs the same processing as in S1401 to generate pseudonymized data Data_P and a random number R_P corresponding to the pseudonymized data Data_P, and assigns Data_P to Data, which represents the latest data, and R_P to the random number R corresponding to Data (S1803). On the other hand, if index data does not exist (S1802: No), it assigns Data_O to Data and R_O to R (S1804).

[0098] The signature verification unit 1002 determines whether a generalized rule table exists for the attributes of Data_O (S1805). If a generalized rule table does not exist (S1805: No), the process proceeds to S1808.

[0099] On the other hand, if a generalization rule exists (S1805:Yes), the signature verification unit 1002 determines whether Data is generalizable (S1806). If Data is not generalizable (S1806:No), the process moves to S1808. On the other hand, if Data is generalizable (S1806:Yes), the signature verification unit 1002 performs the same generalization process as in S1406 to generate Data_G and R_G, and assigns Data_G to Data and R_G to R (S1807). After the process in S1807, the process moves to S1806.

[0100] The signature verification unit 1002 inputs the concatenated value of Data and R to the hash value generation unit 302, uses the obtained hash value as the element hash value (S1808), and terminates the element hash value generation process.

[0101] The anonymized data user terminal 103 can generate the same element hash value list as the one generated by the signature generator terminal 101 in the signature generation process S1303, regardless of the anonymization request, by calculating element hash values ​​for all attribute values ​​in the anonymized data table 803 in this manner.

[0102] As explained above, if the overall hash value TH generated from the sensitive data table 303 by the signature generation process (step S1101) described in Figures 13 to 14 matches the anonymized overall hash value TH' generated by the signature verification process (step S1110) shown in Figures 17 to 18 on the anonymized data table 803 that has undergone the anonymization process (step S1108) described in Figures 15 to 16, then it becomes possible to verify that no fraudulent anonymization process (step S1008) has been performed, that is, to verify the legitimacy of the anonymization.

[0103] Furthermore, in hash value generation, by concatenating random numbers to the original cell value and then hashing it, the amount of computation required to identify the original data from the hash value becomes enormous, thus reducing the risk of the original data being reconstructed from the anonymized data.

[0104] The signature generator terminal 101 generated the signature value 1003 using the overall hash value TH and the signing key 310 to ensure security (step S1305). However, if security is ensured by other means, the signature generator terminal 101 may send the overall hash value TH directly to the anonymized data providing server 102. In this case, the anonymized data providing server 102 will send the overall hash value TH to the anonymized data user terminal 103.

[0105] Furthermore, in the above-described embodiment, the anonymization system was configured with an anonymization data providing device (a configuration including a signature generator terminal 101, which is the first device, and an anonymization data providing server 102, which is the second device), and an anonymization data user device (an anonymization data user terminal 103, which is the third device). An example was described in which the signature generator terminal 101 and the anonymization data providing server 102 are different computers that can communicate with each other.

[0106] However, the anonymized data provision server 102 may perform the processing (step S1101) that the signature generator terminal 101 performs. In other words, in the anonymization system 100, the signature generator terminal 101 and the anonymized data provision server 102 may be configured on the same computer. In this case, data transmission (step S1102) and data acquisition (step S1103) become unnecessary.

[0107] Next, the second embodiment will be described with reference to Figures 19-29. In the first embodiment, an example was described in which cells within sensitive data are used as elements, but the second embodiment is an example of an anonymization system targeting text data. Note that in describing the second embodiment, explanations similar to those in the first embodiment may be omitted.

[0108] <Anonymization System 1900> Figure 19 is an explanatory diagram showing an example of the system configuration of an anonymization system. The anonymization system 1900 includes a signature generator terminal 1901, an anonymized data providing server 1902, and an anonymized data user terminal 1903.

[0109] The signature generator terminal 1901, the anonymized data provider server 1902, and the anonymized data user terminal 1903 are connected to each other so that they can send and receive information via the network 104. The hardware configuration examples for the signature generator terminal 1901, the anonymized data provider server 1902, and the anonymized data user terminal 1903 are the same as those for the signature generator terminal 101, the anonymized data provider server 102, and the anonymized data user terminal 103.

[0110] <Example of functional configuration of signature generator terminal 1901> Figure 20 is a block diagram showing an example of the functional configuration of the signature generator terminal 1901. The signature generator terminal 1901 includes a named entity recognition unit 2001, an auxiliary data generation unit 2002, a signature generation unit 2003, a hash value generation unit 302, sensitive text data 2004, auxiliary data 2005, a generalized rule table 306, signed sensitive text data 2006, and a signature key 308.

[0111] The named entity recognition unit 2001, the auxiliary data generation unit 2002, and the signature generation unit 2003 are specifically implemented, for example, by causing the processor 201 to execute a program stored in the storage device 202 shown in Figure 2. Furthermore, the sensitive text data 2004, the auxiliary data 2005, and the signed sensitive text data 2006 are stored in the storage device 202.

[0112] The named entity recognition unit 2001 extracts named entities contained in the sensitive text data 2004. The auxiliary data generation unit 2002 generates auxiliary data 2005 from the list of named entities extracted by the named entity recognition unit 2001. The signature generation unit 2003 generates a signature value for the sensitive text data 2004.

[0113] Sensitive text data 2004 is text data that includes sensitive data such as personal information. In this embodiment, as an example, the text data is "My name is Taro Hitachi."

[0114] The Auxiliary Data List 2005 stores the start position, end position, classification, index number for identifying the kana-coded data, and random number for all sensitive data contained in the Sensitive Text Data 2004. For example, for the text data "My name is Hitachi Taro," one auxiliary data entry would be stored: "Start position=5, End position=8, Classification=Person's name, Index number=1, Random number=01F399." The classification is information about the classification of attributes, and the classification name can include, for example, a person's name, full name, address, age, gender, etc.

[0115] Signed sensitive text data 2006 is text data to which a signature value has been added compared to sensitive text data 2004. Signed sensitive text data 2004 is sensitive text data to which signature information has been added to the end, resulting in text data such as "My name is Hitachi Taro. ===BEGIN SIGN===9138BAC825===END SIGN===". Here, "===BEGIN SIGN===9138BAC825===END SIGN===" is the signature information.

[0116] <Example of functional configuration for anonymized data provision server 1902> Figure 21 is a block diagram showing an example of the functional configuration of the anonymized data providing server 1902. The anonymized data providing server 1902 includes a web server function 2101, an anonymization processing unit 2102, a hash value generation unit 302, auxiliary data 2005, a generalized rule table 306, signed sensitive text data 2006, signed anonymized text data 2103, verification auxiliary data 2104, and a verification generalized rule table 806.

[0117] The web server function 2101 and the anonymization processing unit 2102 are specifically implemented, for example, by having the processor 201 execute a program stored in the storage device 202 shown in Figure 2. The anonymized text data 2103 and the verification auxiliary data 2104 are also stored in the storage device 202.

[0118] The web server function 2101 maintains a web page containing identification information for sensitive text data, making it accessible from the anonymized data user terminal 1903. The anonymization processing unit 2102 performs verifiable anonymization processing on the sensitive text data 2006.

[0119] The signed anonymized text data 2103 is the result of anonymizing the signed sensitive text data by the anonymization processing unit 2102. For example, if the name of a person is removed from "My name is Hitachi Taro. ===BEGIN SIGN===9138BAC825===END SIGN===", it becomes "My name is [ ]. ===BEGIN SIGN===9138BAC825===END SIGN===".

[0120] The verification auxiliary data list 2104 stores the auxiliary data necessary for signature verification from the auxiliary data list.

[0121] <Example of functional configuration of anonymized data user terminal 1903> Figure 22 is a block diagram showing an example of the functional configuration of an anonymized data user terminal 1903. The anonymized data user terminal 1903 includes a web browser function 2201, a signature verification unit 2202, a hash value generation unit 302, signed anonymized text data 2103, verification auxiliary data 2104, a verification generalized rule table 806, and a verification key 1004.

[0122] The web browser function 2201 and the signature verification unit 2202 are specifically implemented, for example, by causing the processor 201 to execute a program stored in the storage device 202 shown in Figure 2.

[0123] The web browser function 2201 receives the web page published by the anonymized data provision server 1902 and displays it on the anonymized data user terminal 1903. The signature verification unit 2202 verifies the legitimacy of the anonymization using the signed anonymized text data 2103, the verification auxiliary data 2104, the verification generalized rule table 806, and the verification key 1004 as input.

[0124] <Sequence of anonymization system 1900> Figure 23 is a sequence diagram of the anonymization system 1900. The signature generator terminal 1901 takes sensitive text data, generalization rules, and the signature key 310 as input, generates a signature value 1003 for the sensitive text data, and performs signature generation processing to attach it to the sensitive text data (step S2301). Details of the signature generation processing (step S2301) will be described later using Figures 24 to 27.

[0125] Next, the signature generator terminal 1901 transmits the auxiliary data, generalization rules, and the signed sensitive text data and auxiliary data generated by the signature generation process (step S2301) to the anonymized data providing server 102 (step S2302).

[0126] The anonymized data providing server 1902 retrieves the signed sensitive text data, auxiliary data, and generalization rules transmitted from the signature generating terminal 1901 and stores them in the storage device 202 (step S2303).

[0127] Next, the anonymized data provision server 1902 generates a web page accessible via the network 104 using its web server function 2101, and notifies the web browser function 2201 of the anonymized data user terminal 1903 of the web page's URL (step S2304).

[0128] Next, the anonymized data user terminal 1903 uses the web browser function 2201 to obtain a web page containing the identification information of the signed sensitive text data (step S2305).

[0129] Next, the anonymized data user terminal 1903 sets an anonymization request by operating the data user's input device 203 and sends an anonymized data acquisition request, including the anonymization request, to the Web server function 2101 of the anonymized data provision server 1902 (step S2306). In this embodiment, an example is given in which the anonymized data acquisition request includes an anonymization request to "delete names".

[0130] Next, the anonymized data provision server 1902, using the Web server function 2101, obtains the identification information and anonymization request for the signed sensitive text data notified from the anonymized data user terminal 1903, and passes them to the anonymization processing unit 2102. Then, the anonymized data provision server 1902, using the anonymization processing unit 2102, takes the signed sensitive text data and anonymization request as input to perform anonymization processing and generates signed anonymized text data 2103, verification auxiliary data 2104, and verification generalization rules, and passes these to the Web server function 2101 (step S2307). Details of the anonymization processing (step S2307) will be described later in Figure 28.

[0131] The web server function 2101 registers the signed anonymized text data, verification auxiliary data, and verification generalization rules on a downloadable web page for data users, and notifies the URL of this page to the web browser function 2201 of the anonymized data user terminal 1903.

[0132] Next, the anonymized data user terminal 1903 uses its web browser function 2201 to access a download web page for data users by inputting the notified URL, downloads and obtains the signed anonymized text data 2103, the verification auxiliary data 2104, and the verification generalization rule, and stores them in the memory device 202 of the anonymized data user terminal 1903 (step S2308).

[0133] Finally, the anonymized data user terminal 1903, using the signature verification processing unit 2202, takes the signed anonymized text data 2103, the verification auxiliary data 2104, the verification generalization rule, and the verification key 1004 as input and executes a signature verification process to verify the legitimacy of the anonymization process (step S2307) performed by the anonymized data providing server 1902 (step S2309). The anonymized data user terminal 1903 displays "Verification successful" on the output device 204 if the legitimacy of the anonymization process (step S2307) is verified, or "Verification failed" if the legitimacy of the anonymization process (step S2308) is not verified (step S2309).

[0134] As shown in Figure 23, users of anonymized data can verify whether the signed anonymized text data they obtained has been properly anonymized through legitimate anonymization processing (step S2307). This prevents the provision of services that utilize analysis results obtained from fraudulently anonymized data.

[0135] <Signature generation process (step S2301)> Figure 24 is a flowchart showing a detailed example of the signature generation process (S2301) shown in Figure 24.

[0136] The signature generation unit 2003 performs named entity extraction processing (S2401). In the named entity extraction processing, a list of named entity data is obtained for the text to be signed. The named entity data list contains zero or more named entity data and is sorted in ascending order of the starting position of the named entity. The named entity data includes at least the starting position, ending position, and classification of the named entity. The named entity extraction processing may be performed using a known named entity extraction tool, by the signature generator, or by a combination of a named entity extraction tool and manual extraction.

[0137] The signature generation unit 2003 executes the auxiliary data list generation process (S2402). In the auxiliary data generation process, the text to be signed is divided using a named entity data list, and a list is generated that contains one or more auxiliary data items for each divided named entity or non-named entity string (element). The detailed process of the auxiliary data list generation process will be described later in Figure 25.

[0138] The signature generation unit 2003 generates an element hash value list by referring to the sensitive text data to be signed and the auxiliary data list generated in S2402 (S2403). The detailed process of generating the element hash value list will be described later in Figure 27.

[0139] The signature generation unit 2003 executes the overall hash value generation process (S1304).

[0140] Finally, the signature generation unit 2003 generates a signature using the overall hash value TH and the signing key 308, and stores the signed sensitive text data with the signature attached (step S2405).

[0141] <Auxiliary data generation process> Figure 25 is a flowchart showing a detailed example of the processing procedure for the auxiliary data generation process (S2402) shown in Figure 24.

[0142] The signature generation unit 2003 generates an index map (S2501) using the named entity data obtained in S2401. The detailed process of index map generation will be described later in Figure 26.

[0143] Next, the signature generation unit 2003 generates an empty auxiliary data list (S2502).

[0144] Next, the signature generation unit 2003 determines whether the number N of unique expressions obtained in S2401 is 0 (S2503). If N = 0 (S2503: Yes), the signature generation unit 2004 adds auxiliary data with the start position being 0 and the end position being txt_len - 1 to the auxiliary data list (S2504), and ends the auxiliary data generation process (S2402). Here, txt_len represents the number of characters in the confidential text data.

[0145] On the other hand, if N ≠ 0 (S2503: No), the signature generation unit 2003 determines whether the start position of the 0th unique expression is 0 (S2505). If the start position of the 0th unique expression = 0 (S2505: Yes), the process proceeds to S2507. On the other hand, if the start position of the 0th unique expression ≠ 0 (S2505: No), auxiliary data with the start position being 0 and the end position being the start position of the 0th unique expression - 1 is added to the auxiliary data list (S2506).

[0146] The signature generation unit 2003 initializes a variable i indicating which unique expression in the unique expression data to 0 (S2507).

[0147] The signature generation unit 2003 determines whether i < N (S2508). If i < N is not true (S2508: No), the process proceeds to S2513. On the other hand, if i < N is true (S2508: Yes), the signature generation unit 2003 generates auxiliary data for the classification of the i-th unique expression and adds it to the auxiliary data list (S2509).

[0148] The auxiliary data for the unique expression is composed of the position, classification, random number, and index number of the unique expression. The start position, end position, and classification of the unique expression are obtained from the unique expression data obtained in S2401. The index number is obtained by referring to the index map generated in S2501 for the index number paired with the character string of the unique expression. The random number is generated each time the process of S2508 is executed.

[0149] Next, the signature generation unit 2003 determines whether i ≠ N - 1 and the start position of the (i + 1)-th unique expression ≠ the end position of the i-th unique expression + 1 (S2510). If i ≠ N - 1 and the start position of the (i + 1)-th unique expression ≠ the end position of the i-th unique expression + 1 is not satisfied (S2510: No), the process proceeds to S2512. On the other hand, if i ≠ N - 1 and the start position of the (i + 1)-th unique expression ≠ the end position of the i-th unique expression + 1 (S2510: Yes), auxiliary data with the start position set to the end position of the i-th unique expression + 1 and the end position set to the start position of the (i + 1)-th unique expression - 1 is stored in the auxiliary data list (S2511).

[0150] The signature generation unit 2003 increments the variable i (S2512), and the process proceeds to S2508.

[0151] The signature generation unit 2003 determines whether the end position of the (N - 1)-th unique expression = txt_len - 1 (S2513). If the end position of the (N - 1)-th unique expression = txt_len - 1 (S2513: Yes), the auxiliary data generation process (S2401) ends. On the other hand, if the end position of the (N - 1)-th unique expression ≠ txt_len - 1 (S2513: No), auxiliary data with the start position set to the end position of the (N - 1)-th unique expression + 1 and the end position set to txt_len - 1 is stored in the auxiliary data list (S2514), and the auxiliary data generation process (S2401) ends.

[0152] <Index Map Generation Process> FIG. 26 is a flowchart showing a detailed processing procedure example of the index map generation process (S2501) shown in FIG. 25.

[0153] The signature generation unit 2003 initializes a variable i indicating which unique expression it is among the unique expression data to 0 (S2601).

[0154] Next, the signature generation unit 2003 determines whether i < N (S2602). If i < N is not satisfied (S2602: No), the index map generation process (S2501) ends.

[0155] On the other hand, when i < N (S2602: Yes), the signature generation unit 2003 determines whether there is an index map for the classification of the i-th unique expression (S2603). For example, when the i-th unique expression is "Hitachi Taro", it is determined whether there is an index map related to "name". If there is an index map for the classification of the i-th unique expression (S2603: Yes), the process proceeds to S2605. On the other hand, if there is no index map for the classification of the i-th unique expression (S2603: No), the signature generation unit 2003 generates an empty index map for the classification of the i-th unique expression (S2604), and the process proceeds to S2606.

[0156] The signature generation unit 2003 determines whether the character string of the i-th unique expression exists in the index map for the classification of the i-th unique expression (S2605). If the character string of the i-th unique expression exists (S2605: Yes), the process proceeds to S2607. On the other hand, if the character string of the i-th unique expression does not exist (S2605: No), the process proceeds to S2606.

[0157] The signature generation unit 2003 adds a pair of the character string of the i-th unique expression and an index number to the index map for the classification of the i-th unique expression (S2606). Here, the index number is the number of pairs of character strings and index numbers stored in the index map for the classification of the i-th unique expression. That is, when the index map is empty, the index number is 0, and when three pairs are stored, the index numbers of each pair are (1, 2, 3).

[0158] The signature generation unit 2003 increments the variable i (S2607) and proceeds to S2602.

[0159] <Element hash value list generation process> FIG. 27 is a flowchart showing a detailed processing procedure example of the element hash value list generation process (S2403) shown in FIG. 24.

[0160] The signature generation unit 2003 generates an empty element hash list (S2701).

[0161] The signature generation unit 2003 initializes a variable j indicating which expression in the auxiliary data to 0 (S2702).

[0162] The signature generation unit 2003 determines whether j < M (S2703). M is the number of elements in the confidential text data (the total number of strings of specific expressions and non-specific expressions). If j < M is not satisfied (S2703: No), the element hash value list generation process (S2403) ends. On the other hand, if j < M (S2703: Yes), the signature generation unit 2003 extracts an element character string from the confidential text data using the start position and end position of the j-th auxiliary data (S2704).

[0163] The signature generation unit 2003 determines whether the j-th auxiliary data in the auxiliary data contains a random number (S2705). If the j-th auxiliary data contains a random number (S2705: Yes), the element hash value generation process shown in FIG. 14 is performed using the element character string and the j-th auxiliary data, and the obtained element hash value is added to the element hash value list (S2706). On the other hand, if the j-th auxiliary data does not contain a random number (S2705: No), the element character string is input to the hash generation unit, and the obtained hash value is added to the element hash value list (S2707).

[0164] The signature generation unit 2003 increments the variable j (S2708) and transfers to S2703.

[0165] <Anonymization process (step S2307)> FIG. 28 is a flowchart showing a detailed processing procedure example of the anonymization process (S2307) shown in FIG. 23. The anonymization process is performed on the part of the confidential text data other than the signature in the signed confidential text data, and signed anonymized text data is generated by attaching the signature of the signed confidential text data to the anonymized text data obtained as a result of the anonymization process.

[0166] The anonymization processing unit 2102 generates an empty anonymized element string list and empty verification auxiliary data (S2801).

[0167] The anonymization processing unit 2102 initializes to 0 a variable j indicating which expression among the auxiliary data and a len indicating the number of characters of the anonymized text data (S2802).

[0168] The anonymization processing unit 2102 determines whether j < M (S2803). M is the number of elements in the confidential text data (the total number of strings of specific expressions and the number of strings of non-specific expressions).

[0169] If j < M is not satisfied (S2803: No), the anonymization processing unit 2102 generates anonymized text data obtained by concatenating all the element strings stored in the element string list (S2809), and ends the anonymization processing (S2307).

[0170] On the other hand, if j < M (S2703: Yes), the anonymization processing unit 2102 extracts an element string from the confidential text data using the start position and end position of the j-th auxiliary data (S2804).

[0171] The anonymization processing unit 2102 determines whether processing for the classification of the j-th auxiliary data is required according to step S2306 (step S2805).

[0172] If processing is not required (step S2805: No), the anonymization processing unit 802 adds the element string to the anonymized element string list and the j-th auxiliary data to the verification auxiliary data, respectively, and adds the number of characters of the element string to len (S2806), and proceeds to S2808.

[0173] On the other hand, if processing is requested (step S2805: Yes), anonymization processing is performed on the element string (step S2807). Specifically, first the anonymization processing unit 2102 determines which of the following processing is requested for the element string: "pseudonymization," "generalization," or "deletion."

[0174] If "pseudonymization" is requested, the anonymization processing unit 2102 performs the pseudonymization process S1401 in the element hash value generation process shown in Figure 14, and adds an anonymized element string to the anonymized element string list by concatenating an anonymization identification string indicating that it is an anonymized string to the beginning and end of the pseudonymized data Data_P. In addition, auxiliary data with the start position set to "len", the end position to "len + anonymized element string - 1", the classification to "classification of the j-th auxiliary data", and the random number to "random number R_P corresponding to Data_P" is added to the verification auxiliary data. Furthermore, the number of characters in the anonymized element string is added to len.

[0175] If "generalization" is requested, the anonymization processing unit 2102 performs the S1401-S1407 processes of the element hash value generation process shown in Figure 14, but changing the determination in S1404 from "whether further generalization is possible" to "whether the generalization level of the anonymization request has been reached". An anonymized element string is added to the anonymized element string list by concatenating an anonymization identification string indicating that it is an anonymized string to the beginning and end of the generalized data Data_G obtained as a result of the processing. In addition, auxiliary data with the start position set to "len", the end position to "len + anonymized element string - 1", and the random number set to "random number R_G corresponding to Data_G" is added to the auxiliary data for verification. Furthermore, the number of characters in the anonymized element string is added to len.

[0176] If "deletion" is requested, the anonymization processing unit 2102 performs all of S1401 to S1408 of the element hash value generation process shown in Figure 14. An anonymized element string is added to the anonymized element string list by concatenating an anonymization identification string indicating that it is an anonymized string to the beginning and end of the deleted data Data_R obtained as a result of S1408. In addition, auxiliary data is added to the verification auxiliary data with the start position set to "len", the end position to "len + number of characters in the anonymized element string - 1", and the random number set to "random number R_G corresponding to Data_G". Furthermore, the number of characters in the anonymized element string is added to len.

[0177] The anonymization processing unit 2102 increments the variable j (step S2808) and returns to step S2803. The processing in S2307 generates auxiliary verification data, which includes random numbers for anonymized data that are generated when the information is processed in response to the anonymization request.

[0178] <Signature verification process (step S2309)> Figure 29 is a flowchart showing a detailed example of the signature verification process (step S2309) shown in Figure 23.

[0179] The anonymized data user terminal 1903 generates a list of element hash values ​​for the anonymized text data using the anonymized text data portion of the signed anonymized text data 2103 obtained in step S2308 and the verification auxiliary data 2104 (step S2901). Specifically, for an empty list of element hash values, the terminal sequentially references the M auxiliary data contained in the verification auxiliary data and adds the element hash values ​​obtained by the following process to the list of element hash values.

[0180] First, the signature verification unit 2202 extracts element strings from the anonymized text data using the start and end positions of the j-th auxiliary data, and then removes the anonymized identifier string from the extracted element strings. It then determines whether the j-th auxiliary data contains a "random number," and if it does not, it performs the same element hash value generation process as in S2707 in Figure 27. If it does contain a "random number," it performs the same process as the element hash value generation process shown in Figure 18. However, in this embodiment, the references for the index number and random number are changed from index data (index map) and random number data to auxiliary data.

[0181] The anonymized data user terminal 1903 generates the overall hash value H' of the anonymized text data from the element hash value list using the same process as in S1704 described above (step S2902). The processing from S2903 onward is the same as S1704 to S1706 in Figure 17.

[0182] As explained above, if the overall hash value H generated from the sensitive text data 2004 by the signature generation process (step S2301) described in Figures 24 to 27 matches the anonymized overall hash value H' generated by the signature verification process (step S2309) shown in Figure 29 on the signed anonymized text data that has undergone the anonymization process (step S2307) described in Figure 28, it becomes possible to verify that no fraudulent anonymization process (step S2307) has been performed, that is, to verify the legitimacy of the anonymization.

[0183] Furthermore, in this embodiment, as in the first embodiment, the anonymization system is configured with an anonymization data providing device (a device including a signature generator terminal 1901, which is the first device, and an anonymization data providing server 1902, which is the second device) and an anonymization data user device (an anonymization data user terminal 1903, which is the third device). An example was described in which the signature generator terminal 1901 and the anonymization data providing server 1902 are different computers that can communicate with each other.

[0184] However, the anonymized data provision server 1902 may perform the processing (step S2301) that is performed by the signature generator terminal 1901. In other words, in the anonymization system 1900, the signature generator terminal 1901 and the anonymized data provision server 1902 may be configured on the same computer. In this case, data transmission (step S2302) and data acquisition (step S2303) become unnecessary.

[0185] It should be noted that the present invention is not limited to the embodiments described above, but includes various modifications and equivalent configurations within the spirit of the attached claims. For example, the embodiments described above are described in detail to make the present invention easier to understand, and the present invention is not necessarily limited to having all of the described configurations. Furthermore, some of the configurations of one embodiment may be replaced with those of another embodiment. Furthermore, some of the configurations of one embodiment may be added to those of another embodiment. Furthermore, some of the configurations of each embodiment may be added, deleted, or replaced with other configurations.

[0186] Furthermore, each of the aforementioned configurations, functions, processing units, and processing means may be implemented in hardware, for example, by designing them as integrated circuits, or they may be implemented in software by having a processor interpret and execute programs that realize each function.

[0187] Information such as programs, tables, and files that implement each function can be stored in storage devices 202 such as memory, hard disks, and SSDs (Solid State Drives), or on recording media such as IC (Integrated Circuit) cards, SD cards, and DVDs (Digital Versatile Discs).

[0188] Furthermore, the control lines and information lines shown are those deemed necessary for explanation purposes and do not necessarily represent all control lines and information lines required for implementation. In reality, it can be assumed that almost all components are interconnected.

[0189] In the first embodiment, the index number 502 of the name in the index data table 542 in the signature generator terminal 101 and the anonymized data provision server 102 may be editable.

[0190] For example, if the name (name) attribute in sensitive data table 303 contains elements "Hitachi Taro" and "Hitachi," and these elements identify the same person and do not need to be distinguished, then index number 502 may be edited so that both can be given the same pseudonym. In this case, the user can, for example, determine that they are the same content in sensitive data table 303 and edit to assign the same index number to "Hitachi Taro" and "Hitachi." Furthermore, if the same index number is assigned to multiple elements with the same attribute, each computer (101, 102, 103) may generate and process the random number R_P of the pseudonymized data of these multiple elements with the same value. Here, for example, the same random number R_P may be generated by substituting the random number R_P of one element with the random number R_P of another element among the random number R_P of the pseudonymized data of multiple elements.

[0191] In the second embodiment, the index numbers of the index map of the auxiliary data (2004, 2005) in the signature generator terminal 1901 and the anonymized data providing server 1902 may be editable.

[0192] For example, if sensitive text data contains the strings "Hitachi Taro" and "Hitachi," and both are named entities that identify the same person and do not need to be distinguished, the index numbers in the name classification index map may be edited so that they can both be given the same pseudonym. In this case, the user can, for example, determine that they are the same content in the sensitive text data and edit them to assign the same index number to "Hitachi Taro" and "Hitachi." Furthermore, if the same index number is assigned to multiple named entities in the same classification, each computer (1901, 1902, 1903) may generate and process the same random number R_P for the pseudonymized data of these multiple named entities. Here, for example, the same random number R_P may be generated by substituting the random number R_P of one element with the random number R_P of another element among the random number R_P of the pseudonymized data of multiple named entities.

[0193] In the second embodiment, similar to the anonymization request setting 1200 in the first embodiment, an anonymization request setting of unprocessed, pseudonymized, deleted, or generalized may be made for each classification of named entities. [Explanation of Symbols]

[0194] 100 Anonymization Systems 101 Signature Generator Terminal 102 Anonymized Data Provider Server 103 Anonymized Data User Terminal 301 Signature generation section 302 Hash Value Generation Unit 303 Sensitive Data Table 304 Index Data Table 305 Random Number Data Table 306 Generalized Rule Table 307 Signature Value Table 308 Signing Key 803 Anonymized Data Table 804 Verification Index Data Table 805 Random number data table for verification 806 Generalized rule table for verification

Claims

1. An anonymization system comprising an anonymization data providing device and an anonymization data user device capable of communicating with the anonymization data providing device, The anonymized data providing device is When generating values ​​to be used for data validation, for each element in the data, the following processes are performed: a first process that generates pseudonymized data by replacing the original data within the element with an identifiable value, and a random number for pseudonymized data based on the original data and a random number for the original data; a second process that generates generalized data by replacing the original data within the element with a generalized value, and a random number for generalized data based on the pseudonymized data and the random number for the pseudonymized data; a third process that generates deleted data by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and the random number for the generalized data. When anonymizing data in response to an anonymization request from the anonymized data user device, the necessary processing for the anonymization request is performed on each element of the data from the first, second, and third processing, and anonymized data, which is the anonymized data, and random numbers for anonymized data, which are random numbers based on the processing, are generated. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained via communication. During data verification, for each element within the anonymized data, a process from the first, second, and third processes that the anonymized data provider does not perform is performed based on the random number for anonymized data, and a random number for deleted data is generated. The validity of the acquired anonymized data is verified based on the values ​​used for the aforementioned data verification and the generated random numbers for the deleted data. An anonymization system characterized by the following:

2. An anonymization system comprising an anonymization data providing device and an anonymization data user device capable of communicating with the anonymization data providing device, The anonymized data providing device is When generating values ​​to be used for data validation, for each element in the data, the following processes are performed: a first process that generates pseudonymized data by replacing the original data within the element with an identifiable value, and a random number for pseudonymized data based on the original data and a random number for the original data; a second process that generates generalized data by replacing the original data within the element with a generalized value, and a random number for generalized data based on the pseudonymized data and the random number for the pseudonymized data; a third process that generates deleted data by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and the random number for the generalized data. When anonymizing data in response to an anonymization request from the anonymized data user device, the necessary processing for the anonymization request is performed on each element of the data from the first, second, and third processing, and anonymized data, which is the anonymized data, and random numbers for anonymized data, which are random numbers based on the processing, are generated. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained via communication. During data verification, for each element within the anonymized data, one of the first, second, or third processes that the anonymized data providing device does not perform is performed to generate a random number for deletion data. The validity of the acquired anonymized data is verified based on the random numbers for deletion included in the random numbers for anonymized data generated by the anonymized data providing device, and the generated random numbers for deletion. An anonymization system characterized by the following:

3. An anonymization system according to Claim 1 or Claim 2, The anonymized data providing device is A first device that generates values ​​used for data validation, The system comprises a second device for generating the anonymized data and random numbers for the anonymized data, The first apparatus and the second apparatus are These are different computers connected to each other and capable of communication. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained from the second device. An anonymization system characterized by the following:

4. An anonymization system according to claim 1 or claim 2, The anonymized data user device is In the aforementioned anonymization request, it is possible to select one of the following options for each attribute of the data: unprocessed, pseudonymized, deleted, or generalized. An anonymization system characterized by the following:

5. An anonymization system according to claim 1 or claim 2, The anonymized data providing device and the anonymized data user device are In the first process described above, pseudonymized data is generated based on the attribute name of the data and the index number assigned to each element in the same attribute. An anonymization system characterized by the following:

6. An anonymization system according to claim 1 or claim 2, The anonymized data providing device and the anonymized data user device are In the first process described above, the pseudonymized data is generated based on the attribute name of the data and the index number assigned to each element of the same attribute, which can be edited by the user. For elements with the same index number, generate the same random number for pseudonymized data. An anonymization system characterized by the following:

7. An anonymization system comprising an anonymization data providing device and an anonymization data user device capable of communicating with the anonymization data providing device, The anonymized data providing device is When generating values ​​to be used for data verification, the following processes are performed for each named entity in the text data: a first process that generates kana-type data by replacing the original data with an identifiable value, and a random number for kana-type data based on the original data and a random number for the original data; a second process that generates generalized data by replacing the original data with a generalized value, and a random number for generalized data based on the kana-type data and the random number for the kana-type data; a third process that generates deleted data by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and the random number for the generalized data; and a fourth process that generates a hash value of the original data for each non-named entity in the text data. When anonymizing data in response to an anonymization request from the anonymized data user device, the necessary processing for the anonymization request is performed on each named entity in the data from the first, second, and third processing, and anonymized data, which is the anonymized data, and random numbers for anonymized data, which are random numbers based on the processing, are generated. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained via communication. During data verification, for each named entity in the anonymized data, a process not performed by the anonymized data provider among the first, second, and third processes is performed based on the random number for anonymized data to generate a random number for deleted data, and for each non-named entity in the anonymized data, a hash value of the original data is generated. The validity of the acquired anonymized data is verified based on the values ​​used for the aforementioned data verification, the generated random numbers and hash values ​​for the deleted data, and so on. An anonymization system characterized by the following:

8. An anonymization system comprising an anonymization data providing device and an anonymization data user device capable of communicating with the anonymization data providing device, The anonymized data providing device is When generating values ​​to be used for data verification, the following processes are performed for each named entity in the text data: a first process that generates kana-type data by replacing the original data with an identifiable value, and a random number for kana-type data based on the original data and a random number for the original data; a second process that generates generalized data by replacing the original data with a generalized value, and a random number for generalized data based on the kana-type data and the random number for the kana-type data; a third process that generates deleted data by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and the random number for the generalized data; and a fourth process that generates a hash value of the original data for each non-named entity in the text data. When anonymizing data in response to an anonymization request from the anonymized data user device, the necessary processing for the anonymization request is performed on each named entity in the data from the first, second, and third processing, and anonymized data, which is the anonymized data, and random numbers for anonymized data, which are random numbers based on the processing, are generated. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained via communication. During data verification, for each named entity in the anonymized data, perform the processing among the first, second, and third processes that the anonymized data provider does not perform, generate a random number for deletion data, and for each non-named entity in the anonymized data, generate a hash value of the original data. The validity of the acquired anonymized data is verified based on the random numbers for deletion and hash values ​​included in the random numbers for anonymized data generated by the anonymized data providing device, and the generated random numbers for deletion and hash values. An anonymization system characterized by the following:

9. An anonymization system according to claim 7 or claim 8, The anonymized data providing device is A first device that generates values ​​used for data validation, The system comprises a second device for generating the anonymized data and random numbers for the anonymized data, The first apparatus and the second apparatus are These are different computers connected to each other and capable of communication. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained from the second device. An anonymization system characterized by the following:

10. An anonymization system according to claim 7 or claim 8, The anonymized data user device is In the aforementioned anonymization request, it is possible to select one of the following options for the classification unit of a named entity: unprocessed, pseudonymized, deleted, or generalized. An anonymization system characterized by the following:

11. An anonymization system according to claim 7 or claim 8, The anonymized data providing device and the anonymized data user device are In the first process described above, kana data is generated based on the classification name of the named entity and the index number assigned to each named entity within the same classification. An anonymization system characterized by the following:

12. An anonymization system according to claim 7 or claim 8, The anonymized data providing device and the anonymized data user device are In the first process described above, the kana data is generated based on the classification name of the named entity and the index number assigned to each named entity within the same classification, which can be edited by the user. For named entities with the same index number, generate the same random number for kana-encoded data. An anonymization system characterized by the following:

13. An anonymization method using an anonymization data providing device and an anonymization data user device capable of communicating with the anonymization data providing device, The anonymized data providing device is When generating values ​​to be used for data validation, for each element in the data, the following processes are performed: a first process that generates pseudonymized data by replacing the original data within the element with an identifiable value, and a random number for pseudonymized data based on the original data and a random number for the original data; a second process that generates generalized data by replacing the original data within the element with a generalized value, and a random number for generalized data based on the pseudonymized data and the random number for the pseudonymized data; a third process that generates deleted data by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and the random number for the generalized data. When anonymizing data in response to an anonymization request from the anonymized data user device, the necessary processing for the anonymization request is performed on each element of the data from the first, second, and third processing, and anonymized data, which is the anonymized data, and random numbers for anonymized data, which are random numbers based on the processing, are generated. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained via communication. During data verification, for each element within the anonymized data, a process from the first, second, and third processes that the anonymized data provider does not perform is performed based on the random number for anonymized data, and a random number for deleted data is generated. The validity of the acquired anonymized data is verified based on the values ​​used for the aforementioned data verification and the generated random numbers for the deleted data. An anonymization method characterized by the following:

14. An anonymization method using an anonymization data providing device and an anonymization data user device capable of communicating with the anonymization data providing device, The anonymized data providing device is When generating values ​​to be used for data validation, for each element in the data, the following processes are performed: a first process that generates pseudonymized data by replacing the original data within the element with an identifiable value, and a random number for pseudonymized data based on the original data and a random number for the original data; a second process that generates generalized data by replacing the original data within the element with a generalized value, and a random number for generalized data based on the pseudonymized data and the random number for the pseudonymized data; a third process that generates deleted data by replacing the generalized data with a blank space, and a random number for deleted data based on the generalized data and the random number for the generalized data. When anonymizing data in response to an anonymization request from the anonymized data user device, the necessary processing for the anonymization request is performed on each element of the data from the first, second, and third processing, and anonymized data, which is the anonymized data, and random numbers for anonymized data, which are random numbers based on the processing, are generated. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained via communication. During data verification, for each element within the anonymized data, one of the first, second, or third processes that the anonymized data providing device does not perform is performed to generate a random number for deletion data. The validity of the acquired anonymized data is verified based on the random numbers for deletion included in the random numbers for anonymized data generated by the anonymized data providing device, and the generated random numbers for deletion. An anonymization method characterized by the following:

15. An anonymization method according to claim 13 or claim 14, The anonymized data providing device is A first device that generates values ​​used for data validation, The system comprises a second device for generating the anonymized data and random numbers for the anonymized data, The first apparatus and the second apparatus are These are different computers connected to each other and capable of communication. The anonymized data user device is The anonymized data and the random numbers for the anonymized data are obtained from the second device. An anonymization method characterized by the following:

16. An anonymization method according to claim 13 or claim 14, The anonymized data user device is In the aforementioned anonymization request, it is possible to select one of the following options for each attribute of the data: unprocessed, pseudonymized, deleted, or generalized. An anonymization method characterized by the following:

17. An anonymization method according to claim 13 or claim 14, The anonymized data providing device and the anonymized data user device are In the first process described above, pseudonymized data is generated based on the attribute name of the data and the index number assigned to each element in the same attribute. An anonymization method characterized by the following:

18. An anonymization method according to claim 13 or claim 14, The anonymized data providing device and the anonymized data user device are In the first process described above, the pseudonymized data is generated based on the attribute name of the data and the index number assigned to each element of the same attribute, which can be edited by the user. For elements with the same index number, generate the same random number for pseudonymized data. An anonymization method characterized by the following:

Citation Information

Patent Citations

  • Certification method for authenticity of electronic document and electronic document disclosure system

    JP2007129507A

  • Anonymization system and anonymization method

    JP2020077256A

  • Audit-log integrity using redactable signatures

    US20080104407A1