Anonymization system and anonymization method
The anonymization system addresses the challenge of inconsistent identifier assignments across multiple data sets by generating pseudonymized data and performing random number and hash value processes, resulting in more accurate data validity verification.
Patent Information
- Application Number
- JP2023196897
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-30
AI Technical Summary
Existing techniques for anonymization, such as those described in Patent Document 1 and Non-Patent Document 1, focus solely on anonymizing a single data set and do not account for multiple data sets, which can lead to inappropriate results in pseudonymization, including inconsistent identifier assignments across different data sets.
The proposed anonymization system includes a device for providing anonymized data and a user device that can communicate with it. The system generates pseudonymized data by replacing original data with values identifiable across multiple text data, and performs various processes to generate random numbers and hash values for different types of anonymized data, ensuring appropriate verification of data validity.
This approach enables more accurate verification of data validity, particularly when performing pseudonymization and other information processing tasks across multiple data sets, thereby ensuring the integrity and legitimacy of anonymized data.
Smart Images

Figure 2025083161000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an anonymization system and an anonymization method.
Background Art
[0002] With the law on the protection of personal information, the utilization of anonymized processed information in which descriptions that can identify a specific individual are deleted or replaced has been progressing. In the "Act to Amend Part of the Act on the Protection of Personal Information, etc. (commonly known as the Act to Amend the Personal Information Act of 2020)" fully enforced in April 2022, "pseudonymized processed information" is newly established, with partial compliance obligations relaxed on the condition of being limited to internal analysis, and it is considered that the utilization of data will further progress.
[0003] In the utilization of data, the fact that no improper modification has been made to the data (data legitimacy) is important to ensure the legitimacy of the data utilization results. As documents disclosing techniques for verifying the legitimacy of anonymization processing, there are Patent Document 1 and Non-Patent Document 1.
[0004] Patent Document 1 discloses a technique in which, for tabular data, a signature generator selects any one of "unprocessed", "pseudonymized", "generalized", and "deleted" processing for each attribute, and when an anonymization processor performs the processing selected by the signer, signature verification is possible even after processing.
[0005] Non-Patent Document 1 discloses a technique in which, for tabular data and text data, signature verification is possible even when an anonymization processor performs any one of "unprocessed", "pseudonymized", "generalized", and "deleted" processing.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Non-Patent Documents
[0007] [Non-Patent Document 1] Yumiko Fujii et al., "Examination of an Anonymous Processing Process Assurance Method for Text Data", 2023 Symposium on Cryptography and Security (SCIS2023), 2023 / 01 [Summary of the Invention] [Problems to be Solved by the Invention]
[0008] However, both Patent Document 1 and Non-Patent Document 1 focus only on anonymization of a single data, and do not consider multiple data, which may result in inappropriate results in pseudonymization. Specifically, when there are multiple text data, different identifiers may be assigned to the same object, or conversely, the same identifier may be assigned to different objects.
[0009] Therefore, there is a problem of providing a technique for realizing more appropriate verification when performing anonymization considering multiple data. [Means for Solving the Problems]
[0010] One aspect of the anonymization system disclosed in the present application is as follows. That is, this anonymization system includes an anonymized data providing device and an anonymized data user device that can communicate with the anonymized data providing device. When generating values for data verification, the anonymized data providing device performs the following operations for each unique expression in the text data: generating pseudonymized (with name matching) data in which the original data is replaced with values that can be identified across multiple text data, and performing a first process of generating a random number for the pseudonymized (with name matching) data based on the original data and a random number used for the original data; generating pseudonymized (without name matching) data in which the original data is replaced with values that can be identified only within the text data, and performing a second process of generating a random number for the pseudonymized (without name matching) data based on the pseudonymized (with name matching) data and the random number for the pseudonymized (with name matching) data; generating generalized data in which the original data is replaced with generalized values, and performing a third process of generating a random number for the generalized data based on the pseudonymized (without name matching) data and the random number for the pseudonymized (without name matching) data; generating deleted data in which the generalized data is replaced with blanks, and performing a fourth process of generating a random number for the deleted data based on the generalized data and the random number for the generalized data; and performing a fifth process of generating a hash value of the original data for each non-unique expression in the text data. When anonymizing data in response to an anonymization request regarding the anonymization of data from the anonymized data user device, for each element in the data, the anonymization system performs the necessary process among the first process, the second process, the third process, and the fourth process according to the anonymization request, and generates anonymized data, which is the anonymized data, and a random number for the anonymized data, which is the random number based on the process. The anonymized data user device acquires the anonymized data and the random number for the anonymized data via communication. Then, during data verification, for each unique expression in the anonymized text data, the anonymized data user device performs the process that the anonymized data providing device has not performed among the first process, the second process, the third process, and the fourth process, generates a random number for the deleted data, generates a hash value of the original data for each non-unique expression in the anonymized data, and verifies the validity of the acquired anonymized data based on the random number for the deleted data and the hash value generated by the anonymized data providing device and the generated random number for the deleted data and the hash value.
Advantages of the Invention
[0011] According to the present invention, there is provided a technique capable of more appropriately verifying the validity of data even when performing various information processing (generalization, deletion) including pseudonymization in which a unified identifier is assigned across a plurality of data in anonymization. Problems, configurations, and effects other than those described above will be clarified by the following embodiments.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The embodiments are examples for explaining the present invention, and for the sake of clarity of explanation, appropriate omissions and simplifications have been made. The present invention can be implemented in various other forms. Unless otherwise particularly limited, each component may be singular or plural. The positions, sizes, shapes, ranges, etc. of the respective components shown in the drawings may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings. As examples of various types of information, it may be described using expressions such as "table", "list", "queue", etc., but the various types of information may be represented by data structures other than these. For example, various types of information such as "XX table", "XX list", "XX queue", etc. may be referred to as "XX information". When explaining identification information, expressions such as "identification information", "identifier", "name", "ID", "number", etc. are used, but these can be mutually replaced. When there are a plurality of components having the same or similar functions, they may be described by attaching different subscripts to the same reference numeral. Also, when it is not necessary to distinguish these plurality of components, the subscript may be omitted in the description. In the embodiments, there may be a description of the processing performed by executing a program. Here, the computer executes the program by a processor (for example, CPU, GPU), and performs the processing defined by the program while using storage resources (for example, memory) and interface devices (for example, communication ports), etc. Therefore, the subject of the processing performed by executing the program may be the processor. Similarly, the subject of the processing performed by executing the program may be a controller, device, system, computer, or node having a processor. The subject of the processing performed by executing the program may be an arithmetic unit, and may include a dedicated circuit for performing a specific process. Here, the dedicated circuit is, for example, an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), a CPLD (Complex Programmable Logic Device), etc. The program may be installed on a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable storage medium. When the program source is a program distribution server, the program distribution server includes a processor and a storage resource for storing the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. Also, in the embodiments, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0014] In an embodiment, in anonymization, even when performing various information processing (generalization, deletion) including pseudonymization that assigns a unified identifier across a plurality of data, the validity of the acquired data can be verified, and a more convenient anonymization system will be described. By using this anonymization system, the validity of the data is guaranteed. And based on this guarantee, the use of tampered data in various services is suppressed. Therefore, the anonymization system can contribute from a social perspective.
[0015] <Anonymization System> FIG. 1 is an explanatory diagram showing a system configuration example of an anonymization system. The anonymization system 100 verifies that the provided information has not been improperly modified when providing anonymized information to prevent leakage of personal information and confidential information. The anonymization system 100 is a system for a data owner who holds text data including confidential information to provide the data to a data user after anonymizing the information. The anonymization system 100 includes a signature generator terminal 101, an anonymized data providing server 102, an anonymized data user terminal 103, and a name-matching data management server 104.
[0016] The signature generator terminal 101, the anonymized data providing server 102, the anonymized data user terminal 103, and the name-matching data management server 104 are connected via the network 105 so as to be able to transmit and receive information. Note that as the type of the network 105, the Internet, a WAN (Wide Area Network), a LAN (Local Area Network), or the like can be considered. Also, the connection method of the network 105 may be either wired or wireless.
[0017] The signature generator terminal 101 is a terminal used by a signature generator (for example, an administrative agency) who is a data holder, and generates a signature for sensitive data so as to be able to verify the validity of anonymized data obtained by performing anonymization processing on the sensitive data. The anonymized data providing server 102 executes the provision of the signature value of the sensitive data, and the generation and provision of anonymized data in accordance with the requests of the anonymized data users from the sensitive data entrusted by the signature generator. The anonymized data user terminal 103 is a terminal used by the anonymized data user, and executes the verification of the validity of the anonymized data. The name-matching data management server 104 manages the identifier at the time of anonymization of the character string to be anonymized.
[0018] <Hardware configuration example of the computer (signature generator terminal 101, anonymized data providing server 102, anonymized data user terminal 103, and name-matching data management server 104)> FIG. 2 is a block diagram showing an example of the hardware configuration of a computer. The computer 200 includes a processor 201, a storage device 202, an input device 203, an output device 204, and a communication interface (communication IF) 205. The processor 201, the storage device 202, the input device 203, the output device 204, and the communication IF 205 are connected by a bus 206. The processor 201 controls the computer 200. The storage device 202 serves as a working area for the processor 201. Also, the storage device 202 is a non-temporary or temporary recording medium that stores various programs and data. Examples of the storage device 202 include a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), and a flash memory. The input device 203 inputs data. Examples of the input device 203 include a keyboard, a mouse, a touch panel, a numeric keypad, a scanner, a microphone, and a sensor. The output device 204 outputs data. Examples of the output device 204 include a display, a printer, and a speaker. The communication IF 205 connects to the network 105 and transmits and receives data.
[0019] <Functional configuration example of the signature generator terminal 101> FIG. 3 is a block diagram showing a functional configuration example of the signature generator terminal 101. The signature generator terminal 101 includes a unique expression extraction unit 301, a similarity calculation unit 302, a first aggregation unit 303, a second aggregation unit 304, a unique expression data list update unit 305, a first signature generation unit 306, a hash value generation unit 307, confidential text data 308, a unique expression data list 309, signed confidential text data 310, and a first signature key 311.
[0020] The specific implementation of the inherent expression extraction unit 301, similarity calculation unit 302, first clustering unit 303, second clustering unit 304, inherent expression data list update unit 305, first signature generation unit 306, and hash value generation unit 307 is realized by, for example, causing the processor 201 to execute a program stored in the storage device 202 shown in FIG. 2. Also, the confidential text data 308, the inherent expression data list 309, the signed confidential text data 310, and the first signature key 311 are stored in the storage device 202.
[0021] The inherent expression extraction unit 301 extracts the inherent expressions included in the confidential text data. The similarity calculation unit 302 calculates the similarity between the inherent expressions extracted by the inherent expression extraction unit 301.
[0022] The first clustering unit 303 presents candidates for clustering the inherent expressions included in the confidential text data to the output device of the signature generator terminal 101, and performs clustering based on the presence or absence of clustering set by the input device of the signature generator terminal 101.
[0023] The second clustering unit 304 presents candidates for clustering the inherent expressions included in the confidential text data and the inherent expressions stored in the clustering data management server to the output device of the signature generator terminal 101, and performs clustering based on the presence or absence of clustering set by the input device of the signature generator terminal 101.
[0024] The inherent expression data list update unit 305 adds a random number, an anonymization (with clustering) index, and an anonymization (without clustering) index to the inherent expression data list extracted by the inherent expression extraction unit 301, and deletes the character string. The first signature generation unit 306 generates a signature value for the confidential text data. The hash value generation unit 307 generates a hash value by means of a one-way function or the like.
[0025] The confidential text data 308 is text data including confidential data such as personal information. In this embodiment, the text data is "My name is Taro Hitachi."
[0026] The specific expression data list 309 stores the start position, end position, classification, index for anonymization (with name grouping), index for anonymization (without name grouping), and random number for all the confidential data included in the confidential text data 308. For example, for the text data "My name is Taro Hitachi.", one specific expression data such as "start position = 5, end position = 8, classification = personal name, index number for name grouping = 1, index number for anonymization without name grouping = 1_1, random number = 01F399" is stored.
[0027] The signed confidential text data 310 is text data with a signature value assigned to the confidential text data 308. The signed confidential text data has signature information added at the end of the confidential text data, and is, for example, text data such as "My name is Taro Hitachi. ===BEGIN SIGN===9138BAC825===END SIGN===". Here, "===BEGIN SIGN===9138BAC825===END SIGN===" is the signature information.
[0028] The first signature key 311 is key information for encrypting the random number for the confidential text data 308. The first signature key 311 is, for example, the private key in the public key cryptosystem.
[0029] <Functional configuration example of the anonymization data providing server 102> Figure 4 is a block diagram showing a functional configuration example of the anonymization data providing server 102. The anonymization data providing server 102 includes a web server function 401, an anonymization processing unit 402, a second signature generation unit 403, a hash value generation unit 307, a specific expression data list 309, signed confidential text data 310, signed anonymized text data 404, a verification specific expression data list 405, and a second signature key 406.
[0030] The Web server function 401, the anonymization processing unit 402, and the second signature generation unit 403 are specifically realized, for example, by causing the processor 201 to execute a program stored in the storage device 202 shown in FIG. 2. Also, the signed anonymized text data 404, the verification-specific expression data list 405, and the second signature key 406 are stored in the storage device 202.
[0031] The Web server function 401 holds a Web page including the identification information of the confidential text data so as to be accessible from the anonymized data user terminal 103.
[0032] The anonymization processing unit 402 performs a verifiable anonymization process on the confidential text data. The second signature generation unit 403 generates a signature for the text data subjected to the anonymization process by a known digital signature method.
[0033] The signed anonymized text data 404 is the result of further attaching the signature generated by the second signature generation unit 403 to the result of anonymizing the signed confidential text data by the anonymization processing unit 402. For example, when deleting the name "My name is Taro Hitachi. ===BEGIN SIGN===9138BAC825===END SIGN===", it becomes "My name is [ ] is. ===BEGIN SIGN===9138BAC825,A74BE96224===END SIGN===".
[0034] The verification-specific expression data list 405 stores the specific expression data necessary for signature verification among the specific expression data list 309. The second signature key 406 is key information for encrypting a random number for the data obtained by anonymizing the signed confidential text data. The second signature key 406 is, for example, the private key of the public key cryptosystem.
[0035] <Functional configuration example of the anonymized data user terminal 103> FIG. 5 is a block diagram showing a functional configuration example of the anonymized data user terminal 103. The anonymized data user terminal 103 includes a web browser function 501, a first signature verification unit 502, a second signature verification unit 503, a hash value generation unit 307, signed anonymized text data 404, a verification-specific expression data list 405, a first verification key 504, and a second verification key 505.
[0036] Specifically, the web browser function 501, the first signature verification unit 502, and the second signature verification unit 503 are realized by causing the processor 201 to execute a program stored in the storage device 202 shown in FIG. 2, for example. Also, the first verification key 504 and the second verification key 505 are stored in the storage device 202.
[0037] The web browser function 501 receives the web page published by the anonymized data providing server 102 and displays it on the anonymized data user terminal 103. The first signature verification unit 502 verifies the legitimacy of anonymization with the signed anonymized text data 404, the verification-specific expression data list 405, and the first verification key 504 as inputs. The second signature verification unit 503 verifies that the anonymized text data has not been modified with the signed anonymized text data 404 and the second verification key 505 as inputs.
[0038] <Functional Configuration Example of the Name Aggregation Data Management Server 104> FIG. 6 is a block diagram showing a functional configuration example of the name aggregation data management server 104. The name aggregation data management server 104 includes a similarity calculation unit 302, a name aggregation data generation unit 601, an index generation unit 602, a name aggregation data table group 603, and a text counter 604.
[0039] Specifically, the name aggregation data generation unit 601 and the index generation unit 602 are realized by causing the processor 201 to execute a program stored in the storage device 202 shown in FIG. 2, for example. Also, the name aggregation data table group 603 and the text counter 604 are stored in the storage device 202.
[0040] <Name-matching data table 603> FIG. 7 is an explanatory diagram showing an example of the name-matching data table group 603. The name-matching data table 603 is a group of tables for each classification of character strings (such as personal names and place names). For example, it has a name-matching table 710 for personal names and a name-matching table 720 for place names. Each table has a character string and an index as attribute fields.
[0041] The character strings (711, 721) are partial character strings to be processed in the confidential text data 306 of the signature generator terminal 101. The indexes (712, 722) store identifiers assigned when the values included in the character strings (711, 721) are name-matched.
[0042] <Sequence of the anonymization system 100> Next, an example of the processing of the anonymization system will be described with reference to FIG. 8. FIGS. 8A and 8B are sequence diagrams of the anonymization system 100.
[0043] The signature generator terminal 101 executes a named entity extraction process by the named entity extraction unit 301 (step S801). In the named entity extraction process, a named entity data list for the text to be signed is obtained. The named entity data list includes zero or more named entity data and is sorted in ascending order of the start position of the named entity. The named entity data includes at least the character string of the named entity, the start position, the end position, and the classification of the named entity. The named entity extraction process may use a known named entity extraction tool, or may be extracted by the signature generator, or may use a combination of a named entity extraction tool and manual work.
[0044] Next, the signature generator terminal 101 calculates the similarity between the strings of the specific expressions by the similarity calculation unit 302 (step S802). Specifically, first, the specific expression data included in the specific expression extraction data list acquired in S801 is divided according to the classification of the specific expressions. Next, for each classification, the similarity between the strings of the specific expressions belonging to the classification is calculated by a known string similarity calculation method (for example, the Levenshtein distance). The string similarity calculation is performed between the strings of all the specific expressions belonging to the same classification. For example, when there are three types of strings, "Hanako Takahashi", "Taro Suzuki", and "Suzuki" in the "person name" classification, the similarity calculation is performed for each of the three combinations: "Hanako Takahashi" and "Taro Suzuki", "Hanako Takahashi" and "Suzuki", and "Taro Suzuki" and "Suzuki". When there are 0 or 1 specific expressions during the classification, the similarity calculation is omitted.
[0045] Next, the signature generator terminal 101 displays a name alignment specification screen in the document on the signature generator terminal 101 by the first name alignment unit 303, and generates a specific expression string list using the setting of whether to align with respect to the alignment candidates by operating the input device (step S803).
[0046] Here, with reference to FIG. 9, an example of the name alignment specification screen 900 in the document will be described. FIG. 9 is an explanatory diagram showing an example of the name alignment specification screen in the document. The document name alignment specification screen 900 has a text data display area 901, a name alignment specification area 902, and a "next" button 903.
[0047] The text data display area 901 displays the confidential text data to be signed with the alignment candidates emphasized.
[0048] The display content of the clustering designated area 902 changes according to the presence or absence of a string pair (hereinafter referred to as a clustering candidate) whose similarity calculated in S802 is higher than a preset threshold. When there is one or more clustering candidates, the clustering designated area 902 has a display area for the strings of the clustering candidates and radio buttons for "perform" and "do not perform" clustering for each clustering candidate. The radio buttons can select either "perform" or "do not perform" for each clustering candidate. When there is no clustering candidate, the text "no clustering candidate" is displayed in the clustering designated area.
[0049] When the button 903 is pressed, a first clustered proper name string list in which proper names are grouped based on the instruction of whether to perform clustering for each input clustering candidate is generated. For example, among the three proper names that appear in the sensitive text data in the order of "Hanako Takahashi", "Jiro Suzuki", and "Suzuki", when clustering "Jiro Suzuki" and "Suzuki", two proper name string lists {Hanako Takahashi} and {Jiro Suzuki, Suzuki} are generated.
[0050] Returning to FIG. 8, the description continues. Next, the proper name string list generated in S803 is transmitted to the clustering data management server 104 (step S804).
[0051] Next, the name-matching data management server 104 receives the list of unique expression strings transmitted from the signature generator terminal 101, and calculates the similarity between each string in the list of unique expression strings and each string in the name-matching data*** (i.e., the name-matching data managed in the name-matching data table group 603) in the same manner as in S802 using a known string similarity calculation method (step S805). For example, if the name-matching data has a string group of {"Sato Taro", "Sato"}, {"Suzuki Jiro"}, {"Takahashi"}, then the four strings of "Sato Taro", "Sato", "Suzuki Jiro", "Takahashi" and the three strings of "Takahashi Hanako", "Suzuki Jiro", "Suzuki" received from the signature generator terminal 101 are compared respectively. That is, the similarity is calculated in 12 patterns: "Sato Taro" and "Takahashi Hanako", "Sato Taro" and "Suzuki Jiro", "Sato Taro" and "Suzuki", "Taro" and "Takahashi Hanako", "Taro" and "Suzuki Jiro", "Taro" and "Suzuki", "Suzuki Jiro" and "Takahashi Hanako", "Suzuki Jiro" and "Suzuki Jiro", "Suzuki Jiro" and "Suzuki", "Takahashi" and "Takahashi Hanako", "Takahashi" and "Suzuki Jiro", "Takahashi" and "Suzuki".
[0052] Next, the name-matching data management server 104 generates a name-matching candidate list by the name-matching data generation unit 601 and transmits the name-matching candidate list to the signature generator terminal 101 (step S806). Here, the name-matching candidate list includes zero or more name-matching candidates whose similarity calculated in S805 is higher than a preset threshold, and the name-matching candidates are composed of pairs of the unique expression group received from the signature generator terminal and the unique expression group of the name-matching data. When the string is name-matched in step S803 like {"Suzuki Jiro", "Suzuki"}, or when there is a string with a similarity higher than the threshold among "Suzuki Jiro" or "Suzuki", the name-matching data management server 104 regards it as a name-matching candidate. In this embodiment, it is assumed that two name-matching candidates are the pair of {"Takahashi"} in the name-matching data and {"Takahashi Hanako"} in the received data, and the pair of {"Suzuki Jiro"} in the name-matching data and {"Suzuki Jiro", "Suzuki"} in the received data.
[0053] Next, the signature generator terminal 101 receives collation candidates from the collation data management server 104, and the second collation unit 304 displays a document-to-document collation designation screen on the signature generator terminal 101, and generates a document-to-document collation unique expression string list using the setting of whether to collate for the collation candidates by operating the input device (step S807).
[0054] Here, with reference to FIG. 10, an example of the document-to-document collation designation screen 1000 will be described. FIG. 10 is an explanatory diagram showing an example of the document-to-document collation designation screen. The document-to-document collation designation screen 1000 has a text data display area 1001, a collation designation area 1002, and an "Execute" button 1003.
[0055] The text data display area 1001 displays the confidential text data to be signed, after emphasizing the collation candidates.
[0056] The display content of the collation designation area 1002 changes depending on the presence or absence of the collation candidates received in S807. When there is one or more collation candidates, the collation designation area 1002 has a display area for the strings of the collation candidates and radio buttons for "Do" and "Do not" collate for each collation candidate. The radio buttons can select either "Do" or "Do not" for each collation candidate. When there are no collation candidates, the text "No collation candidates" is displayed in the collation designation area.
[0057] When the button 1003 is pressed, a second collation result including whether to collate for each input collation candidate is generated. For example, for the two collation candidates received in S807, collation results of {In document = {Hanako Takahashi}, Outside document = {Takahashi}, Collate = Do not}, {In document = {Jiro Suzuki, Suzuki}, Outside document = {Jiro Suzuki}, Collate = Do} are generated. In this example, the strings in the confidential text data displayed in the text data display area 1001 are displayed as in-document notations, and the strings related to the collation data table managed by the collation data management server 104 are displayed as out-of-document notations.
[0058] Returning to FIG. 8, the description will be continued. The signature generator terminal 101 transmits the clustering result generated in S807 to the clustering data management server 104 (step S808).
[0059] The clustering data management server 104 receives the clustering result from the signature generator terminal 101 and updates the clustering data based on the received clustering result (step S809). Specifically, the clustering data management server 104 sequentially refers to the clustering result. When "clustering = perform", if the string group described in the document contains a string that does not exist in the string group described outside the document, the clustering data management server 104 adds the string to the corresponding string group of the clustering data. In this example, the string "Suzuki" is newly added to the record with index = 2. When "clustering = none", the clustering data management server 104 adds the pair of the in-document string group of the clustering result and the index to the clustering data. For example, when adding the string "Hanako Takahashi" newly, a record with string group = {Hanako Takahashi} and index = 4 is added.
[0060] Next, the clustering data management server 104 generates indexed specific expression string data with the anonymization (with clustering) index (first index) and the anonymization (without clustering) index (second index) added to the specific expression string data received in S805 (step S810). Specifically, for the specific expression string data, the anonymization (with clustering) index is the index of the record in which the string group is included in the clustering data. The anonymization (without clustering) index is a value obtained by concatenating with an underscore the value obtained by incrementing the value stored in the text counter 604 by 1 and the specific expression number sequentially assigned from 1 to the specific expression string data. For example, when 1 is stored in the text counter 604, {string group = {Hanako Takahashi}, anonymization (with clustering) index = 4, anonymization (without clustering) index = 2_1}, {string group = {Jiro Suzuki, Suzuki}, anonymization (with clustering) index = 2, anonymization (without clustering) index = 2_2} are generated, and the value of the text counter 604 is updated to 2.
[0061] The name-matching data management server 104 transmits the indexed unique expression string data generated in S810 to the signature generator terminal 101 (step S811).
[0062] The signature generator terminal 101 updates the unique expression data list using the indexed unique expression string data received from the name-matching data management server 104 (step S812). Specifically, the unique expression data list update unit 305 of the signature generator terminal 101 sequentially refers to the unique expression data list, and for each unique expression data, adds a pseudonymized (with name-matching) index and a pseudonymized (without name-matching) index using the indexed unique expression string data. Also, the unique expression data list update unit 305 generates a random number for each unique expression data and adds the random number. Furthermore, the unique expression data list update unit deletes the unique expression string from the unique expression data.
[0063] Next, the signature generator terminal 101 inputs the confidential text data 308, the unique expression data 309, and the signature key 311 to the first signature generation unit 306, generates a signature value for the confidential text data, and performs a signature generation process for attaching the signature value to the confidential text data (step S813). The details of the signature generation process will be described later with reference to FIGS. 12 to 15.
[0064] Next, the signature generator terminal 101 transmits the signed confidential text data generated in S813 and the unique expression data list to the anonymization data providing server 102 (step S814).
[0065] The anonymization data providing server 102 acquires the signed confidential text data and the unique expression data list transmitted from the signature generator terminal 101 and stores them in the storage device 202 (step S815).
[0066] Next, the anonymized data providing server 102 generates a web page accessible via the network 104 by means of the web server function 401, and notifies the web browser function 501 of the anonymized data user terminal 103 of the URL of the web page (step S816).
[0067] Next, the anonymized data user terminal 103 acquires, by means of the web browser function 501, a web page including the identification information of the signed confidential text data (step S817).
[0068] Next, the anonymized data user terminal 103 sets an anonymization request by operating the input device 203 of the data user, and transmits an anonymized data acquisition request including the anonymization request to the web server function 401 of the anonymized data providing server 102 (step S818). In this embodiment, an example is given in which the anonymization request included in the anonymization data acquisition request is "pseudonymization of personal names (with name grouping)".
[0069] Here, with reference to FIG. 11, an example of the anonymization request setting screen will be described. FIG. 11 is an explanatory diagram showing an example of the anonymization request setting screen. The anonymization request setting screen 1100 is displayed on the output device 204 of the anonymized data user terminal 103, and has a processing rule designation area 1101 and an execution button 1102.
[0070] The processing rule designation area 1101 has radio buttons for options of "no processing", "pseudonymization (with name grouping)", "pseudonymization (without name grouping)", "generalization", or "deletion" for each classification of the specific expression, and any one processing rule can be selected for each attribute. For example, assume that the processing rule for personal names is set to "pseudonymization (with name grouping)" and other classifications are set to "no processing".
[0071] When the execution button 1102 is pressed, an acquisition request for anonymized data including the anonymization request for each input classification is transmitted to the anonymized data providing server 102 (step S818).
[0072] Returning to FIG. 8, the description continues. Next, the anonymized data providing server 102 acquires the identification information of the signed confidential text data and the anonymization request notified from the anonymized data user terminal 103 by the web server function 401, and passes them to the anonymization processing unit 402. Then, the anonymized data providing server 102 executes anonymization processing with the signed confidential text data and the anonymization request as inputs by the anonymization processing unit 402 to generate signed anonymized text data and a verification-specific expression data list (step S819). The details of the anonymization processing (step S819) will be described later with reference to FIG. 16.
[0073] Next, the anonymized data providing server 102 generates a signature value for the anonymized text data using a known digital signature method by the second signature generation unit 403, generates signed anonymized text data with the generated signature value added to the signature information of the signed anonymized text data, and passes the signed anonymized text data and the verification-specific expression data list to the web server function 401 (step S820).
[0074] The web server function 401 registers the signed anonymized text data and the verification-specific expression data list on a download web page for data users, and notifies the URL to the web browser function 501 of the anonymized data user terminal 103.
[0075] Next, the anonymized data user terminal 103 accesses the download web page for data users with the notified URL as an input by the web browser function 501, downloads and acquires the signed anonymized text data and the verification-specific expression data list, and stores them in the storage device 202 of the anonymized data user terminal 103 (step S821).
[0076] Next, the anonymized data user terminal 103 executes a first signature verification process to verify the validity of the anonymization process (step S819) by the anonymized data providing server 102, using the signed anonymized text data, the verification-specific expression data list, and the verification key as inputs to the first signature verification unit 502. If the validity is verified, "Anonymization validity verification succeeded" is displayed on the output device 204, and if the validity is not verified, "Anonymization validity verification failed" is displayed (step S822). Details of the first signature verification process (step S822) will be described later with reference to FIGS. 17 to 18.
[0077] Finally, the anonymized data user terminal 103 executes a process to verify the authenticity of the anonymized data transmitted by the anonymized data providing server 102 using a known digital signature verification method, using the signed anonymized text data and the second verification key as inputs to the second signature verification unit 503. If the signature verification is successful, "Authenticity verification succeeded" is displayed on the output device 204, and if the signature verification fails, "Authenticity verification failed" is displayed (step S823).
[0078] By the process shown in FIG. 8, the anonymized data user can verify whether the obtained signed anonymized text data has been properly anonymized by the proper anonymization process (step S817) by the anonymized data providing server 102. This can prevent the provision of services that use incorrect analysis results obtained from incorrect anonymized data.
[0079] <Signature generation process (step S813)> FIG. 12 is a flowchart showing a detailed processing procedure example of the signature generation process (step S813) shown in FIG. 8.
[0080] The first signature generation unit 306 divides the confidential text data using the specific expression data list, and generates an element data list including one or more pieces of element data generated for each divided specific expression or non-specific expression character string (element) (step S1201). Details of the element data list generation process will be described later with reference to FIG. 13.
[0081] The first signature generation unit 306 generates an element hash value list by referring to the confidential text data and the element data list (step S1202). Details of the element hash value list generation process will be described later with reference to FIG. 14.
[0082] The first signature generation unit 306 generates a hash value (data hash value) for the entire confidential text data from the element hash value list (step S1203). Specifically, the value obtained by concatenating all the element hash values is input to the hash value generation unit 307 to generate the data hash value.
[0083] Finally, the first signature generation unit 306 generates a signature using the data hash value and the first signature key 311, and stores the signed confidential text data with the signature attached to the confidential text data.
[0084] <Element data list generation process (step S1201)> FIG. 13 is a flowchart showing a detailed processing procedure example of the element data list generation process (step S1201).
[0085] The first signature generation unit 306 generates an empty element data list (step S1301).
[0086] Next, the first signature generation unit 306 determines whether the number N of the specific expression data in the specific expression data list obtained in S801 is 0 (step S1302). If N = 0 (step S1302: Yes), the first signature generation unit 306 adds element data with the start position being 0 and the end position being txt_len - 1 to the element data list (step S1303), and ends the element data list generation process (step S1201). Here, txt_len represents the number of characters of the confidential text data.
[0087] On the other hand, when N≠0 (step S1302: No), the first signature generation unit 306 determines whether the start position of the 0th unique expression is 0 (step S1304). When the start position of the 0th unique expression = 0 (step S1304: Yes), the process proceeds to S1306. On the other hand, when the start position of the 0th unique expression ≠ 0 (step S1304: No), element data with the start position set to 0 and the end position set to the start position of the 0th unique expression - 1 is added to the element data list (step S1305).
[0088] The first signature generation unit 306 initializes a variable i indicating which unique expression it is among the unique expression data to 0 (step S1306).
[0089] The first signature generation unit 306 determines whether i < N (step S1307). When i < N is not satisfied (step S1307: No), the process proceeds to S1312. On the other hand, when i < N (step S1307: Yes), the first signature generation unit 306 adds the i-th unique expression data to the element data list (step S1308).
[0090] Next, the first signature generation unit 306 determines whether i ≠ N - 1 and the start position of the (i + 1)-th unique expression ≠ the end position of the i-th unique expression + 1 (step S1309). When i ≠ N - 1 and the start position of the (i + 1)-th unique expression ≠ the end position of the i-th unique expression + 1 is not satisfied (step S1309: No), the process proceeds to S1311. On the other hand, when i ≠ N - 1 and the start position of the (i + 1)-th unique expression ≠ the end position of the i-th unique expression + 1 (step S1309: Yes), element data with the start position set to the end position of the i-th unique expression + 1 and the end position set to the start position of the (i + 1)-th unique expression - 1 is stored in the element data list (step S1310).
[0091] The first signature generation unit 306 increments the variable i (step S1311), and the process proceeds to S1307.
[0092] The first signature generation unit 306 determines that the end position of the (N - 1)-th unique expression = txt_len - 1 (step S1312). If the end position of the (N - 1)-th unique expression = txt_len - 1 (step S1312: Yes), the element data list generation process (step S1201) ends. On the other hand, if the end position of the (N - 1)-th unique expression ≠ txt_len - 1 (step S1312: No), the element data with the start position being the end position of the (N - 1)-th unique expression + 1 and the end position being txt_len - 1 is stored in the element data list (step S1313), and the element data list generation process (step S1201) ends.
[0093] <Element hash value list generation process (step S1202)> FIG. 14 is a flowchart showing a detailed processing procedure example of the element hash value list generation process (step S1202) shown in FIG. 12.
[0094] The first signature generation unit 306 generates an empty element hash list (step S1401).
[0095] The first signature generation unit 306 initializes a variable j indicating which expression it is in the element data list to 0 (step S1402).
[0096] The first signature generation unit 306 determines whether j < M (step S1403). M is the number of elements in the element data list (the total number of unique expression strings and non-unique expression strings in the confidential text data). If j ≥ M (step S1403: No), the element hash value list generation process (step S1202) ends. On the other hand, if j < M (step S1403: Yes), the first signature generation unit 306 extracts an element string from the confidential text data using the start position and end position of the j-th element data (step S1404).
[0097] The first signature generation unit 306 determines whether a random number is included in the j-th element data of the element data list (step S1405). If a random number is included in the j-th element data (step S1405: Yes), using the element character string and the j-th element data, an element hash value generation process described later with reference to FIG. 15 is performed, and the obtained element hash value is added to the element hash value list (step S1406). On the other hand, if a random number is not included in the j-th element data (step S1405: No), the element character string is input to the hash value generation unit, and the obtained hash value is added to the element hash value list (step S1407).
[0098] The first signature generation unit 306 increments the variable j (step S1408) and proceeds to S1203.
[0099] <Element hash value generation process (step S1405)> FIG. 15 is a flowchart showing a detailed processing procedure example of the element hash value generation process (step S1406) shown in FIG. 14.
[0100] The first signature generation unit 306 reads the element character string Data_O and the element data of Data_O (classification, anonymization (with name grouping) index number Index1, anonymization (without name grouping) index number Index2, random number R_O) extracted in S1404, and performs anonymization (with name grouping) processing (step S1501). Specifically, anonymization (with name grouping) data Data_P1 and a random number R_P1 corresponding to Data_P are generated. Data_P1 is a value obtained by concatenating the classification of Data_O and the anonymization (with name grouping) index number Index1. Also, R_P1 is the hash value of the value obtained by concatenating Data_O and R_O. Data_P1 = classification + Index1 R_P1 = Hash(Data_O + R_O)
[0101] Next, the first signature generation unit 306 performs anonymization (no name grouping) processing on the anonymized (no name grouping) data Data_P2 by using the concatenated value of the classification of Data_O and the anonymized (no name grouping) index number Index2, and setting R_P2 as the hash value of the concatenated value of Data_P1 and R_P1 (step S1502). Data_P2 = classification + Index2 R_P2 = Hash(Data_P1 + R_P1)
[0102] Next, the first signature generation unit 306 performs generalization processing on the generalized data Data_G by using the classification of Data_O and setting the random number R_G as the hash value of the concatenated value of Data_P2 and R_P2 (step S1503). Data_G = classification R_G = Hash(Data_P2 + R_P2)
[0103] Finally, the first signature generation unit 306 performs deletion processing by setting the data after deletion Data_R to blank and setting the random number R_R corresponding to the data after deletion Data_R as the hash value of the concatenated value of the generalized data Data_G and the random number R_G (step S1504). The random number R_R is used as the element hash value, and the element hash value generation process ends. Data_R = blank R_R = Hash(Data_G + R_G) = element hash value
[0104] <Anonymization process (step S2307)> FIG. 16 is a flowchart showing a detailed processing procedure example of the anonymization process (step S819) shown in FIG. 8. The anonymization process is performed on the confidential text data part other than the signature in the signed confidential text data, and signed anonymized text data is generated by attaching the signature of the signed confidential text data to the anonymized text data obtained as a result of the anonymization process.
[0105] The anonymization processing unit 402 generates an element data list from the confidential text data and the list of specific expression data by the element data list generation process (step S1201) shown in FIG. 13.
[0106] The anonymization processing unit 402 generates an empty anonymized element string list and empty verification element data (step S1601).
[0107] The anonymization processing unit 402 initializes to 0 a variable j indicating which expression among the element data it is, and a len indicating the number of characters of the anonymized text data (step S1602).
[0108] The anonymization processing unit 402 determines whether j < M (step S1603). M is the number of elements in the confidential text data (the total number of strings of specific expressions and the number of strings of non-specific expressions).
[0109] If j < M is not satisfied (step S1603: No), the anonymization processing unit 402 generates anonymized text data by concatenating all the element strings stored in the anonymized element string list (step S1609), and ends the anonymization processing (step S819).
[0110] On the other hand, if j < M (step S1603: Yes), the anonymization processing unit 402 extracts an element string from the confidential text data using the start position and end position of the j-th element data (step S1604).
[0111] The anonymization processing unit 402 determines whether processing for the classification of the j-th element data is required according to step S818 (step S1605).
[0112] When no processing is required (step S1605: No), the anonymization processing unit 402 updates the anonymization element string list, the verification-specific expression data list, and len (step S1606). Specifically, the anonymization processing unit 402 first adds the element string to the anonymization element string list. Also, the anonymization processing unit 402 sets the start position to "len", the end position to "len + element string - 1", and adds the specific expression data obtained by replicating the pseudonymization (with name clustering) index, pseudonymization (without name clustering) index, classification, and random number from the j-th element data to the verification-specific expression data list respectively. Further, the anonymization processing unit 402 sets len to the number of characters up to the start of the string. After S1606, it proceeds to S1608.
[0113] On the other hand, when processing is required (step S1605: Yes), the anonymization processing unit 402 performs anonymization processing on the element string (step S1607). Specifically, first, the anonymization processing unit 402 determines which of the "pseudonymization (with name clustering)", "pseudonymization (without name clustering)", "generalization", and "deletion" processes is required for the element string.
[0114] When "pseudonymization (with name clustering)" is required, the anonymization processing unit 402 performs the pseudonymization (with name clustering) process of S1501 in the element hash value generation process shown in FIG. 15, and adds the anonymization element string obtained by concatenating the anonymization identification string indicating that it is an anonymized string to the head and tail of the pseudonymized data Data_P1 to the anonymization element string list. Also, the anonymization processing unit 402 sets the start position to "len", the end position to "len + anonymization element string - 1", and the random number to "random number R_P1 corresponding to Data_P1", and adds the specific expression data obtained by replicating the pseudonymization (without name clustering) index and classification from the j-th element data to the verification-specific expression data list. Further, the anonymization processing unit 402 sets len to the number of characters up to the start of the string.
[0115] When "pseudonymization (without name grouping)" is required, the anonymization processing unit 402 performs the processes of S1501 to S1502 in the element hash value generation process shown in FIG. 15, and concatenates an anonymization identification string indicating that it is an anonymized string at the head and tail of the pseudonymized data Data_P2. The anonymized element string thus obtained is added to the anonymized element string list. Also, the anonymization processing unit 402 adds the specific expression data with the start position being "len", the end position being "len + anonymized element string - 1", and the random number being "random number R_P2 corresponding to Data_P2" to the verification specific expression data list with the classification being "classification of the j-th element data". Further, the anonymization processing unit 402 sets len to the number of characters up to the start of the string.
[0116] When "generalization" is required, the anonymization processing unit 402 performs the processes of S1501 to S1503 in the element hash value generation process shown in FIG. 15, and concatenates an anonymization identification string indicating that it is an anonymized string at the head and tail of the generalized data Data_G obtained as a result of the process. The anonymized element string thus obtained is added to the anonymized element string list. Also, the specific expression data with the start position being "len", the end position being "len + anonymized element string - 1", and the random number being "random number R_G corresponding to Data_G" is added to the verification specific expression data list. Further, the anonymization processing unit 402 sets len to the number of characters up to the start of the string.
[0117] When "deletion" is required, the anonymization processing unit 402 performs all of S1501 to S1504 in the element hash value generation process shown in FIG. 15. An anonymized element string obtained by concatenating an anonymization identification string indicating that it is an anonymized string at the head and tail of the deletion data Data_R (which is blank) obtained as a result of S1504 is added to the anonymized element string list. Also, the anonymization processing unit 402 adds the specific expression data with the start position being "len", the end position being "len + number of characters of the anonymized element string - 1", and the random number being "random number R_R corresponding to Data_R" to the verification specific expression data list. Further, the anonymization processing unit 402 sets len to the number of characters up to the start of the string.
[0118] The anonymization processing unit 402 increments the variable j (step S1608) and returns to step S1603. Through the processing related to S1607, a verification-specific expression data list including random numbers for anonymized data generated when processing information in response to an anonymization request is generated.
[0119] <First signature verification process (step S822)> FIG. 17 is a flowchart showing a detailed processing procedure example of the signature verification process (step S822) shown in FIG. 8.
[0120] The anonymized data user terminal 103 generates an element data list of the anonymized text data using the anonymized text data portion of the signed anonymized text data 404 acquired in step S821 and the verification-specific expression data list 405 in the first signature verification unit 502 (step S1701). The generation process of the element data list is the same as the element data list generation process (step S1201) shown in FIG. 13, except that the sensitive text data is changed to anonymized text data and the specific expression data list is changed to the verification-specific expression data list.
[0121] The first signature verification unit 502 generates an element hash value list of the anonymized text data using the anonymized text data and the element data list generated in S1701 (step S1702). Specifically, the first signature verification unit 502 sequentially refers to the M element data included in the verification element data list for an empty element hash value list, and performs a process of adding the element hash value obtained by the following process to the element hash value list. First, the first signature verification unit 502 determines whether the j-th element data contains "random number". If it does not contain "random number" (in the case of element data for non-specific expressions), the first signature verification unit 502 extracts an element character string from the anonymized text data using the start position and end position of the j-th element data, and performs the same element hash value generation process as S1407 in FIG. 14 on the extracted element character string. If it contains "random number" (in the case of element data for specific expressions), the element hash value generation process described later in FIG. 18 is executed.
[0122] The first signature verification unit 502 generates a data hash value H' of the anonymized text data from the element hash value list of the anonymized text data by the same process as S1203 described above (step S1703).
[0123] The first signature verification unit 502 decrypts the signature value for the confidential text data with the first verification key 504, and obtains the data hash value H for the confidential text data (step S1704).
[0124] The first signature verification unit 502 determines whether the data hash value H of the decrypted confidential text data matches the data hash value H' of the anonymized text data (step S1705). If they match (step S1705: Yes), the first signature verification unit 502 outputs, for example, "Successful verification of anonymization validity" to be displayable on the output device 204 (step S1706). On the other hand, if they do not match (step S1705: No), the first signature verification unit 502 outputs, for example, "Failed verification of anonymization validity" to be displayable on the output device 204 (step S1707), and the first signature verification process (step S822) ends.
[0125] <Element hash value generation S1702 for elements of the unique expression of anonymized data> FIG. 18 is a flowchart showing a detailed processing procedure example of the element hash value generation process for element data of a unique expression in the element hash value list generation process (step S1702) shown in FIG. 17.
[0126] The first signature verification unit 502 extracts an element character string from the anonymized text data using the start position and end position of the j-th element data, deletes the anonymization identification character string from the extracted element character string, and determines whether the element character string Data after deletion is blank (step S1801). If Data is blank (step S1801: Yes), the first signature verification unit 502 ends the element hash value generation process with the random number R of the j-th element data as the element hash value.
[0127] On the other hand, when the element string is not blank (step S1801: No), the first signature verification unit 502 determines whether there is a pseudonymization (with name matching) index Index1 in the j-th element data (step S1802).
[0128] When Index1 exists (step S1802: Yes), the first signature verification unit 502 substitutes Data for Data_O and R for R_O respectively, performs the same processing as S1501, and generates pseudonymized data Data_P1 and a random number R_P1 corresponding to the pseudonymized data Data_P1 (step S1803).
[0129] On the other hand, when the index data does not exist (step S1802: No), the first signature verification unit 502 determines whether there is a pseudonymization (without name matching) index Index2 in the j-th element data (step S1804). When Index2 exists (step S1804: Yes), the first signature verification unit 502 substitutes Data for Data_P1 and R for R_P1 respectively (step S1805) and proceeds to S1807. On the other hand, when Index2 does not exist (step S1802: Yes), the first signature verification unit 502 substitutes Data for Data_P2 and R for R_P2 respectively (step S1806) and proceeds to S1808.
[0130] The first signature verification unit 502 performs the same pseudonymization (without name matching) processing as S1502 to generate Data_P2 and R_P2 (step S1807).
[0131] The first signature verification unit 502 performs the same generalization processing as S1503 to generate Data_G and R_G (step S1808).
[0132] The first signature verification unit 502 performs the same deletion processing as S1504 to generate R_R, uses R_R as the element hash value (step S1809), and ends the element hash value generation process.
[0133] By calculating the element hash value in this way for the unique expression elements of the anonymized text data, the anonymized data user terminal 103 can generate the same element hash value list as the element hash value list generated by the signature generator terminal 101 in the first signature generation process S813 regardless of what anonymization requirements are imposed.
[0134] As described above, when the data hash value H generated from the confidential text data 308 by the first signature generation process (step S813) described in FIGS. 12 to 15 matches the anonymized data hash value H' generated by the first signature verification process (step S822) shown in FIGS. 17 to 18 for the signed anonymized text data 404 subjected to the anonymization process (step S819) described in FIG. 16, it can be verified that no improper anonymization process has been performed, that is, the legitimacy of the anonymization can be verified.
[0135] Also, in the second signature generation process (step S820) shown in FIG. 8, a digital signature using a known digital signature is added to the signed anonymized text data, and by verifying the added digital signature in the second signature verification process (step S823), it becomes possible to verify that the signed anonymized data generated by the anonymized data providing server 102 has not been modified.
[0136] Although the embodiments have been described above, the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, for example, for a part of the configuration of the embodiment, addition, deletion, or replacement with other configurations may be made.
[0137] In the above-described embodiment, the anonymization system 100 is configured by an anonymization data providing device (signature generator terminal 101, anonymization data providing server 102), an anonymization data user device (anonymization data user terminal 103), and a name-matching data management server 104. And an example in which they are different computers capable of communicating with each other has been described.
[0138] Here, for example, the signature generator terminal 101 may execute the processing executed by the name-matching data management server 104. That is, in the anonymization system, the signature generator terminal 101 and the name-matching data management server 104 may be configured by the same computer. Further, the anonymization data providing server 102 may execute the processing executed by the signature generator terminal 101 and the name-matching data management server 104. That is, in the anonymization system 100, the signature generator terminal 101, the anonymization data providing server 102, and the name-matching data management server 104 may be configured by the same computer. Also, the name-matching data management server 104 may be configured as a computer different from the anonymization data providing device and the anonymization data user terminal 103.
Explanation of Signs
[0139] 100 Anonymization system 101 Signature generator terminal 102 Anonymization data providing server 103 Anonymization data user terminal 104 Name-matching data management server 301 Specific expression extraction unit 302 Similarity calculation unit 303 First name-matching unit 304 Second name-matching unit 305 Specific expression data list update unit 306 First signature generation unit 307 Hash value generation unit 308 Confidential text data 309 Specific expression data list 310 Confidential text data with signature 311 First signature key 401 Web server function 402 Anonymization processing unit 403 Second signature generation unit 404 Signed anonymized text data 405 Verification-specific expression data list 406 Second signature key 501 Web browser function 502 First signature verification unit 503 Second signature verification unit 504 First verification key 505 Second verification key 601 Grouping data generation unit 602 Index generation unit 603 Grouping data table group 604 Text counter
Claims
1. An anonymization system comprising an anonymized data providing device and an anonymized data user device capable of communicating with the anonymized data providing device, wherein the anonymized data providing device, when generating a value for data verification, for each unique expression in the text data, a pseudonymized (with name matching) data in which the original data is replaced with a value identifiable across a plurality of text data, and a first process for generating a random number for pseudonymized (with name matching) data based on the original data and a random number used for the original data; pseudonymized (without name matching) data in which the original data is replaced with a value identifiable only within the text data, and a second process for generating a random number for pseudonymized (without name matching) data based on the pseudonymized (with name matching) data and the random number for pseudonymized (with name matching) data; generalized data in which the original data is replaced with a generalized value, and a third process for generating a random number for generalized data based on the pseudonymized (without name matching) data and the random number for pseudonymized (without name matching) data; deleted data in which the generalized data is replaced with a blank, and a fourth process for generating a random number for deleted data based on the generalized data and the random number for generalized data; and a fifth process for generating a hash value of the original data for each non-unique expression in the text data, when anonymizing data in response to an anonymization request regarding anonymization of data from the anonymized data user device, for each element in the data, performing the process required by the anonymization request among the first process, the second process, the third process, and the fourth process, generating anonymized data which is the anonymized data, and a random number for anonymized data which is the random number based on the process, wherein the anonymized data user device, acquires the anonymized data and the random number for anonymized data via communication, when verifying data, for each unique expression in the anonymized text data, performing the process not performed by the anonymized data providing device among the first process, the second process, the third process, and the fourth process, generating a random number for deleted data, and for each non-unique expression in the anonymized data, generating a hash value of the original data, verifying the validity of the acquired anonymized data based on the random number for deleted data and the hash value generated by the anonymized data providing device and the random number for deleted data and the hash value generated, An anonymization system characterized by the above.
2. The anonymization system according to Claim 1, wherein the anonymized data providing device, A first device that generates the random number for deletion data and the hash value; A second device that generates the anonymized data and the random number for anonymized data, and the first device and the second device are different computers communicably connected; the anonymized data user device acquires the anonymized data and the random number for anonymized data from the second device, characterizing an anonymization system. **Claim 3** The anonymization system according to claim 1, wherein the data user device can select any one of unprocessed, pseudonymized (with name matching), pseudonymized (without name matching) deletion, and generalization in the classification of specific expressions in the anonymization request. characterizing an anonymization system. **Claim 4** The anonymization system according to claim 1, wherein the anonymized data providing device and the anonymized data user device generate pseudonymized (with name matching) data based on the classification of specific expressions and a first index assigned based on name matching between specific expressions in the same classification in the first process. characterizing an anonymization system. **Claim 5** The anonymization system according to claim 1, wherein the anonymized data providing device and the anonymized data user device generate pseudonymized (without name matching) data based on the classification of specific expressions and a second index assigned based on name matching within each specific expression in the same classification in the second process. characterizing an anonymization system. **Claim 6** The anonymization system according to claim 1, further comprising a name matching data management device that manages data related to name matching, wherein the anonymized data providing device outputs data of a screen used for a first name matching regarding name matching of similar specific expressions in the text data, and the name matching data management device calculates the similarity between the specific expressions to be name - matched and those not to be name - matched in the first name matching and the specific expressions managed by the name matching data management device, wherein the anonymized data providing device outputs data of a screen that shows information on similar specific expressions based on the similarity, in which the specific expressions in the text data and those outside the text data are shown separately, and is used for a second name matching regarding name matching of similar specific expressions, and the name matching data management device In the second clustering, a first index is set for the specific expressions to be clustered, and a second index is set for the specific expressions not to be clustered. The anonymization data providing device and the anonymization data using device In the first process, anonymized (clustered) data is generated based on the classification of specific expressions and the first index assigned to each of the specific expressions in the same classification. The anonymization data providing device and the anonymization data using device In the second process, anonymized (non-clustered) data is generated based on the classification of specific expressions and the second index assigned to each of the specific expressions in the same classification. An anonymization system characterized by the above.
7. An anonymization method performed using an anonymization data providing device and an anonymization data using device capable of communicating with the anonymization data providing device, The anonymization data providing device When generating values for data verification, for each specific expression in the text data, anonymized (clustered) data obtained by replacing the original data with values that can be identified across multiple text data, the first process of generating a random number for anonymized (clustered) data based on the original data and the random number used for the original data, anonymized (non-clustered) data obtained by replacing the original data with values that can be identified only within the text data, the second process of generating a random number for anonymized (non-clustered) data based on the anonymized (clustered) data and the random number for anonymized (clustered) data, generalized data obtained by replacing the original data with generalized values, the third process of generating a random number for generalized data based on the anonymized (non-clustered) data and the random number for anonymized (non-clustered) data, deleted data obtained by replacing the generalized data with blanks, and the fourth process of generating a random number for deleted data based on the generalized data and the random number for generalized data are performed. When anonymizing data in response to an anonymization request regarding the anonymization of data from the anonymization data using device, for each element in the data, perform the processing required by the anonymization request among the first process, the second process, the third process, and the fourth process, generate anonymized data which is the anonymized data and a random number for anonymized data which is the random number based on the processing. The anonymization data using device Acquire the anonymized data and the random number for anonymized data via communication. When validating data, for each element in the anonymized data, perform the processing that the anonymized data providing device has not performed among the first processing, the second processing, the third processing, and the fourth processing, and generate a random number for deletion data. Verify the validity of the obtained anonymized data based on the random number for deletion data generated by the anonymized data providing device and the generated random number for deletion data. An anonymization method characterized by the above.
8. The anonymization method according to claim 7, wherein the anonymized data providing device comprises a first device that generates the random number for deletion data, and a second device that generates the anonymized data and the random number for anonymized data, wherein the first device and the second device are different computers connected communicably, and the anonymized data user device acquires the anonymized data and the random number for anonymized data from the second device. An anonymization method characterized by the above.
9. The anonymization method according to claim 7, wherein the anonymized data user device can select any one of unprocessed, pseudonymized (with name grouping), pseudonymized (without name grouping), deleted, and generalized in the classification unit of the specific expression in the anonymization request. An anonymization method characterized by the above.
10. The anonymization method according to claim 7, wherein the anonymized data providing device and the anonymized data user device generate pseudonymized (with name grouping) data based on the classification of specific expressions and the first index assigned based on name grouping between documents for each specific expression in the same classification in the first processing. An anonymization method characterized by the above.
11. The anonymization method according to claim 7, wherein the anonymized data providing device and the anonymized data user device generate pseudonymized (without name grouping) data based on the classification of specific expressions and the second index assigned based on name grouping within a document for each specific expression in the same classification in the second processing. An anonymization method characterized by the above.
12. The anonymization method according to claim 7, wherein the anonymization method is a method performed using a name grouping data management device that manages data related to name grouping, and the anonymized data providing device outputs data of a screen used for the first name grouping related to name grouping of similar specific expressions in the text data, and the name grouping data management device Calculate the similarity between the specific expressions to be grouped and those not to be grouped in the first grouping, and the specific expressions managed by the grouping data management device. The anonymized data providing device A screen showing information on specific expressions similar based on the similarity, in which the specific expressions in the text data and those outside the text data are shown separately, and outputs data of a screen used for a second grouping related to the grouping of similar specific expressions. The grouping data management device In the second grouping, set a first index for the specific expressions to be grouped and a second index for the specific expressions not to be grouped. The anonymized data providing device and the anonymized data user device In the first process, generate pseudonymized (grouped) data based on the classification of specific expressions and the first index attached to each of the specific expressions in the same classification. The anonymized data providing device and the anonymized data user device In the second process, generate pseudonymized (ungrouped) data based on the classification of specific expressions and the second index attached to each of the specific expressions in the same classification. An anonymization method characterized by the above.
Citation Information
Patent Citations
Anonymization system and anonymization method
JP2023060684A