Data management device, method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON TELEGRAPH & TELEPHONE CORP
- Filing Date
- 2023-01-12
- Publication Date
- 2026-07-31
AI Technical Summary
【0008】 複数の匿名化データを連結したデータの有用性悪化を低減する技術が提供される。
Smart Images

Figure 0007898111000003 
Figure 0007898111000004 
Figure 0007898111000005
Abstract
Claims
1. A data management device that creates concatenated data by concatenating vertically partitioned data, which consists of records having one identifier and one or more attributes, A kana conversion unit is configured to replace the value of the identifier with a kana character through coordinated calculation with other data management devices, An identification unit configured to deidentify the values of one or more of the aforementioned attributes, A replacement unit is configured to create anonymized data with kana names by randomly replacing each record that constitutes the vertically partitioned data after deidentification. The device includes a connecting unit configured to create concatenated data by linking kana-named anonymized data created by the aforementioned other data management device with kana-named anonymized data created by itself using a common kana character. The de-identification unit is, The properties of the shuffle model are satisfied under a given privacy protection strength (ε, δ). 0 , δ 0 A data management device configured to de-identify the values of one or more attributes for each record constituting the vertically partitioned data by local differential privacy with a privacy protection strength of ).
2. The de-identification unit is, The i-th record included in the vertically partitioned data is x i , the i-th record x i The data obtained by removing the pseudonym from the values of the attributes included in is x i ', the number of records included in the vertically partitioned data is m, (ε 0 , δ 0 ) is a mechanism M that satisfies local differential privacy with privacy protection strength. When (i) is set as M (1) (x 1 '),..., M (m) (x m '), it is configured to anonymize the values of the one or more attributes, The aforementioned replacement part is Let π be the substitution function that randomly replaces {1, ..., m}, and c be the kana character contained in the i-th record. 0 When (i) is set, ((c 0 '(π(1)), M (π(1)) (x π(1) ')), ..., (c 0 '(π(m)), M (π(m)) (x π(m) '))) τ The data management device according to claim 1, configured to create the anonymized data with pseudonyms.
3. The property of the aforementioned shuffle model is ε = ln(1 + ((e^ε) 0 -1) / (e^ε 0 +1))((8√(e^ε 0 ln(4 / δ)) / √m)+(8e^ε 0 ) / m)), δ=δ'+(e^ε+1)(1+e^(-ε 0 ) / 2)mδ 0 (However, δ'∈[0,1] is ε 0 The data management device according to claim 2, satisfying ≤ ln(m / (16ln(2 / δ'))).
4. The data management device according to any one of claims 1 to 3, further comprising a data synthesis unit configured to generate random pseudo-data that retains the statistical properties of the concatenated data from the concatenated data or the statistical quantities of the concatenated data by a data synthesis method.
5. The data synthesis unit, The data management device according to claim 4, configured to estimate the posterior distribution of the statistics of the concatenated data by a statistical reconstruction method corresponding to the algorithm for realizing the local difference privacy, calculate the statistics, and generate the pseudo-data from the statistics.
6. A data management device that creates concatenated data by concatenating vertically partitioned data, which consists of records having one identifier and one or more attributes, A kana encoding procedure that replaces the value of the identifier with a kana character through coordinated calculation with other data management devices, A deidentification procedure for deidentifying the value of one or more of the aforementioned attributes, A replacement procedure to create anonymized data with pseudonyms by randomly replacing each record that constitutes the vertically partitioned data after deidentification, The following is a linking procedure to create linked data by linking the anonymized data with pseudonyms created by the aforementioned other data management device and the anonymized data with pseudonyms created by itself using a common pseudonym: The aforementioned de-identification procedure is: The properties of the shuffle model are satisfied under a given privacy protection strength (ε, δ). 0 , δ 0 A data management method that de-identifies the values of one or more attributes for each record constituting the vertically partitioned data by using local differential privacy with a privacy protection strength of ).
7. A data management device that creates concatenated data by concatenating vertically partitioned data, which consists of records having one identifier and one or more attributes, A kana encoding procedure that replaces the value of the identifier with a kana character through coordinated calculation with other data management devices, A deidentification procedure for deidentifying the value of one or more of the aforementioned attributes, A replacement procedure to create anonymized data with pseudonyms by randomly replacing each record that constitutes the vertically partitioned data after deidentification, The procedure involves creating concatenated data by linking the anonymized data with pseudonyms created by the aforementioned other data management device and the anonymized data with pseudonyms created by itself using a common pseudonym, and then executing this procedure. The aforementioned de-identification procedure is: The properties of the shuffle model are satisfied under a given privacy protection strength (ε, δ). 0 , δ 0 A program that de-identifies the values of one or more attributes for each record constituting the vertically partitioned data by using local differential privacy with a privacy protection strength of ).