Data management device, method, and program

JP7898111B2Active Publication Date: 2026-07-31NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NIPPON TELEGRAPH & TELEPHONE CORP
Filing Date
2023-01-12
Publication Date
2026-07-31

AI Technical Summary

Benefits of technology

【0008】 複数の匿名化データを連結したデータの有用性悪化を低減する技術が提供される。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007898111000003
    Figure 0007898111000003
  • Figure 0007898111000004
    Figure 0007898111000004
  • Figure 0007898111000005
    Figure 0007898111000005
Patent Text Reader

Abstract

To provide a data management apparatus, or the like, which reduces deterioration of the utility of data obtained by combining multiple pieces of anonymized data.SOLUTION: A data management apparatus includes a pseudonym conversion unit, a non-identification unit, a replacement unit, and a combining unit. The pseudonym conversion unit performs cooperative computation with another data management apparatus to replace a value of an identifier with a pseudonym. The non-identification unit makes values of one or more attributes non-identifiable. The replacement unit generates pseudonymous non-identifiable data by randomly replacing records constituting vertically split data which has been made non-identifiable. The combining unit generates combined data by combining pseudonymous non-identifiable data generated by the other data management apparatus with the pseudonymous non-identifiable data generated by itself with a common pseudonym. The non-identification unit makes the values of one or more attributes non-identifiable for each of the records constituting the vertically split data using Local Differential Privacy that defines (ε0,δ0) satisfying the property of a shuffle model under a given privacy protection strength (ε,δ), as privacy protection strength.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Claims

1. A data management device that creates concatenated data by concatenating vertically partitioned data, which consists of records having one identifier and one or more attributes, A kana conversion unit is configured to replace the value of the identifier with a kana character through coordinated calculation with other data management devices, An identification unit configured to deidentify the values ​​of one or more of the aforementioned attributes, A replacement unit is configured to create anonymized data with kana names by randomly replacing each record that constitutes the vertically partitioned data after deidentification. The device includes a connecting unit configured to create concatenated data by linking kana-named anonymized data created by the aforementioned other data management device with kana-named anonymized data created by itself using a common kana character. The de-identification unit is, The properties of the shuffle model are satisfied under a given privacy protection strength (ε, δ). 0 , δ 0 A data management device configured to de-identify the values ​​of one or more attributes for each record constituting the vertically partitioned data by local differential privacy with a privacy protection strength of ).

2. The de-identification unit is, The i-th record included in the vertically partitioned data is x i , the i-th record x i The data obtained by removing the pseudonym from the values of the attributes included in is x i ', the number of records included in the vertically partitioned data is m, (ε 0 , δ 0 ) is a mechanism M that satisfies local differential privacy with privacy protection strength. When (i) is set as M (1) (x 1 '),..., M (m) (x m '), it is configured to anonymize the values of the one or more attributes, The aforementioned replacement part is Let π be the substitution function that randomly replaces {1, ..., m}, and c be the kana character contained in the i-th record. 0 When (i) is set, ((c 0 '(π(1)), M (π(1)) (x π(1) ')), ..., (c 0 '(π(m)), M (π(m)) (x π(m) '))) τ The data management device according to claim 1, configured to create the anonymized data with pseudonyms.

3. The property of the aforementioned shuffle model is ε = ln(1 + ((e^ε) 0 -1) / (e^ε 0 +1))((8√(e^ε 0 ln(4 / δ)) / √m)+(8e^ε 0 ) / m)), δ=δ'+(e^ε+1)(1+e^(-ε 0 ) / 2)mδ 0 (However, δ'∈[0,1] is ε 0 The data management device according to claim 2, satisfying ≤ ln(m / (16ln(2 / δ'))).

4. The data management device according to any one of claims 1 to 3, further comprising a data synthesis unit configured to generate random pseudo-data that retains the statistical properties of the concatenated data from the concatenated data or the statistical quantities of the concatenated data by a data synthesis method.

5. The data synthesis unit, The data management device according to claim 4, configured to estimate the posterior distribution of the statistics of the concatenated data by a statistical reconstruction method corresponding to the algorithm for realizing the local difference privacy, calculate the statistics, and generate the pseudo-data from the statistics.

6. A data management device that creates concatenated data by concatenating vertically partitioned data, which consists of records having one identifier and one or more attributes, A kana encoding procedure that replaces the value of the identifier with a kana character through coordinated calculation with other data management devices, A deidentification procedure for deidentifying the value of one or more of the aforementioned attributes, A replacement procedure to create anonymized data with pseudonyms by randomly replacing each record that constitutes the vertically partitioned data after deidentification, The following is a linking procedure to create linked data by linking the anonymized data with pseudonyms created by the aforementioned other data management device and the anonymized data with pseudonyms created by itself using a common pseudonym: The aforementioned de-identification procedure is: The properties of the shuffle model are satisfied under a given privacy protection strength (ε, δ). 0 , δ 0 A data management method that de-identifies the values ​​of one or more attributes for each record constituting the vertically partitioned data by using local differential privacy with a privacy protection strength of ).

7. A data management device that creates concatenated data by concatenating vertically partitioned data, which consists of records having one identifier and one or more attributes, A kana encoding procedure that replaces the value of the identifier with a kana character through coordinated calculation with other data management devices, A deidentification procedure for deidentifying the value of one or more of the aforementioned attributes, A replacement procedure to create anonymized data with pseudonyms by randomly replacing each record that constitutes the vertically partitioned data after deidentification, The procedure involves creating concatenated data by linking the anonymized data with pseudonyms created by the aforementioned other data management device and the anonymized data with pseudonyms created by itself using a common pseudonym, and then executing this procedure. The aforementioned de-identification procedure is: The properties of the shuffle model are satisfied under a given privacy protection strength (ε, δ). 0 , δ 0 A program that de-identifies the values ​​of one or more attributes for each record constituting the vertically partitioned data by using local differential privacy with a privacy protection strength of ).