Identification Risk Assessment Apparatus for Anonymized Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Pk-anonymity overestimates the probability of individual identification in anonymized database records, leading to an inflated risk assessment when an attacker's background knowledge is considered.
Innovation Solution
An identification-estimation risk assessment apparatus calculates a risk assessment value by determining the probability that a record in an original table is identified as corresponding to a record in an anonymized table, using a first value representing the sum of probabilities of all shuffles that result in a matching anonymized table and a second value representing the sum of probabilities for specific record mappings, thereby providing a more accurate risk assessment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If Pk-anonymity is used to assess individual identification risk, then the assessment covers all possible attacker background knowledge, but the identification probability is overestimated and risk is inflated
Solution Approach 1:
The patent applies local quality by fixing the attacker's background knowledge to a specific state (knowing only the anonymized table structure and size) rather than considering all possible knowledge states. This localizes the assessment to a particular scenario, yielding a more precise and less inflated risk estimate while maintaining reliability for that specific context.
Solution Approach 2:
Instead of assessing risk by considering all possible attacker knowledge (Pk-anonymity approach), the patent inverts the approach by fixing the attacker's knowledge to a minimal state and calculating identification probability from that baseline. This inversion provides a more accurate measure by eliminating the overestimation inherent in the comprehensive approach.
2Adaptability or versatility
If all possible attacker background knowledge is considered, then the assessment is comprehensive, but the complexity of calculating identification probability increases
Solution Approach 1:
The patent extracts and isolates the essential elements needed for risk assessment by fixing the attacker's background knowledge to a minimal, well-defined state. This extraction removes the complexity of considering all possible knowledge states while retaining the core assessment functionality, thereby reducing calculation complexity without completely sacrificing comprehensiveness.
3Reliability
If the probability of individual identification is calculated using Pk-anonymity, then the assessment is conservative, but the result is higher than necessary and misleading
Solution Approach 1:
The patent changes the parameters of the risk assessment model by fixing the attacker's background knowledge to a specific state rather than considering all possible states. This parameter change transforms the assessment from a conservative, inflated estimate to a more precise, accurate measurement that reflects the actual risk under defined conditions.
Data Source
AI summary
An identification-estimation risk assessment apparatus includes a first calculation unit for calculating a first value ΣγPr[Δ(t)=t′oγ], which is a sum related to all shuffles of a probability Pr[Δ(t)=t′oγ] that a table Δ(t) obtained by anonymizing the original table t through randomization Δ coincides with a table t′oγ represented as a composite of a shuffle γ and the anonymized table t′; a second calculation unit for calculating a second value Σγ(r)=r′Pr[Δ(t)=t′oγ], which is a sum related to a shuffle that satisfies γ(r)=r′ of the probability Pr[Δ(t)=t′oγ]; and a third calculation unit for calculating a third value Σγ(r)=r′Pr[Δ(t)=t′oγ]/ΣγPr[Δ(t)=t′oγ] as a risk assessment value of a risk that a record number r is identified and estimated as corresponding to a record number r′, based on the first value and the second value.


