Traceable Data Audit Apparatus for Privacy-Preserving Leakage Source Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack a mechanism to trace data once it is sold and distributed, making it impossible for enterprises to verify the origin of their data even if it is leaked, without compromising the data's precision.
Innovation Solution
A traceable data audit apparatus and method that includes a storage unit, interface, and processing unit, which de-identifies sensitive information while adding traceable data during the de-identification process, storing audit logs with consumer identities and evidence, allowing for the identification of data leakage sources by comparing leaked data sets with existing audit logs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If de-identification operation is applied to protect consumer privacy, then consumer privacy is protected, but data traceability is lost
Solution Approach 1:
The patent segments the data processing into two distinct parts: de-identification operation that protects consumer privacy by removing or masking personal information, and a separate traceability operation that adds watermark information or unique identifiers to each released data set. This segmentation allows both privacy protection and traceability to coexist independently.
Solution Approach 2:
The patent introduces an intermediary element - a traceability marker or watermark - that is embedded in the released data set during the de-identification process. This intermediary carries identification information without revealing consumer personal data, thus mediating between privacy protection requirements and traceability needs.
2Adaptability or versatility
If data is released to consumers, then data utility is improved, but data leakage risk increases
Solution Approach 1:
The patent applies preliminary action by embedding traceability markers into the data set before releasing it to consumers. This proactive measure ensures that if data leakage occurs, the source can be identified, thus mitigating the harmful effects of data leakage while maintaining data utility for legitimate uses.
Solution Approach 2:
The patent establishes a feedback mechanism where traceability information from leaked data can be compared against audit logs of released data sets. This feedback loop enables the system to identify leakage sources and respond appropriately, reducing the overall risk associated with data release.
3Loss of information
If audit logs are maintained for all released data, then data traceability is improved, but system complexity increases
Solution Approach 1:
The patent extracts only the essential traceability information - such as unique identifiers or watermark data - from the complete audit logs for comparison purposes. This extraction reduces the amount of data that needs to be processed and stored, thereby reducing system complexity while maintaining traceability capability.
Solution Approach 2:
The patent changes the parameter of audit log storage by storing only critical traceability markers rather than complete data sets. This parameter change significantly reduces storage requirements and processing complexity while preserving the ability to trace data origins when needed.
Data Source
AI summary
A traceable data audit apparatus, method, and non-transitory computer readable storage medium thereof are provided. The traceable data audit apparatus is stored with an original data set. The original data set includes a plurality of records and is defined with a plurality of fields. Each of the records has a plurality of items corresponding to the fields one-on-one. The fields are classified into an identity sensitive subset and an identity insensitive subset. The traceable data audit apparatus generates a released data set by applying a de-identification operation to each of the items corresponding to the fields in the identity sensitive subset and stores an audit log of the original data set. The audit log includes a date, a consumer identity, an identity of the original data set, and a plurality of evidences. Each of the evidence is one of the records of the released data set.


