Data De-identification Apparatus for Cross-Field Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data de-identification technologies fail to integrate data across different fields while ensuring legal compliance, resulting in loss of data utility for specific uses like credit evaluation.
Innovation Solution
A data de-identification apparatus and method that determines identification categories based on industries and data use, transforming data sets to maintain richer information and comply with cross-field regulations, using a processor connected to storage and input interfaces to perform data transformation and de-identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional data de-identification technologies delete, encrypt, or superordinate directly identifiable data, then personal information protection is improved, but data utility for specific evaluations deteriorates
Solution Approach 1:
The patent applies local quality by treating different data fields differently based on their identification risk and utility value. Sensitive fields like names and ID numbers undergo strict de-identification ( deletion or encryption), while less sensitive fields like occupation and education are transformed or anonymized to varying degrees. This selective approach preserves data utility for evaluation purposes while protecting personal information at appropriate levels.
Solution Approach 2:
The patent changes parameters by transforming data representation rather than simply deleting it. For example, numerical data is transformed by masking certain digits (e.g., showing only last 4 digits of phone numbers), categorical data is transformed by generalization (e.g., transforming specific company names to industry categories). These parameter changes maintain the statistical and analytical utility of data while removing personally identifiable information.
2Measurement precision
If data is integrated across different fields, then decision accuracy and value creation are improved, but compliance with legal norms deteriorates
Solution Approach 1:
The patent applies preliminary action by performing de-identification and transformation operations on data before it is integrated across different fields. The system automatically identifies sensitive fields, applies appropriate de-identification techniques, and transforms data into anonymized forms prior to cross-field integration. This ensures that data compliance requirements are met before the data is used for cross-field analysis and decision-making.
Solution Approach 2:
The patent introduces an intermediary de-identification system that acts as a mediator between raw data and cross-field integration processes. This intermediary layer transforms personally identifiable information into anonymized data that can be safely integrated across fields while maintaining the statistical properties needed for accurate analysis and decision-making.
3Loss of information
If data is transformed to maintain richer information, then data utility is improved, but complexity of de-identification process deteriorates
Solution Approach 1:
The patent applies segmentation by dividing the de-identification process into distinct modules: field identification, sensitivity assessment, transformation selection, and data transformation. Each module handles a specific aspect of the de-identification process, making the overall complex system manageable and maintainable. Different transformation techniques are segmented and applied to different types of fields based on their characteristics.
Solution Approach 2:
The patent implements universality by creating a multi-functional de-identification system that can handle various data types (numerical, categorical, text) and apply multiple transformation techniques (masking, generalization, encryption, deletion) through a unified framework. The system automatically selects appropriate transformation methods based on field characteristics, reducing the need for separate processing pipelines for different data types.
Data Source
AI summary
A data de-identification apparatus and method are provided. The data de-identification apparatus stores a data set of a first industry, wherein the data set is defined with a plurality of fields. The data de-identification apparatus receives a first instruction and a second instruction, wherein the first instruction corresponds to a second industry and the second instruction corresponds to a use of data. The data de-identification apparatus determines an identification category for each of the fields according to the first industry, the second industry, and the use of data. The data de-identification apparatus transforms the data set into a transformed data set according to the use of data and then transforms the transformed data set into a de-identification data set according to the identification categories.


