Intermediary Mapping Layer for Non-Standard Dataset De-identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current solutions for de-identifying non-standard datasets are resource-intensive, requiring specialized expertise and manual effort, and often result in over- or under-estimation of disclosure risk, leading to reduced data utility or leakage of sensitive information.
Innovation Solution
An intermediary mapping and de-identification system that automates the conversion of non-standard datasets to a standard schema and variables, reducing the need for manual effort and expertise by encoding complex requirements and allowing for streamlined quality control and auditing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual de-identification processes are used with specialized expertise, then disclosure risk estimation accuracy is improved, but resource consumption and time requirements increase significantly
Solution Approach 1:
The system performs preliminary actions by pre-defining complex de-identification requirements and mapping rules in an intermediary mapping layer before actual de-identification execution. This allows the system to automatically handle complex transformations without requiring manual expert intervention during the actual processing phase, thus maintaining accuracy while improving efficiency.
Solution Approach 2:
The patent introduces an intermediary mapping layer that acts as a mediator between the raw dataset and the de-identification process. This intermediary layer pre-processes and structures the data according to predefined schemas, enabling automated downstream actions to accurately estimate disclosure risk without requiring manual expert analysis, thereby resolving the contradiction between accuracy and efficiency.
2Productivity
If automated de-identification processes are implemented, then processing speed and scalability are improved, but accuracy in risk estimation may deteriorate due to lack of expert judgment
Solution Approach 1:
The system performs preliminary actions by pre-defining complex de-identification requirements and mapping rules in an intermediary mapping layer before actual de-identification execution. This allows the system to automatically handle complex transformations without requiring manual expert intervention during the actual processing phase, thus maintaining accuracy while improving efficiency.
Solution Approach 2:
The patent transforms the de-identification process by changing parameters from manual expert judgment to automated rule-based processing. By pre-defining mapping schemas and transformation rules, the system converts complex qualitative expert knowledge into quantifiable, automated parameters that can be consistently applied at scale without sacrificing accuracy.
3Manufacturing precision
If complex manual ETL processes are performed to find identifiable information and variable connections, then de-identification accuracy is improved, but effort and expertise requirements increase to 5-10 days
Solution Approach 1:
The system performs preliminary actions by pre-defining complex de-identification requirements and mapping rules in an intermediary mapping layer before actual de-identification execution. This allows the system to automatically handle complex transformations without requiring manual expert intervention during the actual processing phase, thus maintaining accuracy while improving efficiency.
Solution Approach 2:
The patent enables the system to perform ETL processes and find variable connections automatically through self-service mechanisms. The intermediary mapping layer contains pre-configured rules and schemas that enable the system to autonomously identify identifiable information, establish variable connections, and perform transformations without human intervention, reducing processing time from 5-10 days to a fraction of that time while maintaining consistency.
4Adaptability or versatility
If non-standard datasets are processed without intermediary mapping, then data utility is preserved, but de-identification reliability deteriorates due to variability in data formats
Solution Approach 1:
The patent introduces an intermediary mapping layer that acts as a mediator between the raw dataset and the de-identification process. This intermediary layer pre-processes and structures the data according to predefined schemas, enabling automated downstream actions to accurately estimate disclosure risk without requiring manual expert analysis, thereby resolving the contradiction between accuracy and efficiency.
Solution Approach 2:
The patent transforms the de-identification process by changing parameters from manual expert judgment to automated rule-based processing. By pre-defining mapping schemas and transformation rules, the system converts complex qualitative expert knowledge into quantifiable, automated parameters that can be consistently applied at scale without sacrificing accuracy.
Data Source
AI summary
Disclosed is a method for an intermediary mapping an de-identification comprising steps of retrieving datasets and meta data from a data source; selecting a target standard; mapping the retrieved datasets and the metadata to the target standard, wherein the datasets and the metadata are mapped to the target standard using one of, a schema mapping, a variable mapping, or a combination thereof; infer one or more of, variable classifications, variable connections, groupings, disclosure risk settings, and de-identification settings using the dataset mapping and metadata; perform a de-identification propagation using the mapped datasets, the mapped metadata, the inferred variable classifications, the inferred variable connections, the inferred groupings, the inferred disclosure risk settings, the inferred de-identification settings, or a combination thereof.


