Automated Data Masking for Unknown Data Types and Format Preservation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional data masking techniques fail to preserve the format of sensitive data elements, leading to errors in application processing and inefficiencies, especially when metadata is uncertain or misleading, and existing format-preserving encryption methods struggle with complex data formats.
Innovation Solution
A novel data masking approach that utilizes irreversible functions and syntactic characterizations to transform data elements while maintaining their format, allowing for secure and efficient masking without encryption, using irreversible functions like hash algorithms and syntactic definitions to generate masked data elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data masking techniques are applied to data elements with specialized formats, then data security is improved, but application compatibility deteriorates and processing errors increase
Solution Approach 1:
The patent changes the parameter of data transformation from traditional masking (which alters data format) to format-preserving masking (which maintains data format). This is achieved by using deterministic transformations that preserve the syntactic structure, data types, and validation rules of the original data, thereby maintaining application compatibility while still providing security through controlled obfuscation of sensitive information.
Solution Approach 2:
The patent applies local quality by selectively masking only specific portions of data elements that contain sensitive information while preserving the overall format and non-sensitive portions. This allows applications to continue processing masked data without errors while the sensitive parts remain protected, resolving the contradiction between security and compatibility.
2Shape
If format preserving encryption is used to handle specialized data formats, then data format preservation is improved, but system complexity increases due to cryptographic material management
Solution Approach 1:
The patent replaces complex cryptographic material management with simpler, disposable masking rules that can be generated and applied without requiring long-term management of cryptographic keys. The masking transformations use deterministic algorithms that do not require sensitive cryptographic material, thereby reducing system complexity while maintaining format preservation.
3Device complexity
If traditional data profiling approaches are used without metadata, then data type determination becomes simpler, but accuracy deteriorates leading to security risks
Solution Approach 1:
The patent applies preliminary action by performing data profiling and determining data types before applying masking transformations. This preliminary analysis includes examining data patterns, structures, and characteristics to accurately identify data types even without metadata. By conducting this analysis in advance, the system ensures accurate data type determination while maintaining a relatively simple process flow.
Data Source
AI summary
A system, method and computer-readable medium for generating a data masking syntactic definition for a data element of an unknown data type, including generating one or more alphabets corresponding to one or more element member positions of the data element based at least in part on element members occurring at each element member position in a plurality of data elements of the unknown type, each alphabet comprising a set of one or more sequential element members that have occurred in the plurality of data elements at an element member position and generating a positional map describing a syntactic structure of the data element by mapping at least one of the one or more alphabets to each element member position of the data element.


