Data Masking System Using Statistical Transformations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In financial and healthcare systems, customer-related information is often exposed in production environments, making it difficult to protect sensitive data from unauthorized access, as malicious outsiders can re-identify entities using public sources, and existing methods fail to effectively mask information without affecting system behavior.
Innovation Solution
A method and system that mask sensitive data by replacing it with fictional data of the same type and format, using discrete transforms to alter statistical distributions, ensuring the masked data maintains original statistics while protecting personal information, and using one-to-one, one-to-many, or many-to-one mapping techniques to obscure identities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If sensitive customer information is displayed in production environment for processing and testing, then system functionality and testing capability are improved, but data security and protection from unauthorized access deteriorate
Solution Approach 1:
The patent creates fictional copies of sensitive data that replicate the statistical properties and format of real customer information. These synthetic data copies enable testing and development activities while the original sensitive data remains protected. The fictional data maintains the same structure, type, and statistical distribution characteristics as the real data, allowing systems to function normally without exposing actual customer information.
Solution Approach 2:
The patent transforms sensitive data by altering its statistical parameters and distribution characteristics while preserving the format and structure. By changing the underlying statistical properties (such as distribution patterns, correlations, and frequency characteristics) while maintaining the apparent data structure, the system enables testing functionality while protecting the actual sensitive information from re-identification attacks.
2Object-affected harmful factors
If data is masked to protect sensitive information, then data security is improved, but statistical integrity and reporting accuracy deteriorate
Solution Approach 1:
The patent carefully modifies statistical parameters of the data while preserving the overall distribution characteristics needed for reporting. By selectively changing certain parameters (such as individual record values) while maintaining aggregate statistical properties (such as means, variances, and distribution shapes), the system achieves both protection of sensitive information and preservation of statistical integrity for auditing and reporting purposes.
Solution Approach 2:
The patent applies partial masking where only specific portions of the data that are most susceptible to re-identification attacks are transformed. By selectively applying transformations to certain data elements while leaving others intact, the system maintains sufficient statistical information for reporting and auditing while protecting the most sensitive identifying characteristics.
3Object-affected harmful factors
If fictional data is used to replace sensitive information, then re-identification resistance is improved, but data utility for analysis deteriorates
Solution Approach 1:
The patent transforms data parameters to create fictional values that resist re-identification while preserving the statistical relationships needed for analysis. By changing individual data point parameters while maintaining the overall statistical distribution and correlation structures, the fictional data becomes resistant to re-identification attacks yet retains sufficient utility for analytical purposes such as trend analysis and pattern recognition.
Data Source
AI summary
A data-masking tool encoded on one or more computing readable storage media that includes a code that uses a combination of fields that uniquely identifies data in a record and utilizing it as a reference to mask original data with substitute values, by either aggregating several into one, mapping one-to-one or expanding one into a set.


