Deterministic Data Obfuscation Preserving Statistical Integrity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Software developers lack access to production data for security reasons, necessitating realistic test data that mimics production data characteristics, while ensuring sensitive information is obfuscated and meeting varying project, privacy, and legal requirements.
Innovation Solution
A method and system for obfuscating data by reading values from a data source, generating deterministic obfuscated values using a key, and storing them in a data storage system, preserving referential integrity and statistical characteristics, with the ability to perform parallel processing and prevent reverse engineering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If production data is used for testing, then test data realism is improved, but data security and privacy are compromised
Solution Approach 1:
The patent creates obfuscated copies of production data that preserve statistical characteristics and data relationships while removing sensitive information. The obfuscation process generates test data that is statistically indistinguishable from real production data but contains no actual sensitive values, thus achieving both realism and security.
Solution Approach 2:
The patent transforms production data by changing specific parameters (sensitive values) while preserving others (statistical characteristics, relationships). The obfuscation process modifies data parameters such as replacing actual names with obfuscated identifiers, altering dates while preserving temporal relationships, and transforming sensitive fields while maintaining data distribution patterns.
2Object-affected harmful factors
If obfuscation is applied to protect sensitive information, then data security is improved, but data utility for testing deteriorates
Solution Approach 1:
The obfuscation process selectively transforms only sensitive parameters while preserving statistical characteristics, data relationships, and structural properties. This allows the data to maintain its utility for testing while protecting sensitive information through controlled parameter transformation.
Solution Approach 2:
The patent applies obfuscation partially - only to sensitive fields and values - while leaving other aspects of the data intact. This partial action preserves the statistical characteristics and relationships needed for testing while providing sufficient protection for sensitive information.
3Stability of the object's composition
If deterministic obfuscation is used to maintain referential integrity, then data relationship preservation is improved, but reversibility risk increases
Solution Approach 1:
The patent introduces an obfuscation key as an intermediary that enables consistent transformation while preventing reverse engineering. The key acts as a mediator that allows deterministic obfuscation for maintaining relationships but adds a layer of security that prevents unauthorized recovery of original values.
Solution Approach 2:
The deterministic obfuscation process applies consistent parameter transformation using cryptographic functions and obfuscation keys. This transforms original values into obfuscated values in a repeatable manner that preserves relationships but makes reversal computationally infeasible without the key.
4Manufacturing precision
If data obfuscation is performed manually, then processing accuracy is improved, but processing efficiency deteriorates
Solution Approach 1:
The patent replaces manual obfuscation processes with automated computational systems that apply cryptographic functions and statistical transformations. This substitution maintains high precision through algorithmic consistency while dramatically improving processing speed and scalability.
Solution Approach 2:
The automated obfuscation process uses computational algorithms to apply complex parameter transformations that would be impractical to perform manually. The system efficiently handles large volumes of data while maintaining statistical accuracy and relationship preservation through programmed logic.
Data Source
AI summary
A method for obfuscating data includes: reading values occurring in one or more fields of multiple records from a data source; storing a key value; for each of multiple of the records, generating an obfuscated value to replace an original value in a given field of the record using the key value such that the obfuscated value depends on the key value and is deterministically related to the original value; and storing the collection of obfuscated data including records that include obfuscated values in a data storage system.


