Name Data Masking Preserving Ethnicity and Gender Semantics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data masking techniques fail to effectively mask name data, losing essential characteristics and relevancy, which disrupts data integrity and makes it unsuitable for thorough testing or demonstration, leading to inappropriate or odd appearances that discourage customers.
Innovation Solution
A method for efficiently determining replacement name data that preserves the semantics of actual name data by classifying script, ethnicity, gender, and form, using a database architecture with data quality services to generate substitute names that maintain ethnic and gender relevance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If substitution or shuffling techniques are used to mask name data, then data privacy is protected, but the semantic characteristics and relevancy of name data are lost
Solution Approach 1:
The patent changes the parameters of name data by preserving semantic characteristics such as ethnicity, gender, and script type while substituting the actual name values. This allows masked data to maintain relevancy for testing scenarios while protecting privacy, resolving the contradiction between privacy protection and information preservation
Solution Approach 2:
The patent applies different masking strategies to different aspects of name data - preserving certain characteristics (ethnicity, gender, script) while changing others (specific name values). This local differentiation maintains the necessary semantic quality for testing while achieving privacy protection
2Reliability
If conventional masking techniques are applied to name data, then privacy concerns are addressed, but data integrity links are broken
Solution Approach 1:
The patent changes masking parameters to preserve data integrity by maintaining semantic characteristics and relationships in the masked data, ensuring that data links and referential integrity are not broken while still achieving privacy protection
3Reliability
If substitution techniques replace names with random values, then privacy is protected, but the masked data appears odd or inappropriate for demonstrations
Solution Approach 1:
The patent changes the approach from random substitution to parameter-based substitution, where replacement names are selected to match the semantic parameters (ethnicity, gender, script) of the original data. This ensures masked data appears realistic and appropriate for customer demonstrations while maintaining privacy protection
4Reliability
If numeric variance or cloaking techniques are used, then certain data types are protected, but name data characteristics are not preserved
Solution Approach 1:
The patent applies specialized masking logic tailored to name data characteristics, preserving semantic information such as ethnicity, gender, and script type. This local customization for name data prevents the loss of critical semantic information that would occur with generic masking techniques
Data Source
AI summary
A system includes reception of name data, determination, for each of a plurality of name properties, of an associated property value based on the name data, determination of a gender classification based on the property values, and, for each property value, generation of a substitute property value based on the property associated with the property value and the gender classification.


