Anonymization Engine for Data Labeling Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data labeling processes face challenges when using data with personally identifiable information (PII), as legal and ethical issues prevent exposure of such data to human labelers, necessitating a method to anonymize data while maintaining accurate labeling capabilities.
Innovation Solution
A system and method for anonymizing data files, including videos, images, and audio clips, by removing PII using a processor-based anonymization engine, generating anonymized files, and allowing human labelers to label these files without exposing them to PII, while storing labels associated with the original data files.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If personally identifiable information is retained in data files for labeling, then labeling accuracy can be maintained, but legal and ethical compliance deteriorates due to exposure risks to human labelers
Solution Approach 1:
The system segments the data processing workflow into two distinct phases: (1) automated PII detection and redaction to create anonymized versions of data files, and (2) human labeling of the anonymized files. This segmentation allows the system to maintain labeling accuracy by preserving non-PII content while eliminating legal and ethical risks by removing PII before human exposure.
Solution Approach 2:
An automated PII detection and redaction system serves as an intermediary between the original data files and human labelers. This intermediary process automatically identifies and redacts PII, generating anonymized data files that can be safely processed by humans without compromising labeling quality or violating compliance requirements.
2Object-affected harmful factors
If personally identifiable information is removed from data files, then legal and ethical compliance is improved, but labeling quality may deteriorate due to loss of contextual information
Solution Approach 1:
The system extracts only the personally identifiable information from data files while preserving all other contextual content. By selectively removing only PII elements (such as names, addresses, phone numbers) and retaining the rest of the data structure and context, the system maintains labeling quality while achieving compliance.
Solution Approach 2:
The redaction process applies local quality changes by selectively modifying only the portions of data containing PII while leaving the rest of the content intact. This targeted approach ensures that contextual information necessary for accurate labeling remains preserved, while only the specific problematic elements are removed or obscured.
3Object-affected harmful factors
If a PII redaction system is implemented, then compliance with legal and ethical standards is improved, but processing time and system complexity increase
Solution Approach 1:
The PII redaction system operates autonomously without requiring manual review or intervention. The automated detection and redaction processes work independently to identify and remove PII, generating anonymized data files ready for labeling. This self-service capability reduces the need for complex manual oversight mechanisms while maintaining compliance.
Solution Approach 2:
The system performs PII redaction as a preliminary step before the labeling process begins. By completing the anonymization process in advance, the system eliminates the need for complex real-time compliance checking during labeling operations, thereby reducing overall system complexity while ensuring compliance is maintained throughout the workflow.
Data Source
AI summary
A method and system are disclosed for anonymizing data for labeling and development purposes. A data storage backend has a database of non-anonymous data that is received from a data source. An anonymization engine of the data storage backend generates anonymized data by removing personally identifiable information from the non-anonymous data. These anonymized data are made available to human labelers who manually provide labels based on the anonymized data using a data labeling tool. These labels are then stored in association with the corresponding non-anonymous data, which can then be used for training one or more machine learning models. In this way, non-anonymous data having personally identifiable information can be manually labelled for development purposes without exposing the personally identifiable information to any human labelers.


