Audio and Visual PII Sanitization for Identity-Neutral ML
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing devices fail to sanitize personally identifiable information (PII) in captured audio and visual data, potentially leading to data leaks and privacy issues, even when the data is processed by identity-neutral machine learning endpoints.
Innovation Solution
Implement a PII sanitizing module within the computing device that removes, obfuscates, or transforms PII in audio and visual data, ensuring privacy preservation while maintaining the usability of the data for machine learning tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If PII is removed or obfuscated from audio/visual data, then privacy protection is improved, but data quality and model training effectiveness deteriorate
Solution Approach 1:
The patent segments the audio/visual data processing into distinct stages: original data capture, PII identification and removal/obfuscation, and sanitized data delivery to ML endpoints. This segmentation allows privacy protection to be applied selectively without compromising the overall utility of the data for non-PII features.
Solution Approach 2:
The patent applies different processing treatments to different portions of the data: PII-containing regions are sanitized while non-PII regions are preserved in their original form. This local quality approach ensures that only the necessary portions are modified, maintaining maximum data utility while achieving privacy goals.
2Loss of information
If PII is preserved in captured data, then data utility for ML training is maintained, but data privacy and security deteriorate
Solution Approach 1:
The patent extracts and removes PII components from the audio/visual data before delivery to ML endpoints. By taking out the harmful PII elements while retaining the beneficial non-PII data, the system maintains data utility for training purposes while eliminating privacy security risks.
3Manufacturing precision
If data is homogenized by removing PII, then model accuracy and size are improved, but data diversity and representativeness deteriorate
Solution Approach 1:
The patent changes the parameters of PII-containing data through removal or obfuscation transformations. This parameter change approach homogenizes the data by eliminating variable PII characteristics while preserving the essential non-PII features, thereby improving model accuracy without completely sacrificing data diversity.
Data Source
AI summary
Techniques for sanitizing personally identifiable information (PII) from audio and visual data are provided. For example, in a scenario where the data comprises an audio signal with speech uttered by a speaker S, these techniques can include removing, obfuscating, or transforming speech related and non-speech related audio cues in the audio signal that can be used to trace the identity of S, while allowing the content of S's speech to remain recognizable. As another example, in a scenario where the data comprises an image or video in which a person P appears, these techniques can include removing, obfuscating, or transforming P's visible biological features and visual indicators of P's location, belongings, or personal data in the image/video, while allowing the general nature of the footage to remain discernable. Through this PII sanitization process, the privacy of individuals portrayed in the audio or visual data can be preserved.


