Audio Sensitive Data Replacement via Voice Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for protecting sensitive data in audio files, such as personal information, often disrupt the natural flow of conversations by replacing key elements with audio tones or removing them entirely, making it difficult to maintain the original conversation's understanding and seamless playback.
Innovation Solution
A computer-implemented method using a signal separator to extract voices, transcribing audio to text, identifying sensitive data, replacing it with semantically meaningful synthetic data voiced by a matching synthetic voice, and outputting a new audio file that preserves the original conversation's integrity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If sensitive data is replaced with audio tones or removed entirely, then sensitive information is protected, but the natural flow and understanding of the conversation deteriorates
Solution Approach 1:
The patent introduces an intermediary system that includes: (1) converting audio to text representation, (2) identifying sensitive data through NLP analysis, (3) generating synthetic replacement text that preserves contextual meaning, and (4) converting the modified text back to audio. This intermediary processing chain allows sensitive data protection while maintaining conversation flow and understanding.
Solution Approach 2:
The patent replaces the traditional mechanical approach of simply muting or removing audio segments with an intelligent text-based processing system. By substituting the audio processing mechanism with text analysis and synthesis (using NLP and generative models), the system can preserve semantic meaning while protecting sensitive information.
2Reliability
If key elements of conversation are removed to protect sensitive data, then privacy security is improved, but seamless playback and conversation integrity worsen
Solution Approach 1:
The patent changes the parameter of data representation from raw audio to text, enabling precise identification and selective replacement of sensitive information. By operating in the text domain where sensitive data can be accurately identified through NLP, the system can protect privacy while preserving the overall conversation structure and integrity.
Solution Approach 2:
The patent creates a synthetic copy of the conversation text with sensitive data replaced by contextually appropriate alternatives. This copied version maintains the same structure, flow, and meaning as the original, allowing seamless playback that preserves conversation integrity while protecting sensitive information.
Data Source
AI summary
In an approach to improve the protection of sensitive data, embodiments extract, by a signal separator component, a voice from an audio file, and transcribe, by a speech-to-text component, the voice in the audio file into text. Further, embodiments identify, by a natural language processing and identification component, sensitive data in the text, wherein identifying sensitive date comprises parsing the text and identifying the sensitive data contained within the text. Additionally, embodiments identify, by a voice locator component, a synthetic voice that matches the voice from the audio file, replace, by a voice replacer component, the identified sensitive data with synthetic data that is semantically and contextually meaningful, wherein the synthetic data is voiced using the synthetic voice, and output a new audio file with the sensitive data replaced by the synthetic data.


