Speech Data Augmentation With Sensitive Content Obscuring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Storing audio encounter information and corresponding text transcriptions for training speech processing systems presents security concerns, as sensitive content can be breached, and conventional data augmentation methods either degrade system performance or fail to protect sensitive information.
Innovation Solution
Generate augmented speech signals that retain acoustic properties while removing sensitive content using text-to-speech and voice style transfer processing, allowing secure data augmentation without storing sensitive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If audio encounter information and text transcriptions are stored for training speech processing systems, then system training data availability is improved, but security risk increases due to sensitive content exposure
Solution Approach 1:
The patent extracts and removes sensitive content (names, dates, locations, and other personally identifiable information) from audio transcriptions while preserving the acoustic properties and linguistic structure. This allows training data to be generated without storing sensitive information, resolving the contradiction between data availability and security risk.
Solution Approach 2:
The system creates synthetic audio data that copies the acoustic characteristics, speech patterns, and linguistic structure of real audio encounters without copying sensitive content. This synthetic data can be used for training while the original sensitive data is either removed or anonymized, maintaining training quality without the security risks of storing sensitive information.
2Quantity of substance
If conventional data augmentation methods are used to generate training data, then data quantity is improved, but system performance degrades due to loss of acoustic properties
Solution Approach 1:
The patent applies parameter changes by modifying audio signals through acoustic property transformations (adding noise, reverberation, pitch variations, speed adjustments) while preserving the underlying linguistic content. This generates diverse training data that maintains acoustic realism, resolving the contradiction between data quantity and system performance.
Solution Approach 2:
The system maintains the continuity of useful action by preserving the acoustic properties and linguistic structure of speech signals throughout the data augmentation process. Unlike methods that destroy acoustic information, this approach continuously retains useful acoustic characteristics while generating varied training examples, ensuring both data quantity and performance improvement.
3Object-affected harmful factors
If audio features are extracted and encrypted to protect sensitive content, then security is improved, but adaptability of the training data is reduced due to feature extraction pipeline restrictions
Solution Approach 1:
The patent segments the audio processing into separate functional components: sensitive content identification, acoustic property extraction, and data augmentation. This segmentation allows the system to protect sensitive information while maintaining adaptability, as the acoustic features can be applied to various training scenarios without being constrained by a fixed feature extraction pipeline.
Solution Approach 2:
The system achieves universality by creating a multi-functional approach that can handle different data scenarios (real audio with sensitive content, synthetic audio, audio with acoustic properties) and apply the same security and augmentation framework. This resolves the contradiction by making the system adaptable to various training needs while maintaining security through consistent sensitive content removal.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method, computer program product, and computing system for receiving an input speech signal. A transcription of the input speech signal may be received. A speaker embedding may be extracted from the input speech signal. Acoustic properties from the input speech signal may be extracted. An obscured transcription may be generated from the transcription, where the obscured transcription includes obscured representations of sensitive content from the transcription. An obscured speech signal may be generated based upon, at least in part, the extracted speaker embedding and the obscured transcription, where the obscured speech signal includes obscured representations of sensitive content from the input speech signal. The obscured speech signal may be augmented based upon, at least in part, the extracted acoustic properties.