Speech Privacy Protection via Phoneme-Structured Interference Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech interference methods, such as using white noise or ultrasonic waves, are not effective in preventing speech privacy leakage due to their simplicity and vulnerability to removal by denoising algorithms.
Innovation Solution
A method for designing an interference noise based on the human speech structure, involving the extraction of voiceprint information, data augmentation, phoneme-level segmentation, and the construction of noise sequences to generate a robust interference noise that is difficult for both human and machine-based recognition systems to decode.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If white noise is used to interfere with recordings, then the microphone recording is disrupted, but the speech privacy cannot be effectively protected because denoising algorithms can remove the noise
Solution Approach 1:
The patent uses composite noise signals that combine multiple characteristics (ultrasonic waves, modulated signals, and speech-like patterns) to create an interference noise that is both disruptive to recordings and resistant to denoising algorithms. This composite approach addresses the simplicity issue while maintaining privacy protection effectiveness.
Solution Approach 2:
The patent dynamically changes noise parameters including frequency, amplitude, and temporal characteristics to adapt to different recording environments and denoising algorithms. By continuously adjusting these parameters, the system maintains effective privacy protection while avoiding the simplicity of static white noise.
2Reliability
If ultrasonic waves are injected to interfere with eavesdropping devices, then unauthorized recordings are disrupted, but the noise pattern is too simple and can be removed by denoising algorithms
Solution Approach 1:
The patent segments the interference noise into multiple independent components (ultrasonic carrier waves, modulated signals, and speech-like patterns), each serving a specific function. This segmentation increases overall complexity while maintaining interference effectiveness against unauthorized recordings.
Solution Approach 2:
The patent embeds multiple noise patterns within each other, with speech-like patterns nested within modulated signals, which are themselves nested within ultrasonic carrier waves. This nested structure creates a complex noise profile that is difficult for denoising algorithms to separate and remove.
3Ease of manufacture
If simple noise patterns are used for interference, then the system is easier to implement, but existing denoising algorithms can effectively remove the noise and speech privacy leaks occur
Solution Approach 1:
The patent implements dynamic noise generation where the interference pattern continuously adapts based on the recording environment and target speech characteristics. This dynamic approach maintains ease of implementation through automated parameter adjustment while preventing privacy leakage by avoiding simple, predictable noise patterns.
Solution Approach 2:
The system incorporates feedback mechanisms that monitor the effectiveness of noise interference and adjust parameters accordingly. This feedback loop ensures that the noise remains complex and effective at preventing privacy leakage while maintaining reasonable implementation simplicity through automated control.
Data Source
AI summary
The present invention discloses a method for designing an interference noise of speech based on the human speech structure, including the following steps: (1): obtaining a large amount of speech data containing different speakers and different speech contents, extracting voiceprint information, and then building an initial speech data set; (2): for each user, obtaining a small amount of speech data of the user, extracting voiceprint information, and then matching the most similar speech data in the initial speech data set; (3): performing data augmentation on the matched speech data; (4): segmenting the augmented speech data with a phoneme segmentation algorithm to form a vowel data set and a consonant data set; (5): constructing three noise sequences based on the vowel data set and the consonant data set, and performing superimposition to obtain an interference noise; and (6): continuously generating and playing randomly generated interference noise, and continuously injecting the interference noise into recordings to implement continuous interference. With the present invention, the interference noise cannot be removed from the speech, thereby avoiding the leakage of user privacy information.


