Zone Audio Speech Masking Using Spectral Band Swapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication technologies fail to effectively prevent unwanted eavesdropping in public spaces without causing unpleasant disturbances, such as increased noise levels, particularly in vehicles and public transport, where confidential conversations can be overheard due to zone-based audio systems.
Innovation Solution
A method and device for generating a masking signal in a zone-based audio system that alters the spectral structure of a speech signal by swapping spectral bands and adding a broadband noise signal, combined with spatial reproduction and distraction signals at specific speech onsets, to reduce speech intelligibility without significantly increasing overall sound levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If loud noise is played to prevent eavesdropping, then speech intelligibility for eavesdroppers is reduced, but overall noise level increases causing unpleasant disturbance
Solution Approach 1:
The patent applies local quality by generating a masking signal with specific spectral characteristics that are locally adapted to the speech signal. Instead of using uniform loud noise, the system creates a masking signal with modified spectral distribution (through techniques like spectral swapping or band-limited noise generation) that targets specific frequency ranges where speech energy is concentrated, thereby achieving effective masking with reduced overall noise level and minimal disturbance to listeners.
2Loss of information
If spectral bands are swapped to generate masking signal, then speech intelligibility is reduced without changing overall energy, but the masking signal does not perfectly match the speech spectrum
Solution Approach 1:
The patent applies parameter changes by systematically modifying the spectral parameters of the masking signal. Techniques include spectral swapping where energy distribution across frequency bands is rearranged, or adjusting the spectral density parameters of generated noise to match the speech signal's spectral envelope. This allows the masking signal to closely resemble the speech spectrum in terms of overall energy distribution while maintaining sufficient difference to prevent intelligibility, resolving the contradiction between spectral matching accuracy and masking effectiveness.
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
The present disclosure relates to a method for masking a speech signal in a zone-based audio system, comprising: capturing a speech signal to be masked in one audio zone; transforming the captured speech signal into spectral bands; swapping spectral values of at least two spectral bands; generating a noise signal based on the swapped spectral values; and outputting the noise signal as a masking signal for the speech signal in another audio zone.