Audio Processing Apparatus for Mono-Channel Speech Intelligibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio processing technologies fail to effectively improve speech intelligibility for mono-channel device users and listeners with multiple simultaneous audio signals, especially when original signals lack spatial cues or originate from similar positions, leading to energetic and informational masking effects.
Innovation Solution
The proposed solution involves an audio processing apparatus and method that employs spectral separation, spatial separation, and temporal separation techniques, including spectral filtering, spatialization filtering, time scaling, and time delaying, to minimize energetic and informational masking effects, thereby enhancing speech intelligibility. This is achieved through the use of a system comprising spectral filters, spatialization filters, time scaling units, and delayers, which can be combined based on conditions such as the number of speech signals, similarity between speakers, and importance of audio signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple audio signals are transmitted through mono-channel devices, then communication coverage is expanded, but speech intelligibility deteriorates due to energetic and informational masking effects
Solution Approach 1:
The audio spectrum is segmented into multiple frequency bands using spectral filtering. Different speech signals are assigned to different frequency bands, allowing mono-channel devices to separate and identify multiple speakers by their spectral characteristics, thereby maintaining speech intelligibility while expanding communication coverage
Solution Approach 2:
The patent introduces temporal dimension by applying time scaling and time delaying to speech signals. This creates temporal separation between overlapping speech signals in the time domain, enabling listeners to distinguish between simultaneous speakers even when they occupy the same frequency spectrum in mono-channel transmission
2Loss of information
If spectral filtering is applied to separate speech signals, then speech intelligibility is improved, but audio signal quality may be degraded
Solution Approach 1:
Different frequency bands are allocated to different speech signals based on their importance and characteristics. Critical speech components are preserved in wider bandwidths while less critical components use narrower bandwidths, optimizing both intelligibility and overall audio quality
Solution Approach 2:
The system dynamically adjusts filtering parameters, time scaling factors, and time delays based on real-time analysis of speech signals. This allows the audio processing apparatus to adapt to varying speech conditions, maintaining high quality while improving intelligibility in different scenarios
3Loss of information
If time scaling and time delaying are applied to speech signals, then temporal separation is achieved, but processing complexity increases
Solution Approach 1:
Time scaling and time delaying operations are applied to speech signals before mixing them in the mono-channel output. This preliminary temporal processing creates separable temporal patterns that simplify subsequent listening and processing, reducing the effective complexity of separating mixed speech signals
Data Source
Figure 1~2
Figure 3~5
Figure 6(a)~8
AI summary
An audio processing method and apparatus are described. In one embodiment, at least one first sub-band of a first audio signal is suppressed to obtain a reduced first audio signal with reserved sub-bands; suppressing at least one second sub-band of the at least one second audio signal to obtain at least one reduced second audio signal with reserved sub-bands; and mixing the reduced first audio signal and the at least one reduced second audio signal. Alternatively, a first spatial auditory property is assigned to a first audio signal so that the first audio signal may be perceived as originating from a first position. Alternatively, rhythmic similarity between at least two audio signals is detected, and time scaling is applied to an audio signal in response to relatively high rhythmic similarity between the audio signal and the other audio signal(s); and then the at least two audio signals are mixed.