Audio Upmixer Speech Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing techniques fail to accurately convert stereo audio signals into multi-channel signals without introducing speech components into the back channels, which can disrupt the listener's experience by incorrectly localizing speech sources.
Innovation Solution
A device and method that includes a speech detector and signal modifier to identify and suppress speech components in the ambience channel, ensuring that speech is only present in the front channels, using an upmixer to generate direct and ambience channels while maintaining the original tone-image experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If stereo signals are upmixed to multi-channel signals using conventional techniques, then all loudspeakers including back loudspeakers are utilized, but speech components are incorrectly distributed to back channels causing localization errors
Solution Approach 1:
The patent segments the audio signal processing into distinct functional blocks: upmixer for channel expansion, speech detector for identifying speech portions, and ambience signal modifier for selective processing. This segmentation allows different processing strategies to be applied to different signal components, ensuring speech remains in front channels while ambience fills back channels.
Solution Approach 2:
The speech detector performs preliminary identification of speech portions in the upmixed signal before the ambience signal modifier processes the back channels. This preliminary action ensures that speech components are identified and protected from being redirected to back loudspeakers, maintaining correct localization before final signal distribution.
2Ease of operation
If speech components are present in ambience channels, then the audio scene delving experience is improved, but speech localization is disrupted and listener experience is disturbed
Solution Approach 1:
The patent extracts speech portions from the ambience channel signal using the speech detector and removes them before the ambience signal modifier processes the back channels. This extraction ensures that only non-speech ambience components are present in back channels, maintaining immersion without localization disruption.
Solution Approach 2:
The patent applies different quality characteristics to different spatial zones: front channels receive full-bandwidth signals containing speech for accurate localization, while back channels receive modified ambience signals with speech components removed for immersive atmosphere without localization interference.
3Device complexity
If the upmixing process distributes all signal components uniformly, then the reproduction is simple, but speech sources are incorrectly localized in back channels
Solution Approach 1:
The patent introduces dynamic processing where the ambience signal modifier adjusts the back channel signals based on real-time detection of speech portions. The processing is adaptive rather than static, modifying the ambience signal dynamically to suppress speech components only when and where they appear, maintaining localization accuracy without excessive complexity.
Data Source
AI summary
In order to generate a multi-channel signal having a number of output channels greater than a number of input channels, a mixer is used for upmixing the input signal to form at least a direct channel signal and at least an ambience channel signal. A speech detector is provided for detecting a section of the input signal, the direct channel signal or the ambience channel signal in which speech portions occur. Based on this detection, a signal modifier modifies the input signal or the ambience channel signal in order to attenuate speech portions in the ambience channel signal, whereas such speech portions in the direct channel signal are attenuated to a lesser extent or not at all. A loudspeaker signal outputter then maps the direct channel signals and the ambience channel signals to loudspeaker signals which are associated to a defined reproduction scheme, such as, for example, a 5.1 scheme.


