Audio Upmixer Speech Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing techniques fail to accurately convert stereo audio signals into multi-channel signals without introducing speech components into the back channels, which can disrupt the listener's experience by incorrectly localizing speech sources.

Innovation Solution

A device and method that includes a speech detector and signal modifier to identify and suppress speech components in the ambience channel, ensuring that speech is only present in the front channels, using an upmixer to generate direct and ambience channels while maintaining the original tone-image experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If stereo signals are upmixed to multi-channel signals using conventional techniques, then all loudspeakers including back loudspeakers are utilized, but speech components are incorrectly distributed to back channels causing localization errors

Engineering Contradiction:
Improvemulti-channel utilizationVSAvoidspeech localization accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the audio signal processing into distinct functional blocks: upmixer for channel expansion, speech detector for identifying speech portions, and ambience signal modifier for selective processing. This segmentation allows different processing strategies to be applied to different signal components, ensuring speech remains in front channels while ambience fills back channels.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The speech detector performs preliminary identification of speech portions in the upmixed signal before the ambience signal modifier processes the back channels. This preliminary action ensures that speech components are identified and protected from being redirected to back loudspeakers, maintaining correct localization before final signal distribution.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If speech components are present in ambience channels, then the audio scene delving experience is improved, but speech localization is disrupted and listener experience is disturbed

Engineering Contradiction:
Improveaudio scene immersionVSAvoidspeech localization disruption
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent extracts speech portions from the ambience channel signal using the speech detector and removes them before the ambience signal modifier processes the back channels. This extraction ensures that only non-speech ambience components are present in back channels, maintaining immersion without localization disruption.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different quality characteristics to different spatial zones: front channels receive full-bandwidth signals containing speech for accurate localization, while back channels receive modified ambience signals with speech components removed for immersive atmosphere without localization interference.

Inventive Principle:
Principle #3Local quality

3Device complexity

If the upmixing process distributes all signal components uniformly, then the reproduction is simple, but speech sources are incorrectly localized in back channels

Engineering Contradiction:
Improvesignal processing complexityVSAvoidsound source localization
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent introduces dynamic processing where the ambience signal modifier adjusts the back channel signals based on real-time detection of speech portions. The processing is adaptive rather than static, modifying the ambience signal dynamically to suppress speech components only when and where they appear, maintaining localization accuracy without excessive complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8731209B2Device and method for generating a multi-channel signal including speech signal processing
Publication Date: 2014.05.20 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US8731209B2 patent drawing
  • US8731209B2 patent drawing
  • US8731209B2 patent drawing

AI summary

In order to generate a multi-channel signal having a number of output channels greater than a number of input channels, a mixer is used for upmixing the input signal to form at least a direct channel signal and at least an ambience channel signal. A speech detector is provided for detecting a section of the input signal, the direct channel signal or the ambience channel signal in which speech portions occur. Based on this detection, a signal modifier modifies the input signal or the ambience channel signal in order to attenuate speech portions in the ambience channel signal, whereas such speech portions in the direct channel signal are attenuated to a lesser extent or not at all. A loudspeaker signal outputter then maps the direct channel signals and the ambience channel signals to loudspeaker signals which are associated to a defined reproduction scheme, such as, for example, a 5.1 scheme.