Spatial Audio Processing for Speech Intelligibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Stereo widening and other spatial processing techniques in sound reproduction devices often decrease the intelligibility of speech and key sounds due to distortions, as they increase the perceived width of audio scenes but compromise clarity.
Innovation Solution
An apparatus and method that utilize a trained machine learning model to separate audio signals into a first portion comprising speech or key sounds and a second portion comprising ambient sounds, processing the first portion without spatial audio processing and the second portion with spatial audio processing, while applying different equalization processes to maintain intelligibility and retain spatial effects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If spatial audio processing (stereo widening) is applied to increase perceived width of audio scene, then spatial perception is improved, but intelligibility of speech and key sounds deteriorates due to distortions
Solution Approach 1:
The audio signal is segmented into two distinct portions: a first portion containing speech and key sounds, and a second portion containing ambient sounds and effects. This segmentation allows different processing strategies to be applied to each portion, preserving intelligibility in the first portion while maintaining spatial effects in the second portion.
Solution Approach 2:
Different quality levels of spatial processing are applied to different parts of the audio signal. The first portion (speech/key sounds) receives minimal or no spatial processing to maintain high intelligibility, while the second portion (ambient sounds) receives full spatial processing to preserve spatial effects. This local differentiation resolves the contradiction between spatial perception and intelligibility.
2Adaptability or versatility
If stereo widening process is applied to ambient sounds, then spatial audio effects are retained, but overall audio scene distortion increases
Solution Approach 1:
The audio signal is divided into a first portion (speech/key sounds) and a second portion (ambient sounds). By applying spatial processing only to the second portion, the patent retains spatial audio effects where needed while avoiding the application of distortion-causing processing to the first portion, thereby reducing overall harmful distortions in the audio scene.
Solution Approach 2:
Spatial processing is applied locally only to the second portion of the audio signal containing ambient sounds, while the first portion containing speech and key sounds is processed with minimal or no spatial effects. This localized application of spatial processing preserves spatial audio effects in appropriate contexts while minimizing overall distortion in the audio scene.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Examples of the disclosure relate to apparatus, methods and computer programs for spatial processing audio scenes with improved intelligibility for speech or other key sounds. In examples of the disclosure at least one audio signal comprising two or more channels is obtained. The audio signal is processed with program code to identify at least a first portion of the audio signal wherein the first portion predominantly comprises audio of interest. The first portion is processed using a first process. The second portion is processed using a second process comprising spatial audio processing. The first process comprises no spatial audio processing or a low level of spatial audio processing compared to the second process and the second portion predominantly comprises a remainder. The processed first portion and second portion can be played back using two or more loudspeakers.