Spatial Audio Mixing Preserving Timbre for Impulsive Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio mixing systems face challenges in effectively generating and mixing volumetric virtual sound sources, particularly with signals that have problematic attributes such as impulsiveness and peakiness, which can lead to undesirable changes in timbre and intelligibility.
Innovation Solution
The system analyzes input audio signals for problematic attributes using analyzers for peakiness, impulsiveness, and speech content. Based on these analyses, it adjusts parameters associated with spatial extent processing and mixes the spatially extended audio signal with the original signal to preserve timbre and maintain natural sound perception.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Volume of moving object
If spatial extent processing is applied to extend the spatial volume of sound sources, then the perceived spatial extent is improved, but the timbre and intelligibility of impulsive and peaky sounds deteriorate
Solution Approach 1:
The system dynamically changes processing parameters based on the detected attributes of the input signal. When impulsive or peaky characteristics are detected, the system adjusts the spatial extent processing parameters to reduce their intensity or apply different processing strategies, thereby preserving timbre accuracy while still providing spatial extension for suitable signals.
Solution Approach 2:
The system employs feedback mechanisms where the output of attribute detection (impulsiveness, peakiness) feeds back into the spatial extent processing stage. This closed-loop control allows the system to adaptively adjust processing strength and methods based on real-time signal characteristics, preventing timbre degradation while maintaining spatial extension benefits.
2Reliability
If spatial extent processing is applied to create volumetric virtual sound sources, then the immersive quality is improved, but the natural perception of impulsive sounds is degraded
Solution Approach 1:
The system applies different processing qualities to different parts of the audio signal based on local characteristics. Impulsive and peaky portions are identified and processed differently from sustained or tonal portions, allowing spatial extension to be applied locally where appropriate while preserving the natural impact of impulsive sounds.
Solution Approach 2:
The spatial extent processing is made dynamic rather than static, continuously adapting to the temporal and spectral characteristics of the input signal. This allows the system to provide immersive spatial extension during suitable passages while automatically reducing or modifying processing during impulsive events to maintain natural perception.
3Volume of moving object
If frequency bands are distributed to different spatial directions to create spatial extent, then the spatial volume is improved, but the coherence and unity of the sound source are reduced
Solution Approach 1:
The system segments the audio signal into frequency bands and processes each band independently for spatial distribution. This segmentation allows precise control over how different frequency components are spatially distributed, enabling the creation of spatial volume while maintaining overall coherence through coordinated processing of all segments.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus for generating at least one audio signal associated with a sound scene is configured to receive at least one audio signal and analyse 101 the audio signal(s) to determine at least one attribute parameter such as peakiness, impulsiveness and voice activity. At least one control signal is determined based on the at least one attribute and is then used to generate a spatially extended audio signal from the at least one audio signal. The initial audio signal(s) and the spatially extended audio signal(s) are then mixed 121 in a proportion based on the control signal, to generate at least one audio signal associated with the sound scene where the timbre of the original audio signal is preserved. The spatial extension may comprise applying at least one of: a vector base amplitude panning; a direct binaural panning, a direct assignment to channel output location; synthesized ambisonics; or wavefield synthesis. The technique is particularly suited to synthesis of sound objects containing percussive or impulsive sounds, and speech. The apparatus may be used in virtual reality or computer game applications.