Diffuse Sound Shaping for BCC Audio Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding technologies face challenges in accurately synthesizing complex auditory scenes with multiple audio sources, as they struggle to preserve the temporal envelope and spatial cues of audio signals, leading to artifacts such as pre-echoes and washed-out transients.
Innovation Solution
The method involves generating cue codes for multiple input audio channels, downmixing them to create fewer transmitted channels, and using envelope shaping to ensure that the synthesized audio signal matches the original temporal envelope, while applying inter-channel time and level differences to recreate the spatial cues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional binaural signal synthesis is used to convert mono audio to binaural signals, then spatial cues can be created, but temporal envelope characteristics are degraded causing artifacts
Solution Approach 1:
The audio signal is divided into multiple frequency bands using filter banks before processing. Each frequency band is processed independently through correlation processing to generate binaural cues, while the temporal envelope is preserved through envelope shaping. This segmentation allows spatial manipulation without degrading temporal characteristics.
Solution Approach 2:
The patent applies envelope shaping that modifies the temporal envelope parameters of the audio signal to match the original characteristics. By adjusting envelope parameters (amplitude modulation characteristics) after correlation processing, the system restores accurate temporal envelope while maintaining the spatial cues generated during processing.
2Adaptability or versatility
If audio signals are processed to generate binaural cues with spatial information, then auditory scene synthesis is improved, but signal correlation is increased causing washed-out transients
Solution Approach 1:
The patent extracts the temporal envelope information from the audio signal and uses it separately for envelope shaping. By separating the envelope extraction from the correlation processing, the system can apply spatial manipulation through correlation while using the extracted envelope to restore temporal characteristics, preventing washed-out transients.
Solution Approach 2:
The system uses envelope shaping as a feedback mechanism where the temporal envelope characteristics of the processed signal are adjusted to match the original signal's envelope. This feedback loop compensates for the correlation-induced degradation of temporal characteristics, maintaining signal quality while preserving spatial cues.
3Manufacturing precision
If multiple audio channels are processed to preserve spatial cues, then audio quality is improved, but processing complexity increases
Solution Approach 1:
The patent combines multiple processing operations (correlation processing for spatial cues and envelope shaping for temporal characteristics) into a unified audio processing system. By merging these operations and using shared resources like filter banks and processing blocks, the system achieves high audio quality while managing processing complexity through efficient integration.
Data Source
AI summary
An input audio signal having an input temporal envelope is converted into an output audio signal having an output temporal envelope. The input temporal envelope of the input audio signal is characterized. The input audio signal is processed to generate a processed audio signal, wherein the processing de-correlates the input audio signal. The processed audio signal is adjusted based on the characterized input temporal envelope to generate the output audio signal, wherein the output temporal envelope substantially matches the input temporal envelope.


