Spatial Audio Parameter Set Reduction Across Time-Frequency Tiles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio encoding methods require a high bit rate for transmitting spatial metadata, especially when multiple directions are supported in a time-frequency tile, which is inefficient for immersive audio applications.
Innovation Solution
The method reduces the number of spatial audio parameter sets across time frames, frequency sub-bands, and sound source directions by analyzing and comparing parameter sets for similarity, and selecting or merging them to meet a target coding rate.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple spatial audio parameter sets are transmitted for each time-frequency tile to support multiple sound source directions, then the accuracy of spatial audio representation is improved, but the bit rate required for transmission increases significantly
Solution Approach 1:
The patent merges multiple spatial audio parameter sets by identifying and combining parameter sets that represent similar sound source directions. The decoder combines parameter sets based on similarity criteria (such as direction vectors and energy ratios) to reconstruct spatial audio parameters, thereby reducing the number of parameter sets that need to be transmitted while maintaining spatial audio accuracy
Solution Approach 2:
The patent changes the parameter representation by introducing similarity metrics and combination rules for spatial audio parameters. Instead of transmitting all parameter sets independently, the system uses parameter similarity (based on direction vectors, energy ratios, and coherence values) to determine which parameter sets can be merged or represented more efficiently, reducing the overall parameter transmission requirement
2Productivity
If the number of spatial audio parameter sets is reduced to lower bit rate, then transmission efficiency is improved, but the ability to represent multiple sound source directions may be compromised
Solution Approach 1:
The patent implements a feedback mechanism where the decoder receives information about parameter set similarities and uses this feedback to intelligently combine parameter sets. The system adjusts the number and representation of parameter sets based on the actual spatial audio content, maintaining multi-direction support where needed while reducing transmission for redundant parameter sets
Solution Approach 2:
The patent makes the spatial audio parameter representation dynamic by allowing the number and detail of parameter sets to vary based on the complexity of the spatial audio scene. The system adaptively determines how many parameter sets to transmit and combine based on similarity thresholds and spatial distribution, optimizing between transmission efficiency and multi-direction representation capability
Data Source
AI summary
There is inter alia disclosed an apparatus for spatial audio encoding comprising: means for analysing a plurality of spatial audio parameter sets associated with a frame of one or more audio signals, wherein the plurality of spatial audio parameter sets are associated with a plurality of subframes, a plurality of frequency sub bands and a plurality of sound source directions for the frame of the one or more audio signals; and means for determining from the analysis of the plurality of spatial audio parameter sets at least one spatial audio parameter set for subframes of the frame of the one or more audio signals.


