Audio Encoding Time-Frequency Tile Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding technologies face limitations in scalability, flexibility, and audio quality, particularly in handling varying speaker setups and bitrates, leading to suboptimal performance in reconstructing inter-object coherence and maintaining transparency during encoding and decoding.
Innovation Solution
A decoder and encoder system that utilizes time-frequency tiles, distinguishing between downmix and non-downmix tiles, and employs upmixing and parametric data to adapt encoding and decoding operations, allowing for flexible and efficient encoding and decoding across different data rates and speaker configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If downmix encoding is used for all time-frequency tiles, then coding efficiency is improved, but audio quality and inter-object coherence deteriorate at higher data rates
Solution Approach 1:
The patent applies different encoding strategies to different time-frequency tiles based on their local characteristics. Some tiles are encoded as downmix tiles (combining multiple audio objects) while others are encoded as non-downmix tiles (preserving individual object integrity). This local differentiation allows the system to optimize coding efficiency in regions where it matters while preserving audio quality in regions where it matters most, thereby resolving the contradiction between overall coding efficiency and local audio quality.
Solution Approach 2:
The encoding approach dynamically adapts to the characteristics of each time-frequency tile and the available bitrate. The system can flexibly switch between downmix and non-downmix encoding for different tiles, and can adjust the proportion of each type based on the coding situation and reproduction requirements. This dynamic adaptability allows the system to maintain optimal audio quality across varying data rates while preserving coding efficiency.
2Adaptability or versatility
If parametric encoding is used, then scalability to different bitrates is improved, but transparency and inter-object coherence are lost
Solution Approach 1:
The patent creates a composite encoding structure that combines both parametric downmix tiles and waveform non-downmix tiles within the same audio bitstream. This composite approach allows the system to leverage the scalability benefits of parametric encoding while simultaneously preserving the transparency and inter-object coherence benefits of waveform encoding in critical regions. The combination of different encoding types within a unified framework resolves the contradiction between scalability and transparency.
3Manufacturing precision
If individual audio objects are encoded separately, then inter-object coherence is maintained, but coding efficiency decreases
Solution Approach 1:
The patent merges multiple audio objects into downmix tiles for time-frequency regions where the objects exhibit strong correlations or where individual encoding would be inefficient. This merging reduces the overall coding complexity and improves coding efficiency. Simultaneously, the patent preserves separate non-downmix tiles for regions where individual object encoding is necessary to maintain inter-object coherence, thereby resolving the contradiction between coding efficiency and inter-object coherence through selective merging.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An encoder (1201) for encoding a pluralityofaudio signals comprises a selector (1303) which selects a subset of time-frequency tiles to be downmixed and a subset oftiles to be non-downmix. A downmix indication is generated which indicates whether tiles are encoded as downmixed encoded tiles or as non-downmixtiles. An encoded signal comprising the encoded tiles and the downmix indication is fed to a decoder (1203) which includes a receiver (1401) for receiving the signal. A generator (1403) generates output signals from the encoded time-frequency tiles where the generation of the output signals includes an upmixing for tiles that are indicated by the downmix indication to be encoded downmixed tiles. The invention may provide more flexible and/or improved encoding/ decoding and may specifically provide improved scalability, especially at higher data rates.