Audio Encoding Time-Frequency Tile Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio encoding technologies face limitations in scalability, flexibility, and audio quality, particularly in handling varying speaker setups and bitrates, leading to suboptimal performance in reconstructing inter-object coherence and maintaining transparency during encoding and decoding.

Innovation Solution

A decoder and encoder system that utilizes time-frequency tiles, distinguishing between downmix and non-downmix tiles, and employs upmixing and parametric data to adapt encoding and decoding operations, allowing for flexible and efficient encoding and decoding across different data rates and speaker configurations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If downmix encoding is used for all time-frequency tiles, then coding efficiency is improved, but audio quality and inter-object coherence deteriorate at higher data rates

Engineering Contradiction:
Improvecoding efficiencyVSAvoidaudio quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent applies different encoding strategies to different time-frequency tiles based on their local characteristics. Some tiles are encoded as downmix tiles (combining multiple audio objects) while others are encoded as non-downmix tiles (preserving individual object integrity). This local differentiation allows the system to optimize coding efficiency in regions where it matters while preserving audio quality in regions where it matters most, thereby resolving the contradiction between overall coding efficiency and local audio quality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The encoding approach dynamically adapts to the characteristics of each time-frequency tile and the available bitrate. The system can flexibly switch between downmix and non-downmix encoding for different tiles, and can adjust the proportion of each type based on the coding situation and reproduction requirements. This dynamic adaptability allows the system to maintain optimal audio quality across varying data rates while preserving coding efficiency.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If parametric encoding is used, then scalability to different bitrates is improved, but transparency and inter-object coherence are lost

Engineering Contradiction:
ImprovescalabilityVSAvoidtransparency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a composite encoding structure that combines both parametric downmix tiles and waveform non-downmix tiles within the same audio bitstream. This composite approach allows the system to leverage the scalability benefits of parametric encoding while simultaneously preserving the transparency and inter-object coherence benefits of waveform encoding in critical regions. The combination of different encoding types within a unified framework resolves the contradiction between scalability and transparency.

Inventive Principle:
Principle #40Composite materials

3Manufacturing precision

If individual audio objects are encoded separately, then inter-object coherence is maintained, but coding efficiency decreases

Engineering Contradiction:
Improveinter-object coherenceVSAvoidcoding efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent merges multiple audio objects into downmix tiles for time-frequency regions where the objects exhibit strong correlations or where individual encoding would be inefficient. This merging reduces the overall coding complexity and improves coding efficiency. Simultaneously, the patent preserves separate non-downmix tiles for regions where individual object encoding is necessary to maintain inter-object coherence, thereby resolving the contradiction between coding efficiency and inter-object coherence through selective merging.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP2870603B1Encoding and decoding of audio signals
Publication Date: 2020.09.30 KONINKLIJKE PHILIPS NV
  • EP2870603B1 patent drawingFigure 1
  • EP2870603B1 patent drawingFigure 2
  • EP2870603B1 patent drawingFigure 3

AI summary

An encoder (1201) for encoding a pluralityofaudio signals comprises a selector (1303) which selects a subset of time-frequency tiles to be downmixed and a subset oftiles to be non-downmix. A downmix indication is generated which indicates whether tiles are encoded as downmixed encoded tiles or as non-downmixtiles. An encoded signal comprising the encoded tiles and the downmix indication is fed to a decoder (1203) which includes a receiver (1401) for receiving the signal. A generator (1403) generates output signals from the encoded time-frequency tiles where the generation of the output signals includes an upmixing for tiles that are indicated by the downmix indication to be encoded downmixed tiles. The invention may provide more flexible and/or improved encoding/ decoding and may specifically provide improved scalability, especially at higher data rates.