3D Audio Encoding with Temporal Virtual Loudspeaker Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing three-dimensional audio signal encoding methods suffer from unstable spatial images and reduced sound quality due to frequent changes in virtual loudspeaker selection between frames, leading to discontinuity and noise in reconstructed audio signals.

Innovation Solution

A method that adjusts virtual loudspeaker selection by retaining previous-frame representative virtual loudspeakers and updating current-frame vote values based on previous-frame values, using adjustment parameters to enhance directional continuity and reduce frequent changes, while selecting a reduced number of representative virtual loudspeakers for encoding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the encoder traverses virtual loudspeakers and selects based on current frame only, then the encoding adaptability is improved, but the spatial image stability deteriorates due to frequent changes between frames

Engineering Contradiction:
Improveencoding adaptabilityVSAvoidspatial image stability
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The encoder performs preliminary action by selecting representative virtual loudspeakers from the previous frame and using them as initial candidates for the current frame. This preliminary selection based on temporal continuity reduces the search space and prevents frequent changes, thereby maintaining spatial image stability while preserving encoding adaptability through subsequent refinement steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The method ensures continuity of useful action by maintaining temporal coherence in virtual loudspeaker selection across frames. The representative virtual loudspeakers selected in the previous frame are carried forward and combined with current frame information, creating a continuous selection process that stabilizes the spatial image while adapting to changing audio content.

Inventive Principle:
Principle #20Continuity of useful action

2Manufacturing precision

If the encoder selects more virtual loudspeakers for each frame, then the sound quality is improved, but the calculation complexity increases

Engineering Contradiction:
Improvesound qualityVSAvoidcalculation complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The encoder extracts only the most representative virtual loudspeakers from the candidate set based on vote values and temporal continuity criteria. By selecting a limited number of representative virtual loudspeakers rather than processing all candidates, the method maintains sound quality while significantly reducing calculation complexity in the subsequent encoding stages.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The method applies local quality by differentiating the treatment of virtual loudspeakers based on their representativeness and temporal stability. Representative virtual loudspeakers that show continuity across frames are prioritized and given higher weight, while less representative candidates are processed with reduced complexity, optimizing the overall encoding efficiency.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP4325485B1Three-dimensional audio signal encoding method and encoder
Publication Date: 2026.02.25 HUAWEI TECH CO LTD
  • EP4325485B1 patent drawingFigure 1
  • EP4325485B1 patent drawingFigure 2(a)~2(b)
  • EP4325485B1 patent drawingFigure 3

AI summary

A three-dimensional audio signal encoding method and apparatus, and an encoder (113) are provided, and relate to the multimedia field. The method includes: The encoder (113) obtains a first quantity of current-frame initial vote values for a current frame of a three-dimensional audio signal (S610). Then, the encoder (113) obtains, based on the first quantity of current-frame initial vote values and a sixth quantity of previous-frame final vote values, a seventh quantity of current-frame final vote values that are of a seventh quantity of virtual loudspeakers and that correspond to the current frame (S620). Further, the encoder (113) selects a second quantity of current-frame representative virtual loudspeakers from the seventh quantity of virtual loudspeakers based on the seventh quantity of current-frame final vote values (S630). The encoder (113) encodes the current frame based on the second quantity of current-frame representative virtual loudspeakers, to obtain a bitstream (S640). In this way, signal directional continuity between frames is enhanced, stability of a spatial image of the reconstructed three-dimensional audio signal is improved, and sound quality of the reconstructed three-dimensional audio signal is ensured.