HOA Audio Encoding with Inter-Frame Virtual Loudspeaker Reuse
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high calculation complexity and low encoding efficiency in higher order ambisonics (HOA) audio encoding due to the lack of consideration of inter-frame spatial correlation between virtual loudspeakers during encoding of virtual loudspeaker and residual signals.
Innovation Solution
Determine the encoding parameter of a current frame based on the encoding parameter of a previous frame when the spatial locations of the virtual loudspeakers match or are adjacent, and include a reuse flag in the bitstream to indicate parameter reuse, thereby improving encoding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If encoding parameters are recalculated for each frame independently, then encoding accuracy is maintained, but calculation complexity increases and encoding efficiency decreases
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing encoding parameters from previous frames. When the current frame's virtual loudspeaker spatial location matches a previous frame's spatial location, the pre-stored encoding parameter is directly reused without recalculation. This resolves the contradiction by maintaining encoding accuracy through parameter reuse while dramatically reducing calculation complexity.
2Productivity
If encoding parameters are reused from previous frames, then calculation complexity is reduced and encoding efficiency is improved, but encoding accuracy may deteriorate
Solution Approach 1:
The patent applies local quality by implementing frame-specific parameter selection: for frames where virtual loudspeaker spatial locations match previous frames, encoding parameters are reused from those previous frames; for frames where spatial locations differ, fresh encoding parameters are calculated. This localized approach ensures encoding accuracy is maintained where needed while improving efficiency where possible.
3Reliability
If detailed spatial information is recorded for high-quality three-dimensional audio, then auditory effect is improved, but data amount increases causing transmission and storage difficulties
Solution Approach 1:
The patent applies copying by reusing encoding parameters from previous frames when spatial locations match, rather than storing and transmitting complete encoding parameter sets for every frame. This creates a compressed representation where only necessary parameter updates are transmitted, significantly reducing data amount while maintaining the detailed spatial information needed for high-quality three-dimensional audio reproduction.
Data Source
Figure 1A~1B
Figure 1C~2A
Figure 2B~3A
AI summary
An audio encoding method and apparatus and an audio decoding method and apparatus are disclosed. During encoding of an audio channel signal of a current frame, whether a first target virtual loudspeaker and a second target virtual loudspeaker corresponding to an audio channel signal of a previous frame of the current frame meet a specified condition is first determined. When the first target virtual loudspeaker and the second target virtual loudspeaker meet the specified condition, a first encoding parameter of the audio channel signal of the current frame is determined based on a second encoding parameter of the audio channel signal of the previous frame, so that the audio channel signal of the current frame is encoded based on the first encoding parameter to obtain an encoding result, and the encoding result is written into a bitstream.