Scene Audio Encoding With Virtual Speaker Attributes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing three-dimensional audio technologies face challenges in encoding and decoding high-order Ambisonics (HOA) signals due to the large data amount, which complicates transmission and storage, and current methods result in low encoding performance and audio quality.
Innovation Solution
A scene audio encoding method that directly encodes a first audio signal and attribute information of a target virtual speaker, allowing for reconstruction of the scene audio signal without calculating virtual speaker signals and residual signals, thereby reducing data amount and encoding complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If the HOA signal is encoded using conventional methods, then the audio quality is improved, but the encoding complexity and data amount increase
Solution Approach 1:
The patent segments the HOA signal encoding process into two independent parts: a first audio signal with K channels and attribute information of a target virtual speaker. This segmentation allows the encoder to process only the essential information needed for reconstruction, reducing overall encoding complexity while maintaining audio quality through the virtual speaker's spatial attributes.
Solution Approach 2:
The patent extracts only the necessary information from the complete HOA signal by identifying and encoding a target virtual speaker's attributes (position, orientation, radius) and a reduced audio signal. This extraction eliminates redundant data while preserving the core audio quality, directly reducing encoding complexity without sacrificing reconstruction fidelity.
2Manufacturing precision
If the HOA signal is encoded with more channels to improve audio quality, then the audio quality is improved, but the data amount and bit rate increase
Solution Approach 1:
The patent applies local quality by encoding only the essential attributes of the target virtual speaker (position, orientation, radius) and a reduced audio signal with K channels, rather than encoding all C1 channels. This localized encoding approach maintains high audio quality for the virtual speaker's contribution while significantly reducing the overall data amount transmitted.
Solution Approach 2:
The patent changes the parameter representation from full-channel HOA signals to a reduced set of K channels combined with virtual speaker attributes. This parameter transformation allows the system to represent the same audio information with fewer data elements, reducing bit rate while maintaining or improving audio quality through the virtual speaker model.
3Adaptability or versatility
If the scene audio signal is converted into virtual speaker signal and residual signal for encoding, then the encoding flexibility is improved, but the encoding complexity increases
Solution Approach 1:
Instead of converting the HOA signal into virtual speaker signal and residual signal as done in conventional methods, the patent inverts the approach by directly encoding a reduced audio signal with K channels and virtual speaker attributes. This inversion simplifies the encoding process by eliminating the need for complex signal conversion while maintaining encoding flexibility through the virtual speaker model.
Data Source
AI summary
A scene audio encoding method includes obtaining a to-be-encoded scene audio signal, where the scene audio signal includes an audio signal with C1 channels, determining attribute information of a target virtual speaker based on the scene audio signal, and encoding a first audio signal in the scene audio signal and the attribute information to obtain a first bitstream. The first audio signal is an audio signal with K channels in the scene audio signal, and K is a positive integer less than or equal to C1.


