3D Audio Encoding with Representative Virtual Speaker Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high calculation complexity of compressing three-dimensional audio signals using existing encoders hinders efficient data storage and transmission, requiring a solution to reduce computational load while maintaining audio quality.
Innovation Solution
A method and apparatus that select a reduced number of representative virtual speakers based on vote values to encode three-dimensional audio signals, using a candidate virtual speaker set and iterative voting to improve compression efficiency and reduce calculation complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a plurality of preconfigured virtual speakers are used to compress the three-dimensional audio signal, then the compression efficiency is improved, but the calculation complexity increases
Solution Approach 1:
The patent applies partial action by selecting only the top K virtual speakers with the highest vote values from the candidate set, rather than using all preconfigured virtual speakers for compression. This reduces the number of virtual speakers involved in the encoding process from the full set to a subset, thereby decreasing calculation complexity while maintaining acceptable compression efficiency
Solution Approach 2:
The patent changes the parameter of virtual speaker quantity dynamically based on the audio signal characteristics. By adjusting K (the number of selected virtual speakers) according to the vote values obtained from the audio signal, the system adapts the complexity of the encoding process to match the actual compression needs, resolving the contradiction between compression efficiency and calculation complexity
2Measurement precision
If all coefficients of the current frame are used to vote for each virtual speaker, then the selection accuracy is improved, but the calculation load increases
Solution Approach 1:
The patent extracts only the essential information needed for virtual speaker selection by using a reduced set of representative coefficients rather than processing all coefficients from the current frame. This extraction approach maintains the accuracy of virtual speaker selection while significantly reducing the calculation load required to process the audio signal
Data Source
AI summary
This application discloses a three-dimensional audio signal coding method and apparatus, and an encoder, and relates to the multimedia field. The method includes: After determining a first quantity of virtual speakers and a first quantity of vote values based on a current frame of a three-dimensional audio signal, a candidate virtual speaker set, and a voting round quantity, the encoder selects a second quantity of representative virtual speakers for the current frame from the first quantity of virtual speakers based on the first quantity of vote values, and further encodes the current frame based on the second quantity of representative virtual speakers for the current frame to obtain a bitstream. This achieves efficient data compression.


