Perceptual Audio Coding via Reinforcement Learning Bit Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding technologies face challenges in dynamically allocating bits across subbands to achieve target bitrates while maintaining perceptual quality, especially in varying environmental conditions, due to the complexity of psychoacoustic models and the difficulty in capturing environmental factors.
Innovation Solution
The use of semi-supervised machine learning algorithms, specifically reinforcement learning, frames perceptual audio coding as a sequential decision-making problem, allowing for adaptive bit distribution across subbands to achieve target bitrates and ensure perceptual quality by interacting with environmental conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional psychoacoustic models are used for bit allocation, then perceptual quality can be maintained, but the system complexity and computational burden increase significantly
Solution Approach 1:
The patent replaces traditional psychoacoustic models with a reinforcement learning-based neural network system. The neural network learns optimal bit allocation strategies through interaction with the environment, substituting complex psychoacoustic calculations with a trained model that makes decisions based on learned patterns, thereby reducing computational burden while maintaining perceptual quality
Solution Approach 2:
The patent transforms the bit allocation problem from a deterministic optimization based on psychoacoustic parameters to a sequential decision-making problem where the neural network learns to allocate bits based on environmental states and rewards, changing the fundamental parameters and approach of the system
2Productivity
If dynamic bit allocation is implemented to adapt to environmental conditions, then compression efficiency improves, but the difficulty of capturing and responding to environmental factors increases
Solution Approach 1:
The patent implements a reinforcement learning framework where the neural network receives feedback in the form of rewards based on the perceptual quality of the encoded audio. This feedback mechanism allows the system to automatically adapt to environmental conditions and learn optimal bit allocation strategies without requiring explicit detection and measurement of all environmental factors
Solution Approach 2:
The neural network autonomously learns and adapts to environmental conditions through self-interaction with the encoding environment. The system serves itself by learning from rewards and penalties, eliminating the need for external control or manual tuning to respond to environmental changes
3Adaptability or versatility
If semi-supervised machine learning is used for bit allocation, then adaptability to environmental changes improves, but the extent of automation and training requirements increase
Solution Approach 1:
The patent employs semi-supervised learning where the neural network is pre-trained with labeled data containing ground truth bit allocation information. This preliminary training provides a head start, allowing the system to adapt to environmental changes more effectively while reducing the extent of automation needed during the learning phase, as the network already possesses baseline knowledge from supervised training
Data Source
AI summary
In general, techniques are described by which to perform perceptual audio coding as sequential decision making problems. A source device comprising a memory and a processor may be configured to perform the techniques. The memory may store at least a portion of the audio data. The processor may apply a filter to the audio data to obtain subbands of the audio data. The processor may adapt a controller according to a machine learning algorithm, the controller configured to determine bit distributions across the subbands of the audio data. The processor may specify, based on the bit distributions and in a bitstream representative of the audio data, one or more indications representative of the subbands of the audio data, and output the bitstream via a wireless connection in accordance with a wireless communication protocol.


