Neural Network Audio Encoding With Perceptual Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs based on psychoacoustic models face challenges in maintaining audio signal quality while managing model complexity, leading to inefficiencies in encoding and decoding processes.
Innovation Solution
The implementation of a neural network-based audio signal encoding and decoding method using a perceptually weighted error function, which adjusts model complexity based on human auditory characteristics through a training process that generates a masking threshold, weight matrix, and weighted error function to optimize audio signal processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional psychoacoustic models are used for audio encoding, then audio signal quality can be maintained, but model complexity increases leading to encoding inefficiency
Solution Approach 1:
The patent transforms the audio encoding problem by changing the error function parameters to be perceptually weighted based on human auditory characteristics. Instead of using traditional uniform error metrics, the invention applies psychoacoustic weighting factors that reflect human perception sensitivity across different frequencies and masking conditions, thereby achieving better quality with the same model complexity or same quality with reduced complexity
2Reliability
If traditional uniform error functions are used in neural network training, then training is simpler, but audio perceptual quality is not optimized
Solution Approach 1:
The patent applies local quality by making the error function weights frequency-dependent and locally adapted to human auditory characteristics. Different frequency regions are weighted differently based on psychoacoustic principles - regions where human ears are more sensitive receive higher weights, while regions subject to masking effects receive lower weights. This localized weighting optimizes perceptual quality without uniformly increasing complexity across all frequency bands
3Productivity
If model complexity is reduced to improve encoding efficiency, then processing speed increases, but audio signal quality deteriorates
Solution Approach 1:
The patent changes the fundamental parameter of the error function from uniform to perceptually weighted, which allows the model to focus computational resources on the most perceptually important aspects of audio quality. This parameter transformation enables simpler models to achieve the same perceptual quality as complex models, or better quality with the same complexity, because the weighting directs optimization toward human-perceptible features rather than treating all frequency components equally
Data Source
AI summary
Provided is a training method of a neural network that is applied to an audio signal encoding method using an audio signal encoding apparatus, the training method including generating a masking threshold of a first audio signal before training is performed, calculating a weight matrix to be applied to a frequency component of the first audio signal based on the masking threshold, generating a weighted error function obtained by correcting a preset error function using the weight matrix, and generating a second audio signal by applying a parameter learned using the weighted error function to the first audio signal.


