Neural Network Audio Encoding With Perceptual Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio codecs based on psychoacoustic models face challenges in maintaining audio signal quality while managing model complexity, leading to inefficiencies in encoding and decoding processes.

Innovation Solution

The implementation of a neural network-based audio signal encoding and decoding method using a perceptually weighted error function, which adjusts model complexity based on human auditory characteristics through a training process that generates a masking threshold, weight matrix, and weighted error function to optimize audio signal processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional psychoacoustic models are used for audio encoding, then audio signal quality can be maintained, but model complexity increases leading to encoding inefficiency

Engineering Contradiction:
Improveaudio signal qualityVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the audio encoding problem by changing the error function parameters to be perceptually weighted based on human auditory characteristics. Instead of using traditional uniform error metrics, the invention applies psychoacoustic weighting factors that reflect human perception sensitivity across different frequencies and masking conditions, thereby achieving better quality with the same model complexity or same quality with reduced complexity

Inventive Principle:
Principle #35Parameter changes

2Reliability

If traditional uniform error functions are used in neural network training, then training is simpler, but audio perceptual quality is not optimized

Engineering Contradiction:
Improveperceptual audio qualityVSAvoidtraining complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by making the error function weights frequency-dependent and locally adapted to human auditory characteristics. Different frequency regions are weighted differently based on psychoacoustic principles - regions where human ears are more sensitive receive higher weights, while regions subject to masking effects receive lower weights. This localized weighting optimizes perceptual quality without uniformly increasing complexity across all frequency bands

Inventive Principle:
Principle #3Local quality

3Productivity

If model complexity is reduced to improve encoding efficiency, then processing speed increases, but audio signal quality deteriorates

Engineering Contradiction:
Improveencoding efficiencyVSAvoidaudio signal quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent changes the fundamental parameter of the error function from uniform to perceptually weighted, which allows the model to focus computational resources on the most perceptually important aspects of audio quality. This parameter transformation enables simpler models to achieve the same perceptual quality as complex models, or better quality with the same complexity, because the weighting directs optimization toward human-perceptible features rather than treating all frequency components equally

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11416742B2Audio signal encoding method and apparatus and audio signal decoding method and apparatus using psychoacoustic-based weighted error function
Publication Date: 2022.08.16 ELECTRONICS & TELECOMM RES INST
  • US11416742B2 patent drawing
  • US11416742B2 patent drawing
  • US11416742B2 patent drawing

AI summary

Provided is a training method of a neural network that is applied to an audio signal encoding method using an audio signal encoding apparatus, the training method including generating a masking threshold of a first audio signal before training is performed, calculating a weight matrix to be applied to a frequency component of the first audio signal based on the masking threshold, generating a weighted error function obtained by correcting a preset error function using the weight matrix, and generating a second audio signal by applying a parameter learned using the weighted error function to the first audio signal.