Audio Encoder Switching Between Frequency and Speech Coding Domains
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio coding systems face challenges in switching between frequency-domain and speech coding modes at low bitrates, resulting in poor quality transitions and increased overhead, especially when restarting speech coders, which leads to prolonged steady-state attainment and introduction of distortions.
Innovation Solution
The proposed solution involves considering state information of filters after reset to shorten the start-up period and providing additional information for the decoder to warm up prediction synthesis filters before switching, allowing for smoother transitions between coding domains and reducing artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech coders are restarted frequently to adapt to changing signal conditions, then adaptability is improved, but the time to reach steady state increases and signal quality deteriorates
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing filter state information before a restart is needed. When a speech coder restart is anticipated, the filter states are saved in advance, allowing the coder to quickly restore to a steady state condition rather than重新开始 from initial conditions, thus reducing the time to reach steady state while maintaining adaptability.
Solution Approach 2:
The patent uses copying by creating and storing copies of filter state information from the previous steady state. These state copies are preserved and reused when restarts occur, allowing the speech coder to rapidly reconstruct its processing state without undergoing a prolonged startup period, thereby reducing steady-state attainment time while preserving adaptability.
2Adaptability or versatility
If speech coders are restarted frequently to adapt to changing signal conditions, then adaptability is improved, but signal quality deteriorates due to transient distortions
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing filter state information before a restart is needed. When a speech coder restart is anticipated, the filter states are saved in advance, allowing the coder to quickly restore to a steady state condition rather than重新开始 from initial conditions, thus reducing the time to reach steady state while maintaining adaptability.
Solution Approach 2:
The patent uses copying by creating and storing copies of filter state information from the previous steady state. These state copies are preserved and reused when restarts occur, allowing the speech coder to rapidly reconstruct its processing state without undergoing a prolonged startup period, thereby reducing steady-state attainment time while preserving adaptability.
3Ease of operation
If filter states are reset to initial conditions after a restart, then the speech coder can begin processing, but the start-up period is prolonged and artifacts are introduced
Solution Approach 1:
The patent uses copying by creating and storing copies of filter state information from the previous steady state. These state copies are preserved and reused when restarts occur, allowing the speech coder to rapidly reconstruct its processing state without undergoing a prolonged startup period, thereby reducing steady-state attainment time while preserving adaptability.
Solution Approach 2:
The patent applies preliminary anti-action by preemptively saving filter states before restart conditions occur. This preliminary preservation of state information counteracts the harmful effect of prolonged startup and artifact generation that would otherwise result from resetting to initial conditions, thereby eliminating switching artifacts while maintaining restart capability.
4Productivity
If the MDCT filter bank is used for critical sampling, then coding efficiency is improved, but speech quality at low bitrates deteriorates
Solution Approach 1:
The patent applies dynamics by implementing domain switching capability that allows the coding system to dynamically transition between frequency-domain coding (MDCT) and speech-domain coding (LPC/ACELP) based on the characteristics of the input signal. This dynamic adaptation enables the system to use MDCT for music signals where it provides high coding efficiency, while switching to speech-domain coding for speech signals where it maintains high quality at low bitrates, thus resolving the contradiction between coding efficiency and speech quality.
Solution Approach 2:
The patent uses parameter changes by modifying the coding domain parameter based on signal type detection. When speech signals are detected, the system changes from frequency-domain MDCT coding to speech-domain LPC/ACELP coding, adjusting the coding parameters to match the signal characteristics. This parameter change allows the system to achieve high coding efficiency for music while maintaining high speech quality at low bitrates.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An audio encoder (100) adapted for encoding frames of a sampled audio signal to obtain encoded frames, wherein a frame comprises a number of time domain audio samples, comprising a predictive coding analysis stage (110) for determining information on coefficients of a synthesis filter and information on a prediction domain frame based on a frame of audio samples. The audio encoder (100) further comprises a frequency domain transformer (120) for transforming a frame of audio samples to the frequency domain to obtain a frame spectrum and an encoding domain decider (130). Moreover, the audio encoder (100) comprises a controller (140) for determining an information on a switching coefficient when the encoding domain decider decides that encoded data of a current frame is based on the information on the coefficients and the information on the prediction domain frame when encoded data of a previous frame was encoded based on a previous frame spectrum.