Scalable Audio Coding with Adaptive Temporal Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing scalable coding methods suffer from pre-echo and post-echo issues due to low temporal resolution in higher layers, leading to decreased sound quality, and existing solutions either increase processing complexity or result in wasted encoded data.
Innovation Solution
A coding apparatus and method that determines the start or end point of an active speech portion and selectively excludes specific bands from the coding target in the higher layer, using the lower layer's decoded signal to improve temporal masking and reduce echo effects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If transform coding is applied in the higher layer with low temporal resolution, then coding efficiency is improved, but pre-echo and post-echo occur due to distortion propagation across the entire frame
Solution Approach 1:
The patent applies local quality by making the temporal resolution of transform coding variable rather than uniform. In the higher layer, transform coding is performed with different temporal resolutions depending on the time region: long temporal resolution (e.g., 20ms frame) in regions where speech is present to improve coding efficiency, and short temporal resolution (e.g., 5ms sub-frame) in regions where speech is absent or transitioning to suppress pre-echo and post-echo. This localized adaptation of coding parameters resolves the contradiction between overall coding efficiency and local distortion suppression.
2Object-affected harmful factors
If the coding interval is shortened to improve temporal resolution, then pre-echo is suppressed, but the frame length becomes too short to achieve high coding efficiency in transform coding
Solution Approach 1:
The patent employs dynamics by making the temporal resolution of transform coding adaptive and variable rather than fixed. The coding interval dynamically changes based on speech activity detection: during speech transitions or silence periods, the interval shortens to suppress pre-echo; during steady-state speech, the interval lengthens to maximize coding efficiency. This dynamic adjustment allows the system to optimize both pre-echo suppression and coding efficiency at different time points within the same audio signal processing.
3Productivity
If the higher layer uses a long frame length for transform coding, then coding efficiency is improved, but distortion propagates over a wider time range causing more noticeable echo
Solution Approach 1:
The patent applies local quality by segmenting the audio signal processing into different temporal regions with different coding characteristics. In regions where speech is actively present, long frame lengths are used for transform coding to achieve high coding efficiency. In regions where speech is absent or transitioning (attack and release phases), short frame lengths are used to limit distortion propagation. This localized differentiation ensures that long frames improve efficiency only where beneficial, while short frames suppress echo where distortion would be noticeable.
Data Source
AI summary
Disclosed are an encoding device and a decoding device which suppress the occurrence of pre-echo artifacts and post-echo artifacts caused by a high layer having a low temporal resolution, and which implement high subjective quality encoding and decoding. An encoding device (100) carries out scalable coding comprising a low layer, and a high layer having a lower temporal resolution than that of the low layer. A start point detection unit (or end point detection unit) (150) determines the start point (or end point) of sections of the decoded low layer signal which have audio, and when the start point (or end point) is determined, a second layer encoding unit (160) selects a bandwidth to be excluded from encoding on the basis of the spectral energy from the decoded first layer signal, excludes the selected bandwidth, and encodes an error signal.


