Scalable Audio Encoding with Two-Stage Error Band Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing scalable coding schemes face challenges in accurately specifying error bands across the full frequency band with high computational complexity, especially when the number of layers increases, leading to suboptimal speech quality due to predetermined subband settings and limited subband position adjustments.
Innovation Solution
The proposed solution involves an encoding and decoding apparatus that employs a two-step process to identify error bands: a first step searches for the band with the maximum error across a wider bandwidth with a coarse step size, and a second step refines the target frequency band within that range using a narrower step size, allowing for accurate specification of error bands with reduced computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a predetermined subband configuration is used for second layer encoding, then the device complexity is reduced, but the speech quality improvement is limited because error components cannot be accurately targeted
Solution Approach 1:
The patent divides the frequency band search into two segments: a first coarse search across the full bandwidth to identify candidate bands, and a second fine search within the candidate band to precisely locate the target frequency. This segmentation allows accurate error band specification without requiring exhaustive search of the entire frequency range, thus reducing device complexity while maintaining measurement precision.
Solution Approach 2:
The patent performs a preliminary coarse search to identify candidate frequency bands before conducting the final precise search. By pre-identifying the approximate location of maximum error bands, the system avoids the computational complexity of searching the entire frequency spectrum with fine granularity, thereby reducing device complexity while preserving the ability to accurately specify error bands.
2Measurement precision
If a full band search with fine step size is performed to accurately identify error bands, then the speech quality is improved, but the computational complexity increases significantly
Solution Approach 1:
The patent segments the frequency band search into two distinct phases: a first phase with wide bandwidth and coarse step size to identify candidate bands, and a second phase with narrow bandwidth and fine step size to precisely locate the target frequency within the candidate band. This segmentation enables accurate error band specification while minimizing computational complexity by limiting fine-grained search to only necessary frequency regions.
Solution Approach 2:
Instead of performing a complete fine-step search across the entire frequency band, the patent applies partial action by conducting the computationally intensive fine search only within the candidate band identified by the coarse search. This approach achieves sufficient measurement precision for accurate error band specification while avoiding the excessive computational complexity of a full band fine search.
3Adaptability or versatility
If the number of encoding layers is increased to improve speech quality, then the coding flexibility and quality are enhanced, but the computational complexity to specify subband positions increases substantially
Solution Approach 1:
The patent applies segmentation to the multi-layer encoding process by dividing the frequency band specification into coarse and fine search stages. This segmentation allows each layer to efficiently specify its target subband positions without requiring exhaustive searches, thereby maintaining coding flexibility across multiple layers while controlling the computational complexity that would otherwise increase substantially with each additional layer.
Solution Approach 2:
The patent uses preliminary coarse searches to identify candidate bands for each encoding layer before performing precise subband position specification. This preliminary action enables multiple layers to be configured with appropriate subband positions without each layer requiring independent exhaustive searches, thus maintaining adaptability and coding flexibility while preventing computational complexity from increasing substantially with the number of layers.
4Device complexity
If predetermined subbands are used for second layer encoding, then the device complexity is reduced, but the speech quality improvement is insufficient when error components are concentrated in non-predetermined bands
Solution Approach 1:
The patent introduces dynamics by making the target frequency band selection adaptive rather than fixed. The system dynamically identifies the frequency band with maximum error components through coarse and fine searches, allowing the second layer encoding to adapt to the actual error distribution in the input signal. This dynamic approach enables speech quality improvement by targeting actual error locations while maintaining relatively low device complexity through the two-step search methodology.
Solution Approach 2:
The patent implements feedback by using the error signal characteristics (maximum error components) to guide the selection of target frequency bands for second layer encoding. The system analyzes the error between original and first-layer decoded signals, identifies bands with maximum error energy, and uses this feedback information to dynamically configure the second layer encoding parameters, thereby achieving speech quality improvement without substantially increasing device complexity.
Data Source
Figure 1A~1C
Figure 2
Figure 3
AI summary
Disclosed is an encoding device which can accurately specify a band having a large error among all the bands by using a small calculation amount. The device includes: a first position identification unit (201) which uses a first layer error conversion coefficient indicating an error of decoding signal for an input signal so as to search for a band having a large error in a relatively wide bandwidth in all the bands of the input signal and generates first position information indicating the identified band; a second position identification unit (202) which searches for a target frequency band having a large error in a relatively narrow bandwidth in the band identified by the first position identification unit (201) and generates second position information indicating the identified target frequency band; and an encoding unit (203) which encodes a first layer decoding error conversion coefficient contained in the target frequency band. The first position information, the second position information, and the encoding unit are transmitted to a communication partner.