Wideband Speech Coding Split-Band Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech coding technologies are limited in transmitting wideband voice communications efficiently, particularly in extending the frequency range beyond traditional PSTN limits, which affects intelligibility and quality, especially for applications like VoIP and cellular telephony.
Innovation Solution
A wideband speech coding system employing a split-band coding scheme that separates the frequency range into narrowband and highband components, using different coding modes and rates for each band to efficiently encode and decode speech signals, allowing for transmission of a wider frequency range without significant quality loss.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speech coding is used to limit bandwidth to 300-3400 Hz, then transmission efficiency is maintained, but speech intelligibility and quality deteriorate due to loss of high-frequency information
Solution Approach 1:
The speech signal is divided into two separate frequency bands: narrowband (300-3400 Hz) and highband (3400-8000 Hz). Each band is processed independently with appropriate coding schemes, allowing efficient transmission of the complete wideband signal by segmenting the frequency spectrum into manageable portions that can be encoded separately and recombined at the receiver.
2Measurement precision
If wideband frequency range (50 Hz to 8 kHz) is transmitted, then speech quality and intelligibility are improved, but transmission bandwidth requirements increase
Solution Approach 1:
The wideband frequency range is segmented into narrowband and highband portions, each transmitted using optimized coding schemes. The narrowband portion uses traditional efficient codecs while the highband portion uses specialized highband extension techniques, allowing the system to achieve wideband quality without transmitting the entire wideband range at full resolution simultaneously.
Solution Approach 2:
Different quality levels and coding approaches are applied to different frequency regions. The narrowband portion receives standard coding treatment while the highband portion receives enhanced processing tailored to its specific characteristics, optimizing overall speech quality while managing bandwidth consumption efficiently through region-specific optimization.
3Quantity of substance
If different coding modes are used for active and inactive frames, then average bit rate is reduced, but complexity of frame classification and coding selection increases
Solution Approach 1:
The speech coder dynamically switches between different coding modes based on the activity detection of each frame. For inactive frames, a simplified low-bit-rate coding mode is used, while for active frames, a higher-bit-rate mode is employed. This dynamic adaptation allows the system to optimize bit rate efficiency by matching the coding complexity to the actual speech content requirements of each frame.
Data Source
Figure 1
Figure 2
Figure 3A~3B
AI summary
Applications of dim-and-burst techniques to coding of wideband speech signals are described. Reconstruction of a highband portion of a frame of a wideband speech signal using information from a previous frame is also described.