Speech Encoding Modes and Rates via Closed-Loop Re-Decision
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech coders face challenges in maintaining high voice quality at low bit rates, particularly in wireless telephony and satellite communications, where they often introduce perceptually significant distortion due to limited bit rates and channel errors.
Innovation Solution
A device that dynamically adjusts encoding modes and rates using closed-loop re-decision mechanisms, representing speech signals by amplitude and phase components to optimize bit allocation and improve quality, allowing for flexible switching between different encoding modes and rates based on real-time analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If speech is transmitted by simply sampling and digitizing, then speech quality is maintained at conventional analog telephone levels, but data rate increases to sixty-four kilobits per second
Solution Approach 1:
The patent extracts only the essential information needed to represent speech by analyzing speech parameters (pitch, amplitude, spectral characteristics) and transmitting only these extracted features rather than the complete raw waveform, achieving compression while maintaining quality
Solution Approach 2:
The patent transforms the speech signal from raw waveform data into transformed parameters (cepstral coefficients, spectral envelope, pitch period) that capture the essential speech characteristics with fewer bits, changing the representation to achieve compression
2Quantity of substance
If speech compression is applied to reduce data rate, then data rate is significantly reduced, but voice quality deteriorates due to perceptually significant distortion
Solution Approach 1:
The patent implements feedback mechanisms where the receiver provides information about channel conditions and decoding results back to the transmitter, enabling adaptive adjustment of compression parameters to maintain quality while minimizing data rate
Solution Approach 2:
The patent dynamically adjusts compression parameters and speech coding modes based on real-time analysis of speech characteristics and channel conditions, allowing the system to optimize the balance between data rate and quality for each specific situation
3Device complexity
If fixed encoding modes and rates are used, then device complexity is reduced, but adaptability to varying channel conditions and speech characteristics is limited
Solution Approach 1:
The patent implements dynamic mode selection where the encoder can switch between different speech coding modes (CELP, PPP, NELP) and encoding rates based on real-time analysis of speech characteristics and channel conditions, providing adaptability without requiring all modes to be simultaneously active
Solution Approach 2:
The patent divides the speech coding system into separate modular components (mode decision unit, encoding unit, rate control unit) that can independently operate and select from different modes, allowing flexibility while maintaining manageable complexity through functional separation
Data Source
AI summary
In a device configurable to encode speech performing an closed loop re-decision may comprise representing a speech signal by amplitude components and phase components for a current frame and a past frame. In a first closed loop stage, a first set of compressed components and a first set of uncompressed components for a current frame may be generated. A first set of features may be generated by comparing current and past frame amplitude and/or phase components. In a second closed loop stage, a second set of compressed components for the current frame may be generated by compressing the first set of compressed components and compressing the first set of uncompressed components. Generation of a second set of features may be based on the second set of compressed components from the current frame and a combination of amplitude and/or phase components from the past frame.


