Scalable Stereo Speech Coding with Monaural Intermediary
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech coding methods using a monaural-stereo scalable configuration face inefficiencies when inter-channel correlation between stereo signals is low, leading to deteriorated prediction performance and coding efficiency.
Innovation Solution
A speech coding apparatus with a core layer for monaural signal encoding and an extension layer for stereo signal encoding, where a monaural signal is generated from both channel signals and used to synthesize prediction signals for each channel, enhancing prediction gain even with low inter-channel correlation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If pitch prediction between channels is used to encode stereo speech, then coding efficiency is improved when inter-channel correlation is high, but prediction performance deteriorates when inter-channel correlation is low
Solution Approach 1:
The patent introduces a monaural signal as an intermediary between the two stereo channels. Instead of directly predicting one channel from the other (which fails when correlation is low), the system predicts each channel from the monaural signal derived from both channels. This intermediary approach maintains prediction performance even when direct inter-channel correlation is low, while still achieving efficient coding through the scalable monaural-stereo configuration.
2Adaptability or versatility
If a monaural-stereo scalable configuration is used, then adaptability is improved for different communication scenarios, but device complexity increases compared to simple monaural encoding
Solution Approach 1:
The patent segments the stereo encoding process into two independent layers: a core monaural layer and an extension stereo layer. Each layer can be independently decoded, allowing the system to adapt to different bandwidth and processing requirements. This segmentation provides scalability where receivers can choose to decode only the monaural signal or both stereo signals, balancing adaptability with manageable device complexity.
Solution Approach 2:
The monaural signal generated in the core layer serves multiple functions: it is both the final output for monaural mode and the basis for predicting stereo channels in extension mode. This multi-functionality reduces overall system complexity by reusing the same encoded data for different output scenarios, rather than maintaining separate encoding paths.
Data Source
AI summary
A sound coding device having a monaural/stereo scalable structure and capable of efficiently coding stereo sound. even when the correlation between the channel signals of a stereo signal is small. In a core layer coding block of this device, a monaural signal generating section generates a monaural signal from first and second-channel sound signal, a monaural signal coding section codes the monaural signal, and a monaural signal decoding section greatest a monaural decoded signal from monaural signal coded data and outputs it to an expansion layer coding block. In the expansion layer coding block, a first-channel prediction signal synthesizing section synthesizes a first-channel prediction signal from the monaural decoded signal and a first-channel prediction filter digitizing parameter and a second-channel prediction signal synthesizing section synthesizes a second-channel prediction signal from the monaural decoded signal and second-channel prediction filter digitizing parameter.


