Scalable Stereo Encoding via Shared Codebooks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing scalable coding techniques for stereo signals result in increased coding rate and circuit scale, leading to deteriorated speech quality, as they employ separate adaptive and fixed codebooks for each channel, which is not efficiently scalable for both stereo and monaural communication.
Innovation Solution
A scalable coding apparatus that predicts excitation signals for both channels using a monaural signal, reducing the number of codebooks and codebook indices required, thereby decreasing the coding rate and circuit scale while maintaining high speech quality by using a monaural coding section to generate a predicted excitation for both channels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If separate adaptive and fixed codebooks are used for each channel in stereo signal encoding, then speech quality can be maintained, but coding rate increases and circuit scale becomes larger
Solution Approach 1:
The patent combines the adaptive codebooks of both channels into a single shared adaptive codebook, and combines the fixed codebooks into a single shared fixed codebook. This merging reduces the number of codebooks from four (two per channel) to two total, thereby reducing circuit scale and coding rate while maintaining speech quality through shared resource utilization.
Solution Approach 2:
The patent creates universal codebooks that serve multiple channels simultaneously. The shared adaptive codebook and shared fixed codebook are used by both first and second channels during encoding and decoding processes, making these codebooks multi-functional resources that reduce overall system complexity while maintaining performance.
2Reliability
If separate adaptive and fixed codebooks are used for each channel in stereo signal encoding, then speech quality can be maintained, but coding rate increases
Solution Approach 1:
The patent merges the codebook structures across channels, reducing the total number of codebook indices that need to be transmitted. By sharing adaptive and fixed codebooks between channels, the coding rate is reduced because fewer unique codebook references need to be encoded and transmitted for stereo signals compared to separate per-channel codebooks.
3Device complexity
If the number of adaptive codebooks and fixed codebooks is reduced, then coding rate and circuit scale can be reduced, but speech quality of decoded signal deteriorates
Solution Approach 1:
The patent merges codebooks in a way that preserves speech quality by ensuring both channels can access the same comprehensive codebook resources. The shared adaptive codebook and shared fixed codebook contain sufficient diversity and coverage to support both channels, preventing quality deterioration despite the reduction in total codebook count.
Solution Approach 2:
The patent uses copying mechanisms where excitation signals and codebook structures from one channel can be referenced or copied to another channel. This allows efficient utilization of the reduced number of codebooks while maintaining the ability to reconstruct high-quality speech signals for both channels through intelligent signal processing.
4Reliability
If independent CELP coding is performed per channel, then speech quality can be maintained, but coding rate increases and circuit scale becomes larger
Solution Approach 1:
The patent merges the CELP coding processes by sharing codebook resources between channels. Instead of maintaining completely independent CELP coding systems for each channel, the invention combines the adaptive and fixed codebook resources, reducing the total information that needs to be coded and transmitted while preserving the quality benefits of CELP coding through shared resource access.
Data Source
AI summary
A scalable encoding device capable of reducing an encoding rate to reduce a circuit scale while preventing sound quality deterioration of a decoded signal. An extension layer is coarsely divided into a system for processing a first channel and a system for processing a second channel. A sound source predictor for processing the first channel predicts a drive sound source signal of the first channel from a drive sound source signal of a monaural signal, and outputs the predicted drive sound source signal through a multiplier to a first CELP encoder. A sound source predictor for processing the second channel predicts the drive sound source signal of the second channel from the drive sound source signal of the monaural signal and the output from the first CELP encoder, and outputs the predicted drive sound source signal through a multiplier to a second CELP encoder. The first and second CELP encoders perform CELP encoding operations of the individual channels using individual predicted drive sound source signals.


