Scalable Stereo Speech Coding with Monaural Intermediary

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech coding methods using a monaural-stereo scalable configuration face inefficiencies when inter-channel correlation between stereo signals is low, leading to deteriorated prediction performance and coding efficiency.

Innovation Solution

A speech coding apparatus with a core layer for monaural signal encoding and an extension layer for stereo signal encoding, where a monaural signal is generated from both channel signals and used to synthesize prediction signals for each channel, enhancing prediction gain even with low inter-channel correlation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If pitch prediction between channels is used to encode stereo speech, then coding efficiency is improved when inter-channel correlation is high, but prediction performance deteriorates when inter-channel correlation is low

Engineering Contradiction:
Improvecoding efficiencyVSAvoidprediction performance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces a monaural signal as an intermediary between the two stereo channels. Instead of directly predicting one channel from the other (which fails when correlation is low), the system predicts each channel from the monaural signal derived from both channels. This intermediary approach maintains prediction performance even when direct inter-channel correlation is low, while still achieving efficient coding through the scalable monaural-stereo configuration.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If a monaural-stereo scalable configuration is used, then adaptability is improved for different communication scenarios, but device complexity increases compared to simple monaural encoding

Engineering Contradiction:
ImproveadaptabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the stereo encoding process into two independent layers: a core monaural layer and an extension stereo layer. Each layer can be independently decoded, allowing the system to adapt to different bandwidth and processing requirements. This segmentation provides scalability where receivers can choose to decode only the monaural signal or both stereo signals, balancing adaptability with manageable device complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The monaural signal generated in the core layer serves multiple functions: it is both the final output for monaural mode and the basis for predicting stereo channels in extension mode. This multi-functionality reduces overall system complexity by reusing the same encoded data for different output scenarios, rather than maintaining separate encoding paths.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7945447B2Sound coding device and sound coding method
Publication Date: 2011.05.17 III HOLDINGS 12 LLC
  • US7945447B2 patent drawing
  • US7945447B2 patent drawing
  • US7945447B2 patent drawing

AI summary

A sound coding device having a monaural/stereo scalable structure and capable of efficiently coding stereo sound. even when the correlation between the channel signals of a stereo signal is small. In a core layer coding block of this device, a monaural signal generating section generates a monaural signal from first and second-channel sound signal, a monaural signal coding section codes the monaural signal, and a monaural signal decoding section greatest a monaural decoded signal from monaural signal coded data and outputs it to an expansion layer coding block. In the expansion layer coding block, a first-channel prediction signal synthesizing section synthesizes a first-channel prediction signal from the monaural decoded signal and a first-channel prediction filter digitizing parameter and a second-channel prediction signal synthesizing section synthesizes a second-channel prediction signal from the monaural decoded signal and second-channel prediction filter digitizing parameter.