Steganographic Audio Sub-stream Embedding for Backwards Compatible 3D Teleconferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current 3D audio teleconferencing systems require higher bit rates and are not backwards compatible with legacy mono-only terminal devices, limiting adoption due to increased upgrade costs and compatibility issues.

Innovation Solution

A multi-party control unit (MCU) embeds lower-quality sub-streams representing sounds from other terminal devices into mixed audio data streams using steganography, allowing both mono and stereo playback compatibility without increasing bitrates, enabling 3D audio teleconferencing while supporting legacy devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If 3D audio teleconferencing is implemented using traditional methods, then spatial audio quality is improved, but bit rate increases and compatibility with legacy mono-only devices is lost

Engineering Contradiction:
Improvespatial audio qualityVSAvoidcompatibility with legacy devices
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The audio signal is segmented into multiple components: a high-quality mixed audio stream and embedded sub-streams containing spatial position information for each participant. This segmentation allows legacy devices to process only the mixed stream while 3D-capable devices can extract and utilize the embedded sub-streams for spatial rendering

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The sub-streams containing spatial information are nested within the mixed audio data stream using steganographic embedding techniques. This nesting allows the spatial information to be transported within the existing audio infrastructure without requiring separate communication channels or increasing overall bit rate

Inventive Principle:
Principle #7Nested doll (Nesting)

2Measurement precision

If 3D audio teleconferencing is implemented using traditional methods, then spatial audio quality is improved, but device complexity and upgrade costs increase

Engineering Contradiction:
Improvespatial audio qualityVSAvoidupgrade requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system provides backward compatibility services automatically - legacy mono-only devices simply play the mixed audio stream without needing to know about or process the embedded sub-streams. 3D-capable devices automatically detect and extract the embedded spatial information, enabling spatial audio without requiring manual configuration or complex upgrade procedures

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The mixed audio data stream serves multiple functions simultaneously: it provides high-quality audio for all devices while also carrying embedded spatial information for 3D-capable devices. This multi-functionality eliminates the need for separate audio streams or complex device-specific processing

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If sub-streams are embedded into mixed audio data streams using steganography, then both mono and stereo playback compatibility is achieved without increasing bitrates, but processing complexity increases

Engineering Contradiction:
Improveplayback compatibilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The steganographic embedding processes only the essential spatial information parameters rather than the entire audio signal. This partial action approach embeds only the necessary position data in the sub-streams, reducing the computational burden of both embedding and extraction operations while maintaining compatibility

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP2959669B1Teleconferencing using steganographically-embedded audio data
Publication Date: 2019.04.03 QUALCOMM INC
  • EP2959669B1 patent drawingFigure 1
  • EP2959669B1 patent drawingFigure 2
  • EP2959669B1 patent drawingFigure 3

AI summary

A multi-party control unit (MCU) generates, based on audio data streams that represent sounds associated terminal devices, a mixed audio data stream. In addition, the MCU modifies the mixed mono audio data to steganographically embed sub-streams that include representations of the mono audio data streams. A terminal device receives the modified mixed audio data stream. When the terminal device is configured for stereo playback, the terminal device performs an inverse steganographic process to extract, from the mixed audio data stream, the sub-streams. The terminal device generates and outputs multi-channel audio data based on the extracted sub-streams and the mixed audio data stream. When the terminal device is not configured for stereo playback, the terminal device outputs sound based on the mixed audio data stream without extracting the embedded sub-streams.