Multi-channel Stereo Audio Encoder with Channel Unmixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing CD+G audio format has limited capabilities for storing and displaying additional graphics data, requiring a special player for karaoke applications and lacking flexibility in mixing and unmixing audio channels, which restricts its use in professional and multi-channel audio applications.

Innovation Solution

A method and device for encoding and decoding multiple independent mono audio channels into a stereo recording, using a restricted set of additional parameters to allow for mixing and unmixing, enabling applications like karaoke, play-along, and true quadraphonic audio reproduction, while integrating MIDI data and lyrics for enhanced functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple independent mono audio channels are mixed into a stereo recording to enable multi-channel applications, then audio versatility and functionality are improved, but the complexity of encoding and decoding increases

Engineering Contradiction:
Improveaudio versatilityVSAvoidencoding and decoding complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The audio signal is segmented into multiple independent mono channels (lead vocal, lead instrument, backing vocals, instruments) that are mixed into stereo. The encoder separates these channels using spectral analysis and masking thresholds, allowing independent manipulation of each channel while maintaining stereo compatibility. This segmentation enables versatile audio applications without requiring complex full multi-channel encoding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A lead track identifier and masking threshold act as intermediaries between the original multi-channel master and the stereo recording. The encoder uses the lead track identifier to select which channels to emphasize, and the masking threshold to determine how much of each channel can be extracted without audible artifacts. This intermediary approach simplifies the encoding process compared to direct multi-channel storage.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If a restricted set of additional parameters is used to store channel information, then storage efficiency is improved, but the precision of channel reconstruction may be reduced

Engineering Contradiction:
Improvestorage efficiencyVSAvoidchannel reconstruction precision
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

Instead of storing complete information for perfect channel reconstruction, the system stores a restricted set of parameters (lead track identifier, masking thresholds, channel mixing matrices) that provide sufficient information for practical channel separation. The masking thresholds are calculated to ensure that extracted channels remain below perceptible levels, providing adequate precision without requiring exhaustive parameter storage.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system transforms the representation of audio channel information by using perceptual parameters (masking thresholds based on human hearing characteristics) rather than raw signal parameters. This parameter transformation allows efficient storage while maintaining reconstruction quality that is imperceptible to human listeners, effectively resolving the precision-storage tradeoff.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If separate lead tracks are stored independently for karaoke applications, then audio quality is improved, but storage space and system complexity increase

Engineering Contradiction:
Improveaudio qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Multiple audio channels (including lead vocal and lead instrument tracks) are merged into a single stereo recording with embedded parameter information. The lead tracks are not stored as separate physical files but are reconstructed from the stereo mix using the stored mixing matrices and masking thresholds. This merging approach maintains audio quality while eliminating the need for separate track storage and associated file management complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The stereo recording serves multiple functions simultaneously: it can be played as a standard stereo track, or the embedded parameters can be used to extract and manipulate individual channels for karaoke, instrumental isolation, or other audio processing applications. This multi-functionality eliminates the need for separate lead track files while maintaining the ability to access high-quality isolated channels when needed.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8009837B2Multi-channel compatible stereo recording
Publication Date: 2011.08.30 GOER DYNAMICS BV
  • US8009837B2 patent drawing
  • US8009837B2 patent drawing
  • US8009837B2 patent drawing

AI summary

An encoder for mixing a plurality of independent mono audio channels into a stereo recording and generating a restricted set of additional parameters used to master an audio track of a storage device is described. The plurality of independent mono audio channels are constructed such that the storage device can be played using an optical disk player so that in a first mode all of the plurality of independent mono audio channels are played as the stereo recording and in a second mode at least one of the plurality of independent mono audio channels can be unmixed and the stereo recording played with at least one mono audio channel removed. A corresponding decoder and an audio system comprising such encoder and decoder are also described.