Stereo Sound Decoding Using Time Domain Up-Mixing and Factor Beta
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current stereo sound encoding technologies face challenges in producing high-quality stereo sound at low bit-rates, especially in complex audio scenes with low correlation between sound signals, fluctuating background noise, and interfering talkers, often requiring high bit-rates to maintain quality, which is inefficient and affects sound intelligibility.
Innovation Solution
A stereo sound decoding method and system that uses time domain up-mixing with LP filter coefficients and a factor β to determine the contributions of primary and secondary channels, allowing for efficient bit allocation and improved sound quality at lower bit-rates, specifically using a modified EVS encoder for scalable bit-rate allocation between channels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If dual mono encoding is used to transmit stereo information, then stereo sound quality is maintained, but the bit-rate needs to be doubled
Solution Approach 1:
The patent combines the encoding of left and right stereo channels into a single integrated encoding process. Instead of independently encoding both channels (dual mono), the system encodes them together using a unified model that exploits correlations between channels, thereby reducing the total bit-rate required while preserving stereo quality.
Solution Approach 2:
The patent creates a universal encoding framework that can handle both mono and stereo content with a single encoder. The system adaptively switches between mono and stereo modes based on the audio content, allowing the same encoding infrastructure to serve multiple functions and reduce overall bit-rate requirements.
2Quantity of substance
If very low bit-rate is used for each channel to keep overall bit-rate reasonable, then bit-rate is reduced, but sound quality is affected
Solution Approach 1:
By merging the encoding of both stereo channels into a single integrated process, the system can allocate bits more efficiently across the combined signal. The unified encoder exploits inter-channel correlations to represent the stereo information more compactly, maintaining sound quality at lower overall bit-rates compared to independent channel encoding.
Solution Approach 2:
The patent employs parametric stereo encoding where instead of transmitting full waveform data for both channels, it transmits a mono signal plus parameters describing the stereo spatial characteristics. This parameter-based approach dramatically reduces bit-rate while preserving perceptual sound quality through efficient representation of spatial audio information.
3Quantity of substance
If parametric stereo encoding is used to reduce bit-rate, then bit-rate is reduced, but efficiency is insufficient at low bit-rates
Solution Approach 1:
The patent implements a dynamic encoding system that adaptively adjusts between different encoding modes (mono, stereo, parametric stereo) based on the characteristics of the audio content. The encoder dynamically selects the most efficient representation for each segment of audio, optimizing encoding efficiency across varying bit-rate conditions and audio scenarios.
Solution Approach 2:
The patent segments the audio signal into different processing stages and components (mono signal extraction, stereo parameter derivation, spatial encoding). By dividing the encoding process into distinct segments that can be independently optimized, the system achieves higher overall encoding efficiency compared to monolithic parametric stereo approaches.
4Measurement precision
If panning factor adaptation is made too fast, then spatial positioning is accurate, but it becomes disturbing to the listener
Solution Approach 1:
The patent implements dynamic adaptation of panning factors with controlled temporal characteristics. The system adjusts spatial parameters in response to audio content changes while applying smoothing and temporal filtering to prevent abrupt transitions. This dynamic approach maintains accurate speaker positioning while avoiding listener disturbance through controlled adaptation rates.
Solution Approach 2:
The patent applies preliminary smoothing and temporal filtering to panning factor adjustments before they are applied to the encoded signal. By cushioning rapid changes in spatial parameters through predictive algorithms and temporal averaging, the system prevents disturbing abrupt transitions while maintaining accurate spatial positioning over time.
5Object-affected harmful factors
If panning factor adaptation is made too slow, then listener comfort is maintained, but it does not reflect the real position of speakers
Solution Approach 1:
The patent implements multi-rate adaptive filtering for panning factor adjustment, using different time constants for different aspects of spatial parameter adaptation. Fast adaptation tracks speaker position changes accurately, while slower smoothing filters prevent listener discomfort from rapid transitions. This dynamic multi-timescale approach balances position accuracy with listener comfort.
Solution Approach 2:
The patent employs feedback mechanisms that monitor both the accuracy of speaker position representation and the stability of spatial parameters. The system uses this feedback to adjust adaptation rates dynamically, ensuring that panning factors converge to accurate speaker positions while maintaining listener comfort through controlled transition speeds based on real-time audio scene analysis.
6Reliability
If minimum bit-rate of 24 kb/s is used for wideband signals, then speech quality is decent, but it increases the bit-rate requirement
Solution Approach 1:
The patent transforms the encoding approach from transmitting full waveform data to transmitting parametric representations of the audio signal. By encoding mono signal parameters and stereo spatial parameters separately, the system achieves efficient bit-rate allocation that maintains speech quality at lower overall bit-rates than traditional dual mono or parametric stereo approaches.
Solution Approach 2:
The patent segments the bit-rate allocation between mono signal encoding and stereo parameter encoding. By dividing the total bit-rate budget into dedicated portions for different encoding components and optimizing each segment independently, the system achieves efficient use of available bits to maintain speech quality at lower overall bit-rates.
Data Source
AI summary
A stereo sound decoding method and system decode left and right channels of a stereo sound signal, using received encoding parameters comprising encoding parameters of a primary channel, encoding parameters of a secondary channel, and a factor β. The primary channel encoding parameters comprise LP filter coefficients of the primary channel. The primary channel is decoded in response to the primary channel encoding parameters. The secondary channel is decoded using one of a plurality of coding models, wherein at least one of the coding models uses the primary channel LP filter coefficients to decode the secondary channel. The decoded primary and secondary channels are time domain up-mixed using the factor β to produce the decoded left and right channels of the stereo sound signal, wherein the factor β determines respective contributions of the primary and secondary channels upon production of the left and right channels.


