Audio Encoder Decoder Overlapping Block Time-Domain Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional coding strategies face challenges in achieving optimal performance for both speech and music signals, as perceptual audio coders fail to match the quality of speech coders at low bit rates, and using speech coders for music results in significant quality impairments, while existing filterbanks like MDCT struggle with time-domain aliasing.

Innovation Solution

A combined time-domain and frequency-domain encoding and decoding approach that transforms time-domain data to the frequency domain for efficient decoding, combining overlapping blocks to minimize time aliasing and adapt to coding domain changes, allowing for seamless switching between audio codecs for music and speech content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If perceptual audio coders use MDCT filterbanks for frequency-domain encoding, then smooth cross-fade between processing blocks is achieved and blocking artifacts are avoided, but time-domain aliasing occurs and performance for speech signals at low bit rates deteriorates

Engineering Contradiction:
Improveaudio qualityVSAvoidtime-domain aliasing
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

The audio signal is divided into multiple processing blocks that are encoded separately in the frequency domain using MDCT, then reconstructed in the time domain through overlap-add of inverse MDCT transforms. This segmentation allows frequency-domain processing benefits while managing time-domain aliasing through structured block processing and overlapping.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary time-domain reconstruction stage that processes the frequency-domain encoded blocks through inverse MDCT transforms and overlap-add operations. This intermediary step serves as a mediator between frequency-domain encoding and time-domain output, resolving the aliasing issue while preserving encoding efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If speech coders use predictive approach in time domain, then excellent performance for speech signals at low bit rates is achieved, but performance for general audio signals/music deteriorates with significant quality impairments

Engineering Contradiction:
Improvespeech coding performanceVSAvoidaudio signal type adaptability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal audio coding system that can handle both speech and music signals through a single frequency-domain encoding framework. The MDCT-based encoder processes all audio types uniformly, eliminating the need for separate speech and music coders while maintaining high performance across different signal types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the fundamental encoding parameter from time-domain prediction to frequency-domain transform coding. This parameter change enables the encoder to adapt to different audio characteristics through frequency-selective processing, achieving versatility across speech and music while maintaining efficiency through perceptual modeling in the frequency domain.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If conventional concepts use layered combination with all partial coders active, then both time-domain and frequency-domain contributions are combined, but device complexity increases and switching between coding modes becomes cumbersome

Engineering Contradiction:
Improvecoding performanceVSAvoidencoder structure
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges the time-domain and frequency-domain coding approaches into a unified frequency-domain encoding system. By combining the perceptual modeling benefits of frequency-domain analysis with efficient transform coding, the system achieves high performance without requiring separate parallel coding paths or complex switching mechanisms.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified frequency-domain encoder serves multiple functions: it provides perceptual audio coding for music, maintains compatibility with speech coding requirements, and enables smooth transitions between different audio types. This multi-functional design eliminates the need for complex layered structures while maintaining coding effectiveness.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11961530B2Encoder, decoder and methods for encoding and decoding data segments representing a time-domain data stream
Publication Date: 2024.04.16 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US11961530B2 patent drawing
  • US11961530B2 patent drawing
  • US11961530B2 patent drawing

AI summary

An apparatus for decoding data segments representing a time-domain data stream, a data segment being encoded in the time domain or in the frequency domain, a data segment being encoded in the frequency domain having successive blocks of data representing successive and overlapping blocks of time-domain data samples. The apparatus includes a time-domain decoder for decoding a data segment being encoded in the time domain and a processor for processing the data segment being encoded in the frequency domain and output data of the time-domain decoder to obtain overlapping time-domain data blocks. The apparatus further includes an overlap/add-combiner for combining the overlapping time-domain data blocks to obtain a decoded data segment of the time-domain data stream.