Audio Encoder Cross-Processor for Seamless Domain Switching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal encoding and decoding technologies face challenges in maintaining audio quality, particularly for non-speech signals with prominent high-frequency content, due to band limitations and reduced accuracy in frequency domain encoders.

Innovation Solution

The proposed solution combines a time domain encoding/decoding processor with a frequency domain encoding/decoding processor that includes a gap filling functionality, allowing for full-band core encoding/decoding and seamless switching between encoding strategies using a cross-processor for continuous initialization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a frequency domain encoder is used for non-speech signals with prominent high-frequency content, then bandwidth extension can be achieved, but the accuracy and audio quality are reduced due to band limitations and inability to separately encode prominent harmonics

Engineering Contradiction:
Improvebandwidth extension capabilityVSAvoidencoding accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The audio signal is divided into multiple frequency bands, with different encoding strategies applied to each band. Frequency domain encoding is used for bands containing prominent harmonics to preserve accuracy, while other bands use bandwidth extension techniques. This segmentation allows simultaneous achievement of high accuracy for critical bands and bandwidth extension for overall signal.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different encoding quality levels are applied to different frequency regions based on their characteristics. Bands with prominent harmonics receive high-quality frequency domain encoding to preserve local accuracy, while other bands use lower-resolution bandwidth extension. This local quality differentiation resolves the contradiction by maintaining precision where needed while achieving bandwidth extension elsewhere.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If bandwidth extension is applied in frequency domain encoding, then the encoded signal can cover a wider frequency range, but the encoder must downsample the signal which reduces the maximum frequency content

Engineering Contradiction:
Improvefrequency range coverageVSAvoidmaximum frequency content
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The encoder dynamically adjusts the sampling rate and encoding strategy based on the input signal characteristics. For signals with prominent high-frequency content, the encoder maintains higher sampling rates and applies frequency domain encoding without aggressive downsampling. This dynamic adaptation allows the system to preserve maximum frequency content while still achieving bandwidth extension where appropriate.

Inventive Principle:
Principle #15Dynamics

3Productivity

If a time domain encoder is used for speech signals, then encoding efficiency is improved, but the encoder cannot effectively handle non-speech signals with prominent high-frequency harmonics

Engineering Contradiction:
Improveencoding efficiencyVSAvoidsignal type handling capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The encoding system is designed with multi-functionality to handle both speech and non-speech signals effectively. It incorporates both time domain and frequency domain encoding capabilities, allowing it to adapt to different signal types. The system automatically selects or combines encoding methods based on signal characteristics, achieving universal applicability while maintaining efficiency for speech and accuracy for high-frequency content.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Loss of energy

If both frequency domain and time domain encoding branches are used with bandwidth extension, then bitrate efficiency is increased, but the system becomes inflexible due to band limitations imposed by bandwidth extension procedures

Engineering Contradiction:
Improvebitrate efficiencyVSAvoidencoding flexibility
Core Design Contradiction:
Loss of energyVSAdaptability or versatility

Solution Approach 1:

The system dynamically selects and switches between frequency domain and time domain encoding branches based on the characteristics of the input signal. Rather than applying fixed bandwidth extension limitations, the encoder adapts its strategy in real-time, choosing the most appropriate method for each signal segment. This dynamic approach maintains bitrate efficiency while preserving encoding flexibility.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250124935A1Audio encoder and decoder using a frequency domain processor, a time domain processor, and a cross processor for continuous initialization
Publication Date: 2025.04.17 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US20250124935A1 patent drawing
  • US20250124935A1 patent drawing
  • US20250124935A1 patent drawing

AI summary

An audio encoder for encoding an audio signal includes: a first encoding processor for encoding a first audio signal portion in a frequency domain, wherein the first encoding processor includes: a time frequency converter for converting the first audio signal portion into a frequency domain representation having spectral lines up to a maximum frequency of the first audio signal portion; a spectral encoder for encoding the frequency domain representation; a second encoding processor for encoding a second different audio signal portion in the time domain; a cross-processor for calculating, from the encoded spectral representation of the first audio signal portion, initialization data of the second encoding processor, so that the second encoding processing is initialized to encode the second audio signal portion immediately following the first audio signal portion in time in the audio signal; a controller configured for analyzing the audio signal and for determining, which portion of the audio signal is the first audio signal portion encoded in the frequency domain and which portion of the audio signal is the second audio signal portion encoded in the time domain; and an encoded signal former for forming an encoded audio signal including a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second audio signal portion. d