Audio Encoder Switching Between Frequency and Speech Coding Domains

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio coding systems face challenges in switching between frequency-domain and speech coding modes at low bitrates, resulting in poor quality transitions and increased overhead, especially when restarting speech coders, which leads to prolonged steady-state attainment and introduction of distortions.

Innovation Solution

The proposed solution involves considering state information of filters after reset to shorten the start-up period and providing additional information for the decoder to warm up prediction synthesis filters before switching, allowing for smoother transitions between coding domains and reducing artifacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech coders are restarted frequently to adapt to changing signal conditions, then adaptability is improved, but the time to reach steady state increases and signal quality deteriorates

Engineering Contradiction:
ImproveadaptabilityVSAvoidsteady-state attainment time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing filter state information before a restart is needed. When a speech coder restart is anticipated, the filter states are saved in advance, allowing the coder to quickly restore to a steady state condition rather than重新开始 from initial conditions, thus reducing the time to reach steady state while maintaining adaptability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating and storing copies of filter state information from the previous steady state. These state copies are preserved and reused when restarts occur, allowing the speech coder to rapidly reconstruct its processing state without undergoing a prolonged startup period, thereby reducing steady-state attainment time while preserving adaptability.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If speech coders are restarted frequently to adapt to changing signal conditions, then adaptability is improved, but signal quality deteriorates due to transient distortions

Engineering Contradiction:
ImproveadaptabilityVSAvoidsignal quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing filter state information before a restart is needed. When a speech coder restart is anticipated, the filter states are saved in advance, allowing the coder to quickly restore to a steady state condition rather than重新开始 from initial conditions, thus reducing the time to reach steady state while maintaining adaptability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating and storing copies of filter state information from the previous steady state. These state copies are preserved and reused when restarts occur, allowing the speech coder to rapidly reconstruct its processing state without undergoing a prolonged startup period, thereby reducing steady-state attainment time while preserving adaptability.

Inventive Principle:
Principle #26Copying

3Ease of operation

If filter states are reset to initial conditions after a restart, then the speech coder can begin processing, but the start-up period is prolonged and artifacts are introduced

Engineering Contradiction:
Improverestart capabilityVSAvoidswitching artifacts
Core Design Contradiction:
Ease of operationVSObject-generated harmful factors

Solution Approach 1:

The patent uses copying by creating and storing copies of filter state information from the previous steady state. These state copies are preserved and reused when restarts occur, allowing the speech coder to rapidly reconstruct its processing state without undergoing a prolonged startup period, thereby reducing steady-state attainment time while preserving adaptability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies preliminary anti-action by preemptively saving filter states before restart conditions occur. This preliminary preservation of state information counteracts the harmful effect of prolonged startup and artifact generation that would otherwise result from resetting to initial conditions, thereby eliminating switching artifacts while maintaining restart capability.

Inventive Principle:
Principle #9Preliminary anti-action

4Productivity

If the MDCT filter bank is used for critical sampling, then coding efficiency is improved, but speech quality at low bitrates deteriorates

Engineering Contradiction:
Improvecoding efficiencyVSAvoidspeech quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent applies dynamics by implementing domain switching capability that allows the coding system to dynamically transition between frequency-domain coding (MDCT) and speech-domain coding (LPC/ACELP) based on the characteristics of the input signal. This dynamic adaptation enables the system to use MDCT for music signals where it provides high coding efficiency, while switching to speech-domain coding for speech signals where it maintains high quality at low bitrates, thus resolving the contradiction between coding efficiency and speech quality.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent uses parameter changes by modifying the coding domain parameter based on signal type detection. When speech signals are detected, the system changes from frequency-domain MDCT coding to speech-domain LPC/ACELP coding, adjusting the coding parameters to match the signal characteristics. This parameter change allows the system to achieve high coding efficiency for music while maintaining high speech quality at low bitrates.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2311034B1Audio encoder and decoder for encoding frames of sampled audio signals
Publication Date: 2015.11.04 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP2311034B1 patent drawingFigure 1
  • EP2311034B1 patent drawingFigure 2
  • EP2311034B1 patent drawingFigure 3

AI summary

An audio encoder (100) adapted for encoding frames of a sampled audio signal to obtain encoded frames, wherein a frame comprises a number of time domain audio samples, comprising a predictive coding analysis stage (110) for determining information on coefficients of a synthesis filter and information on a prediction domain frame based on a frame of audio samples. The audio encoder (100) further comprises a frequency domain transformer (120) for transforming a frame of audio samples to the frequency domain to obtain a frame spectrum and an encoding domain decider (130). Moreover, the audio encoder (100) comprises a controller (140) for determining an information on a switching coefficient when the encoding domain decider decides that encoded data of a current frame is based on the information on the coefficients and the information on the prediction domain frame when encoded data of a previous frame was encoded based on a previous frame spectrum.