Neural Network Bandwidth Extension for Speech Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio codecs, such as AMR-NB, limit the frequency range of speech signals, making them less intelligible and less pleasant, and upgrading to wider-band codecs like EVS requires significant network changes, while blind bandwidth extension methods often fail to significantly improve speech quality.

Innovation Solution

A neural network processor generates a parametric representation of the enhancement frequency range, which is used to process a raw signal generated by a separate raw signal generator, optimizing audio quality and complexity without requiring full neural network processing for the entire frequency range.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If blind bandwidth extension is applied to extend the frequency range without additional bits, then network infrastructure changes are avoided, but speech quality improvement is insufficient

Engineering Contradiction:
Improvebandwidth extension capabilityVSAvoidspeech quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent introduces an excitation signal as an intermediary element that mediates between the narrowband input signal and the wideband output signal. By generating and shaping this excitation signal in the enhancement frequency range, the system achieves quality improvement without requiring full neural network processing of the entire frequency range, thus resolving the contradiction between adaptability and quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If full neural network processing is applied to generate the entire enhancement frequency range, then speech quality is improved, but computational complexity increases

Engineering Contradiction:
Improvespeech qualityVSAvoidneural network complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the bandwidth extension process into two distinct parts: (1) generation of an excitation signal using a raw signal generator, and (2) shaping of this excitation signal using a separate neural network processor. This segmentation allows the neural network to process only the parametric representation of the enhancement frequency range rather than the entire signal, significantly reducing computational complexity while maintaining quality improvement.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If spectral folding or translation is used to generate the excitation signal, then bandwidth extension is achieved, but algorithmic delay increases

Engineering Contradiction:
Improvebandwidth extension capabilityVSAvoidalgorithmic delay
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent employs a raw signal generator that creates the excitation signal autonomously based on the narrowband input signal characteristics, without requiring complex spectral folding or translation operations. This self-service approach generates the enhancement frequency range signal directly, minimizing processing steps and reducing algorithmic delay while maintaining bandwidth extension capability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP3701527B1Apparatus, method or computer program for generating a bandwidth-enhanced audio signal using a neural network processor
Publication Date: 2023.08.30 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP3701527B1 patent drawingFigure 1
  • EP3701527B1 patent drawingFigure 2a~2b
  • EP3701527B1 patent drawingFigure 2c

AI summary

An apparatus for generating a bandwidth enhanced audio signal from an input audio signal (50) having an input audio signal frequency range, comprises: a raw signal generator (10) configured for generating a raw signal (60) having an enhancement frequency range, wherein the enhancement frequency range is not included in the input audio signal frequency range; a neural network processor (30) configured for generating a parametric representation (70) for the enhancement frequency range using the input audio frequency range of the input audio signal and a trained neural network (31 ); and a raw signal processor (20) for processing the raw signal (60) using the parametric representation (70) for the enhancement frequency range to obtain a processed raw signal (80) having frequency components in the enhancement frequency range, wherein the processed raw signal (80) or the processed raw signal and the input audio signal frequency range of the input audio signal represent the bandwidth enhanced audio signal.