Blind Bandwidth Extension Neural Network for Speech Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech communication technologies, such as VoIP and cellular networks, face challenges in maintaining high speech quality and intelligibility due to the limitations of narrowband codecs, which lack the transmission of higher frequency components, leading to suboptimal acoustic quality and intelligibility. Additionally, existing blind bandwidth extension methods struggle to efficiently process speech signals in real-time on embedded systems like mobile phones.

Innovation Solution

The development of deep neural network-based systems using convolutional and recurrent neural networks for end-to-end adversarial blind bandwidth extension, which decompose speech signals into spectral envelopes and excitation signals, and apply adversarial training with spectral normalization to enhance the quality of narrowband speech by artificially regenerating missing frequency components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If narrowband speech transmission is used to reduce data transmission, then transmission bandwidth is saved, but speech quality and intelligibility deteriorate

Engineering Contradiction:
Improvetransmission bandwidthVSAvoidspeech quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary BBWE system that processes narrowband speech signals to reconstruct missing high-frequency components. The neural network acts as a mediator between the compressed narrowband signal and the desired wideband output, generating artificial high-frequency content based on learned patterns from training data, thus improving speech quality without increasing transmission bandwidth

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the speech signal parameters by converting narrowband frequency content (200-3400 Hz) into wideband content (50-7000 Hz). The neural network learns to map narrowband spectral features to wideband spectral features, effectively changing the frequency domain parameters to reconstruct missing high-frequency components

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If complex BBWE algorithms are used to improve speech quality, then speech quality improves, but computational complexity increases

Engineering Contradiction:
Improvespeech qualityVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical signal processing methods (filter banks, spectral folding, nonlinear transformations) with a data-driven neural network approach. The neural network is trained offline to learn the complex mapping from narrowband to wideband speech, and during runtime, it directly generates wideband output from narrowband input, significantly reducing real-time computational complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent performs preliminary training of the neural network offline using large datasets of wideband speech. This preliminary action allows the network to learn complex bandwidth extension patterns during training, so that during actual deployment, the network can perform bandwidth extension with minimal real-time computation, as the complex processing has already been done during offline training

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If traditional BBWE methods are used, then bandwidth extension is achieved, but algorithmic delay increases

Engineering Contradiction:
Improvebandwidth extension qualityVSAvoidalgorithmic delay
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the speech signal into short frames (e.g., 20-25 ms) and processes each frame independently or with minimal overlap through the neural network. This segmentation allows for efficient batch processing and reduces the overall algorithmic delay compared to processing long speech segments, while maintaining quality through the neural network's ability to capture temporal dependencies within each frame

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230016637A1Apparatus and Method for End-to-End Adversarial Blind Bandwidth Extension with one or more Convolutional and/or Recurrent Networks
Publication Date: 2023.01.19 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US20230016637A1 patent drawing
  • US20230016637A1 patent drawing
  • US20230016637A1 patent drawing

AI summary

An apparatus for processing a narrowband speech input signal by conducting bandwidth extension of the narrowband speech input signal to obtain a wideband speech output signal according to an embodiment is provided. The apparatus includes a signal envelope extrapolator including a first neural network, wherein the first neural network is configured to receive as input values of the first neural network a plurality of samples of a signal envelope of the narrowband speech input signal, and configured to determine as output values of the first neural network a plurality of extrapolated signal envelope samples. Moreover, the apparatus includes an excitation signal extrapolator configured to receive a plurality of samples of an excitation signal of the narrowband speech input signal, and configured to determine a plurality of extrapolated excitation signal samples. Furthermore, the apparatus includes a combiner configured to generate the wideband speech output signal such that the wideband speech output signal is bandwidth extended with respect to the narrowband speech input signal depending on the plurality of extrapolated signal envelope samples and depending on the plurality of extrapolated excitation signal samples.