Neural Network Bandwidth Extension for Speech Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs, such as AMR-NB, limit the frequency range of speech signals, making them less intelligible and less pleasant, and upgrading to wider-band codecs like EVS requires significant network changes, while blind bandwidth extension methods often fail to significantly improve speech quality.
Innovation Solution
A neural network processor generates a parametric representation of the enhancement frequency range, which is used to process a raw signal generated by a separate raw signal generator, optimizing audio quality and complexity without requiring full neural network processing for the entire frequency range.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If blind bandwidth extension is applied to extend the frequency range without additional bits, then network infrastructure changes are avoided, but speech quality improvement is insufficient
Solution Approach 1:
The patent introduces an excitation signal as an intermediary element that mediates between the narrowband input signal and the wideband output signal. By generating and shaping this excitation signal in the enhancement frequency range, the system achieves quality improvement without requiring full neural network processing of the entire frequency range, thus resolving the contradiction between adaptability and quality.
2Manufacturing precision
If full neural network processing is applied to generate the entire enhancement frequency range, then speech quality is improved, but computational complexity increases
Solution Approach 1:
The patent segments the bandwidth extension process into two distinct parts: (1) generation of an excitation signal using a raw signal generator, and (2) shaping of this excitation signal using a separate neural network processor. This segmentation allows the neural network to process only the parametric representation of the enhancement frequency range rather than the entire signal, significantly reducing computational complexity while maintaining quality improvement.
3Adaptability or versatility
If spectral folding or translation is used to generate the excitation signal, then bandwidth extension is achieved, but algorithmic delay increases
Solution Approach 1:
The patent employs a raw signal generator that creates the excitation signal autonomously based on the narrowband input signal characteristics, without requiring complex spectral folding or translation operations. This self-service approach generates the enhancement frequency range signal directly, minimizing processing steps and reducing algorithmic delay while maintaining bandwidth extension capability.
Data Source
Figure 1
Figure 2a~2b
Figure 2c
AI summary
An apparatus for generating a bandwidth enhanced audio signal from an input audio signal (50) having an input audio signal frequency range, comprises: a raw signal generator (10) configured for generating a raw signal (60) having an enhancement frequency range, wherein the enhancement frequency range is not included in the input audio signal frequency range; a neural network processor (30) configured for generating a parametric representation (70) for the enhancement frequency range using the input audio frequency range of the input audio signal and a trained neural network (31 ); and a raw signal processor (20) for processing the raw signal (60) using the parametric representation (70) for the enhancement frequency range to obtain a processed raw signal (80) having frequency components in the enhancement frequency range, wherein the processed raw signal (80) or the processed raw signal and the input audio signal frequency range of the input audio signal represent the bandwidth enhanced audio signal.