Blind Bandwidth Extension Neural Network for Speech Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech communication technologies, such as VoIP and cellular networks, face challenges in maintaining high speech quality and intelligibility due to the limitations of narrowband codecs, which lack the transmission of higher frequency components, leading to suboptimal acoustic quality and intelligibility. Additionally, existing blind bandwidth extension methods struggle to efficiently process speech signals in real-time on embedded systems like mobile phones.
Innovation Solution
The development of deep neural network-based systems using convolutional and recurrent neural networks for end-to-end adversarial blind bandwidth extension, which decompose speech signals into spectral envelopes and excitation signals, and apply adversarial training with spectral normalization to enhance the quality of narrowband speech by artificially regenerating missing frequency components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If narrowband speech transmission is used to reduce data transmission, then transmission bandwidth is saved, but speech quality and intelligibility deteriorate
Solution Approach 1:
The patent introduces an intermediary BBWE system that processes narrowband speech signals to reconstruct missing high-frequency components. The neural network acts as a mediator between the compressed narrowband signal and the desired wideband output, generating artificial high-frequency content based on learned patterns from training data, thus improving speech quality without increasing transmission bandwidth
Solution Approach 2:
The patent transforms the speech signal parameters by converting narrowband frequency content (200-3400 Hz) into wideband content (50-7000 Hz). The neural network learns to map narrowband spectral features to wideband spectral features, effectively changing the frequency domain parameters to reconstruct missing high-frequency components
2Manufacturing precision
If complex BBWE algorithms are used to improve speech quality, then speech quality improves, but computational complexity increases
Solution Approach 1:
The patent replaces traditional mechanical signal processing methods (filter banks, spectral folding, nonlinear transformations) with a data-driven neural network approach. The neural network is trained offline to learn the complex mapping from narrowband to wideband speech, and during runtime, it directly generates wideband output from narrowband input, significantly reducing real-time computational complexity
Solution Approach 2:
The patent performs preliminary training of the neural network offline using large datasets of wideband speech. This preliminary action allows the network to learn complex bandwidth extension patterns during training, so that during actual deployment, the network can perform bandwidth extension with minimal real-time computation, as the complex processing has already been done during offline training
3Manufacturing precision
If traditional BBWE methods are used, then bandwidth extension is achieved, but algorithmic delay increases
Solution Approach 1:
The patent segments the speech signal into short frames (e.g., 20-25 ms) and processes each frame independently or with minimal overlap through the neural network. This segmentation allows for efficient batch processing and reduces the overall algorithmic delay compared to processing long speech segments, while maintaining quality through the neural network's ability to capture temporal dependencies within each frame
Data Source
AI summary
An apparatus for processing a narrowband speech input signal by conducting bandwidth extension of the narrowband speech input signal to obtain a wideband speech output signal according to an embodiment is provided. The apparatus includes a signal envelope extrapolator including a first neural network, wherein the first neural network is configured to receive as input values of the first neural network a plurality of samples of a signal envelope of the narrowband speech input signal, and configured to determine as output values of the first neural network a plurality of extrapolated signal envelope samples. Moreover, the apparatus includes an excitation signal extrapolator configured to receive a plurality of samples of an excitation signal of the narrowband speech input signal, and configured to determine a plurality of extrapolated excitation signal samples. Furthermore, the apparatus includes a combiner configured to generate the wideband speech output signal such that the wideband speech output signal is bandwidth extended with respect to the narrowband speech input signal depending on the plurality of extrapolated signal envelope samples and depending on the plurality of extrapolated excitation signal samples.


