Cascaded Subpixel CNNs for Speech Bandwidth Extension
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current communication systems face limitations in voice quality due to bandwidth constraints, particularly in Voice over Internet Protocol (VoIP) sessions, where narrowband signals with limited frequency ranges result in poor voice quality, and existing methods for bandwidth extension suffer from latency and distortion, making real-time processing challenging.
Innovation Solution
The implementation of cascaded subpixel convolutional neural networks (CNNs) to perform bandwidth extension, which receives narrowband audio data and generates wideband audio data by interpolating missing samples, effectively extending the speech bandwidth from 4 kHz to 8 kHz or higher, using machine learning techniques inspired by image super-resolution algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If narrowband signals are used for VoIP communication, then bandwidth consumption is reduced, but voice quality deteriorates
Solution Approach 1:
The patent applies parameter changes by transforming the narrowband signal parameters (frequency range 300-3400 Hz) into wideband signal parameters (frequency range 50-7000 Hz) through neural network processing. The system changes the spectral characteristics of the audio signal to achieve better voice quality while maintaining efficient bandwidth utilization in the communication channel.
2Manufacturing precision
If traditional bandwidth extension methods are used, then speech bandwidth is extended, but latency increases
Solution Approach 1:
The patent replaces traditional mechanical signal processing methods (filter banks, iterative optimization algorithms) with a neural network-based system. This substitution enables parallel processing of spectral components and eliminates computationally intensive iterative procedures, significantly reducing processing latency while achieving bandwidth extension from 300-3400 Hz to 50-7000 Hz.
3Manufacturing precision
If traditional bandwidth extension methods are used, then speech bandwidth is extended, but distortion increases
Solution Approach 1:
The patent introduces an intermediary neural network processing stage that acts as a bridge between narrowband input and wideband output. The neural network learns the complex mapping relationship between limited input frequencies and the full spectrum of human speech, generating missing spectral components without introducing distortion. This intermediary processing preserves speech naturalness while achieving bandwidth extension.
4Speed
If real-time processing is implemented, then communication responsiveness is improved, but processing complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-training the neural network offline to learn the complex transformation from narrowband to wideband signals. During real-time communication, the pre-trained model performs rapid inference without requiring complex computational operations. This separates the computationally intensive learning phase from the real-time processing phase, achieving both speed and simplicity in the communication system.
Data Source
AI summary
A system configured to improve a voice quality during a communication session by performing bandwidth extension on a narrowband speech signal to generate a wideband speech signal with higher audio quality. For example, a system can extend a speech bandwidth from a narrowband signal having a first bandwidth (e.g., 4 kHz) to a wideband signal having a second bandwidth (e.g., 8 kHz or higher). To perform bandwidth extension, the system may include cascaded neural networks, such as two or more sub-pixel convolutional neural networks (CNNs) connected in series. In some examples, a first sub-pixel CNN may extend the speech bandwidth from 4 kHz to 6 kHz and a second sub-pixel CNN may extend the speech bandwidth from 6 kHz to 8 kHz. Alternatively, the system may use three or more cascaded neural networks and/or may extend the speech bandwidth above 8 kHz without departing from the disclosure.


