AI Speech Bandwidth Extension with Continual Learning Retention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing bandwidth extension technologies using artificial intelligence suffer from performance degradation during continual learning, particularly when transitioning between different tasks.
Innovation Solution
A framework integrating masked speech modeling (MSM) pre-training and bandwidth extension (BWE) fine-tuning through continual learning, utilizing a combination of loss functions including Fisher information matrix and exponential moving average, to maintain and enhance model performance across tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If continual learning is performed to adapt to new data distributions, then adaptability is improved, but learning effect decreases significantly
Solution Approach 1:
The patent applies preliminary action by pre-training the AI model on a large corpus of speech data before fine-tuning on specific tasks. This pre-training establishes a robust foundation that prevents catastrophic forgetting during continual learning. The model learns general speech representations first, then adapts to specific bandwidth extension tasks while maintaining previously acquired knowledge.
Solution Approach 2:
The patent utilizes parameter changes by dynamically adjusting learning rates and optimization parameters during continual learning. Different learning rates are applied to different parameter groups, with lower learning rates for parameters critical to previous tasks and higher rates for new task-specific parameters. This enables the model to adapt to new data distributions while preserving previously learned capabilities.
2Adaptability or versatility
If model parameters are updated continuously to learn new tasks, then adaptability is improved, but performance on previous tasks degrades
Solution Approach 1:
The patent applies local quality by treating different parameter groups with different update strategies. Critical parameters that govern general speech processing capabilities are updated with lower learning rates or constrained more strictly, while task-specific parameters are updated more aggressively. This localized differential update approach allows the model to learn new tasks while preserving performance on previous tasks.
Solution Approach 2:
The patent implements feedback mechanisms by monitoring performance on both new and previous tasks during continual learning. When performance degradation is detected on previous tasks, the system adjusts learning parameters or applies regularization techniques to restore performance. This feedback loop ensures that adaptability to new tasks does not come at the cost of forgetting previous capabilities.
3Measurement precision
If bandwidth extension is performed to improve speech clarity, then speech quality is improved, but computational complexity increases
Solution Approach 1:
The patent replaces traditional mechanical signal processing methods with an AI-based neural network approach. Instead of using complex filter banks and spectral manipulation techniques, the model learns direct mappings from narrowband to wideband signals through training. This substitution reduces computational complexity while maintaining or improving speech clarity, as the neural network can perform bandwidth extension in a single forward pass rather than through multiple iterative processing stages.
Data Source
AI summary
A method performed by an electronic device using artificial intelligence according to an embodiment of the disclosure, the method may include receiving a first narrowband signal and outputting a wideband signal by using the first narrowband signal as input in a pre-trained first artificial intelligence algorithm model.


