Voice Wake-up Model Training via Decoding Module Configuration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice wake-up systems often awaken multiple devices simultaneously due to shared wake-up words, leading to inaccurate self-defined voice wake-up functionality.
Innovation Solution
A method and apparatus for training a voice wake-up model using a base model with an encoder-decoder architecture, where voice recognition training data is used to adjust the decoding module's configuration, followed by voice wake-up training to enhance the model's accuracy in recognizing user-defined wake-up words, thereby reducing false alarms and improving recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a shared wake-up word is used across multiple devices, then device simplicity and ease of operation are improved, but device differentiation and wake-up accuracy deteriorate
Solution Approach 1:
The patent applies local quality by enabling each device to have a unique wake-up word or customized wake-up phrase while maintaining the same basic wake-up functionality. The system allows individual devices to be configured with specific wake-up words (e.g., different names for different devices) while using the same underlying technical framework, thus achieving both ease of operation and device differentiation.
Solution Approach 2:
The patent implements dynamics by allowing the wake-up word to be configurable and changeable based on user needs. The system can adapt to different wake-up words dynamically rather than being fixed, enabling users to customize wake-up phrases for different devices or scenarios while maintaining the same core technology.
2Device complexity
If multiple devices use the same wake-up word, then system simplicity is improved, but false alarm rate increases
Solution Approach 1:
The patent resolves this contradiction by implementing device-specific wake-up words or customizable wake-up phrases. Each device can be configured with its own unique wake-up word, which prevents false alarms while maintaining system simplicity. The underlying technical architecture remains unified, but the wake-up trigger is localized to each device's specific configuration.
3Adaptability or versatility
If voice recognition training is performed on base model, then model general capability is improved, but wake-up specific accuracy deteriorates
Solution Approach 1:
The patent applies preliminary action by first training the base model with general voice recognition capabilities using diverse voice data, then subsequently fine-tuning the model with wake-up specific training data. This two-stage approach ensures the model first learns general voice understanding and then specializes in accurate wake-up word detection, resolving the contradiction between general capability and specific accuracy.
Solution Approach 2:
The patent implements parameter changes by adjusting model parameters and training configurations between the base model training phase and the wake-up specific training phase. The system modifies training parameters, loss functions, and data weighting to prioritize wake-up word accuracy while retaining the benefits of general voice recognition training.
Data Source
AI summary
The present disclosure provides a method and an apparatus for training a voice wake-up model, a method and an apparatus for voice wake-up, a device and a storage medium, which relates to the field of artificial intelligence and particularly to the field of deep learning and voice technology. A specific implementation lies in: acquiring voice recognition training data and voice wake-up training data that are created, and firstly performing training on a base model according to the voice recognition training data to obtain a model parameter of the base model when a model loss function converges; then updating, based on a model configuration instruction, a configuration parameter of a decoding module in the base model to obtain a first model; and finally performing training on the first model according to the voice wake-up training data to obtain a trained voice wake-up model when the model loss function converges.


