Voice Wake-up Model Training via Decoding Module Configuration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice wake-up systems often awaken multiple devices simultaneously due to shared wake-up words, leading to inaccurate self-defined voice wake-up functionality.

Innovation Solution

A method and apparatus for training a voice wake-up model using a base model with an encoder-decoder architecture, where voice recognition training data is used to adjust the decoding module's configuration, followed by voice wake-up training to enhance the model's accuracy in recognizing user-defined wake-up words, thereby reducing false alarms and improving recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a shared wake-up word is used across multiple devices, then device simplicity and ease of operation are improved, but device differentiation and wake-up accuracy deteriorate

Engineering Contradiction:
Improveease of wake-up word configurationVSAvoidwake-up word recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent applies local quality by enabling each device to have a unique wake-up word or customized wake-up phrase while maintaining the same basic wake-up functionality. The system allows individual devices to be configured with specific wake-up words (e.g., different names for different devices) while using the same underlying technical framework, thus achieving both ease of operation and device differentiation.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamics by allowing the wake-up word to be configurable and changeable based on user needs. The system can adapt to different wake-up words dynamically rather than being fixed, enabling users to customize wake-up phrases for different devices or scenarios while maintaining the same core technology.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If multiple devices use the same wake-up word, then system simplicity is improved, but false alarm rate increases

Engineering Contradiction:
Improvesystem complexityVSAvoidfalse alarm rate
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent resolves this contradiction by implementing device-specific wake-up words or customizable wake-up phrases. Each device can be configured with its own unique wake-up word, which prevents false alarms while maintaining system simplicity. The underlying technical architecture remains unified, but the wake-up trigger is localized to each device's specific configuration.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If voice recognition training is performed on base model, then model general capability is improved, but wake-up specific accuracy deteriorates

Engineering Contradiction:
Improvevoice recognition capabilityVSAvoidwake-up word detection accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by first training the base model with general voice recognition capabilities using diverse voice data, then subsequently fine-tuning the model with wake-up specific training data. This two-stage approach ensures the model first learns general voice understanding and then specializes in accurate wake-up word detection, resolving the contradiction between general capability and specific accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements parameter changes by adjusting model parameters and training configurations between the base model training phase and the wake-up specific training phase. The system modifies training parameters, loss functions, and data weighting to prioritize wake-up word accuracy while retaining the benefits of general voice recognition training.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230317060A1Method and apparatus for training voice wake-up model, method and apparatus for voice wake-up, device, and storage medium
Publication Date: 2023.10.05 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20230317060A1 patent drawing
  • US20230317060A1 patent drawing
  • US20230317060A1 patent drawing

AI summary

The present disclosure provides a method and an apparatus for training a voice wake-up model, a method and an apparatus for voice wake-up, a device and a storage medium, which relates to the field of artificial intelligence and particularly to the field of deep learning and voice technology. A specific implementation lies in: acquiring voice recognition training data and voice wake-up training data that are created, and firstly performing training on a base model according to the voice recognition training data to obtain a model parameter of the base model when a model loss function converges; then updating, based on a model configuration instruction, a configuration parameter of a decoding module in the base model to obtain a first model; and finally performing training on the first model according to the voice wake-up training data to obtain a trained voice wake-up model when the model loss function converges.