Speech Wakeup Model Using DNN and CTC for Low-Power Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech wakeup technologies rely heavily on keyword-specific speech data, limiting their accessibility and flexibility, and are not well-suited for use in low-power devices.

Innovation Solution

A speech wakeup model trained with both general speech data and keyword-specific speech data using a Deep Neural Network (DNN) and Connectionist Temporal Classifier (CTC), allowing for user-defined keywords and efficient operation on low-power devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keyword-specific speech data is used for training, then speech wakeup accuracy is improved, but data accessibility and system flexibility deteriorate

Engineering Contradiction:
Improvespeech wakeup accuracyVSAvoidsystem flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent trains the speech wakeup model using general speech data that can serve multiple purposes - both as general speech recognition data and as keyword-specific data through data augmentation techniques. This universal approach allows the same training corpus to improve speech wakeup accuracy while maintaining system flexibility and adaptability across different keywords and scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If complex speech wakeup models are used, then speech wakeup accuracy is improved, but computational burden and power consumption increase

Engineering Contradiction:
Improvespeech wakeup accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs extensive model training and optimization in advance using general speech data, creating a pre-trained speech wakeup model that can be deployed on low-power devices. By doing the computationally intensive work beforehand, the model achieves high accuracy while requiring minimal computational resources during actual speech wakeup operations on power-constrained devices.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If keyword-specific speech data is used for training, then speech wakeup accuracy is improved, but data accessibility deteriorates

Engineering Contradiction:
Improvespeech wakeup accuracyVSAvoiddata accessibility
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent uses general speech data that is widely available from public corpora and diverse sources, making data acquisition easy and accessible. Through data augmentation and universal training approaches, this easily accessible general data is transformed into effective training material that achieves speech wakeup accuracy previously requiring scarce keyword-specific data.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3579227B1Voice wake-up method and device and electronic device
Publication Date: 2021.06.09 ADVANCED NEW TECHNOLOGIES CO LTD
  • EP3579227B1 patent drawingFigure 1~2
  • EP3579227B1 patent drawingFigure 3~5
  • EP3579227B1 patent drawingFigure 6~7

AI summary

A speech wakeup method, apparatus, and electronic device are disclosed in embodiments of this specification. The method includes: implementing speech wakeup by using a speech wakeup model that includes a Deep Neural Network (DNN) and a Connectionist Temporal Classifier (CTC). The speech wakeup model can be obtained by training with general speech data.