Continuous Voice Wake-Up Processing for Terminal Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice wake-up systems in terminal devices can only achieve one-time wake-up, leading to low wake-up accuracy for continuous voice data, as they fail to recognize and respond to multiple instances of wake-up words in user input, resulting in incorrect voice recognition.

Innovation Solution

A method and apparatus that collect and recognize user voice data, performing a wake-up operation each time a wake-up word is detected, allowing for continuous wake-up and controlling voice recognition operations based on the occurrence of wake-up words, with the ability to send voice data after the wake-up word to a server for further recognition, and determining a weight value for reliability, to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the terminal device performs voice wake-up only once and then enters waiting state, then the device structure is simple, but the wake-up accuracy for continuous voice data is low

Engineering Contradiction:
Improvewake-up accuracyVSAvoidwake-up processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements continuous wake-up processing by maintaining the voice recognition state after the initial wake-up. The system continuously monitors for wake-up words in the voice stream rather than entering a static waiting state, enabling multiple wake-up cycles within a single voice interaction session. This continuous action approach resolves the contradiction by improving wake-up accuracy through repeated detection opportunities without requiring fundamentally new hardware.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent performs preliminary voice data collection and recognition preparation before the actual wake-up word is detected. The system pre-positions the voice recognition module in a ready state and maintains audio stream buffering, so that when wake-up words appear multiple times in continuous speech, the system can immediately process them without delay. This preliminary preparation enables accurate multi-instance wake-up recognition without adding significant processing complexity.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the terminal device recognizes the entire voice data as one continuous recognition task, then the processing is simple, but the wake-up accuracy decreases when wake-up word appears multiple times

Engineering Contradiction:
Improvewake-up accuracyVSAvoidvoice recognition efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the continuous voice data into distinct recognition tasks based on wake-up word occurrences. Each time a wake-up word is detected, the system creates a new recognition task boundary, separating the voice data into segments before and after each wake-up word. This segmentation allows the system to accurately identify multiple wake-up instances and process each segment appropriately, resolving the contradiction between simple processing and accurate multi-word recognition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic adjustment of the voice recognition state based on real-time wake-up word detection. The system transitions between different processing modes: collecting voice data, detecting wake-up words, and executing recognition tasks. This dynamic state management allows the system to adapt its processing strategy based on whether wake-up words are detected multiple times in a sequence, optimizing both accuracy and efficiency without fixed rigid processing rules.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If the terminal device performs voice recognition on all voice data including before the wake-up word, then the processing is comprehensive, but the recognition results become incorrect when wake-up word appears multiple times

Engineering Contradiction:
Improvewake-up accuracyVSAvoidvoice data recognition accuracy
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent extracts and separates the wake-up word detection function from the general voice recognition function. By identifying wake-up words as distinct markers that delimit recognition tasks, the system can extract and remove these marker segments from the continuous voice stream. This extraction allows the system to process only the relevant voice data segments (after each wake-up word) for actual recognition, eliminating incorrect recognition results that would occur if the entire continuous stream including multiple wake-up words was processed as one task.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11037560B2Method, apparatus and storage medium for wake up processing of application
Publication Date: 2021.06.15 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11037560B2 patent drawing
  • US11037560B2 patent drawing

AI summary

The present disclosure provides a method, an apparatus and a storage medium for a wake-up processing of an application, a first voice data input by a user is collected and recognized, and a wake-up operation is performed on a target application each time when it is recognized that a wake-up word of the target application is included in the first voice data, where the wake-up word of the target application appears one or more times in the first voice data. The method, apparatus and storage medium for the wake-up processing of the application provided by the present disclosure can wake up the target application when the wake-up word appears one or more times in the first voice data input by the user, thereby improving a wake-up accuracy of the application.