Continuous Voice Wake-Up Processing for Terminal Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice wake-up systems in terminal devices can only achieve one-time wake-up, leading to low wake-up accuracy for continuous voice data, as they fail to recognize and respond to multiple instances of wake-up words in user input, resulting in incorrect voice recognition.
Innovation Solution
A method and apparatus that collect and recognize user voice data, performing a wake-up operation each time a wake-up word is detected, allowing for continuous wake-up and controlling voice recognition operations based on the occurrence of wake-up words, with the ability to send voice data after the wake-up word to a server for further recognition, and determining a weight value for reliability, to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the terminal device performs voice wake-up only once and then enters waiting state, then the device structure is simple, but the wake-up accuracy for continuous voice data is low
Solution Approach 1:
The patent implements continuous wake-up processing by maintaining the voice recognition state after the initial wake-up. The system continuously monitors for wake-up words in the voice stream rather than entering a static waiting state, enabling multiple wake-up cycles within a single voice interaction session. This continuous action approach resolves the contradiction by improving wake-up accuracy through repeated detection opportunities without requiring fundamentally new hardware.
Solution Approach 2:
The patent performs preliminary voice data collection and recognition preparation before the actual wake-up word is detected. The system pre-positions the voice recognition module in a ready state and maintains audio stream buffering, so that when wake-up words appear multiple times in continuous speech, the system can immediately process them without delay. This preliminary preparation enables accurate multi-instance wake-up recognition without adding significant processing complexity.
2Measurement precision
If the terminal device recognizes the entire voice data as one continuous recognition task, then the processing is simple, but the wake-up accuracy decreases when wake-up word appears multiple times
Solution Approach 1:
The patent segments the continuous voice data into distinct recognition tasks based on wake-up word occurrences. Each time a wake-up word is detected, the system creates a new recognition task boundary, separating the voice data into segments before and after each wake-up word. This segmentation allows the system to accurately identify multiple wake-up instances and process each segment appropriately, resolving the contradiction between simple processing and accurate multi-word recognition.
Solution Approach 2:
The patent implements dynamic adjustment of the voice recognition state based on real-time wake-up word detection. The system transitions between different processing modes: collecting voice data, detecting wake-up words, and executing recognition tasks. This dynamic state management allows the system to adapt its processing strategy based on whether wake-up words are detected multiple times in a sequence, optimizing both accuracy and efficiency without fixed rigid processing rules.
3Measurement precision
If the terminal device performs voice recognition on all voice data including before the wake-up word, then the processing is comprehensive, but the recognition results become incorrect when wake-up word appears multiple times
Solution Approach 1:
The patent extracts and separates the wake-up word detection function from the general voice recognition function. By identifying wake-up words as distinct markers that delimit recognition tasks, the system can extract and remove these marker segments from the continuous voice stream. This extraction allows the system to process only the relevant voice data segments (after each wake-up word) for actual recognition, eliminating incorrect recognition results that would occur if the entire continuous stream including multiple wake-up words was processed as one task.
Data Source
AI summary
The present disclosure provides a method, an apparatus and a storage medium for a wake-up processing of an application, a first voice data input by a user is collected and recognized, and a wake-up operation is performed on a target application each time when it is recognized that a wake-up word of the target application is included in the first voice data, where the wake-up word of the target application appears one or more times in the first voice data. The method, apparatus and storage medium for the wake-up processing of the application provided by the present disclosure can wake up the target application when the wake-up word appears one or more times in the first voice data input by the user, thereby improving a wake-up accuracy of the application.

