Speech Control Method Using Pre-Wake Audio Segments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current intelligent speech interaction systems face challenges with inaccurate wake-up word detection and poor reliability in speech recognition, leading to false alarms and inconsistent recognition results due to the imprecision of wake-up detection algorithms.
Innovation Solution
A speech control method that acquires target audio data including audio collected before and after wake-up, performs speech recognition on this data, and controls devices based on instructions from a second audio segment following wake-up word detection, with the second segment either being later or overlapping with the first segment, to enhance recognition accuracy and reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If wake-up detection algorithm is used to detect wake-up word, then the system can respond to user commands, but the detection accuracy is poor leading to false alarms
Solution Approach 1:
The audio data is divided into multiple segments: first audio segment (before wake-up), second audio segment (after wake-up), and third audio segment (overlapping portion). This segmentation allows the system to analyze different time periods separately, improving the precision of wake-up word detection by focusing on the second audio segment while using the first segment for context, thereby reducing false alarms while maintaining reliable response to user commands.
2Measurement precision
If speech recognition is performed on audio data collected after wake-up only, then processing is simple and fast, but recognition accuracy is poor due to incomplete audio context
Solution Approach 1:
The system performs preliminary action by collecting and preparing the first audio segment (before wake-up) in advance. Although the main speech recognition is performed on the second audio segment (after wake-up), the pre-collected first audio segment provides essential context that improves recognition accuracy. This preliminary collection of audio data before wake-up occurs allows the system to maintain simple and fast processing while achieving better recognition results through contextual information.
3Measurement precision
If the second audio segment completely overlaps with the first audio segment, then all audio context is captured, but processing time and computational load increase
Solution Approach 1:
The system applies partial action by having the second audio segment extend beyond the first audio segment rather than completely overlapping it. The second audio segment includes the wake-up moment and continues for a duration that captures the essential instruction context without unnecessarily extending the processing window. This partial extension provides sufficient context for accurate instruction recognition while avoiding excessive processing time and computational load that would result from complete overlap or longer durations.
Data Source
AI summary
The disclosure provides a speech control method, a speech control apparatus, an electronic device, and a storage medium. The method includes: acquiring target audio data sent by a client, the target audio data including audio data collected by the client within a target duration before wake-up and audio data collected by the client after wake-up; performing speech recognition on the target audio data; and controlling the client based on an instruction recognized from a second audio segment of the target audio data in response to recognizing a wake-up word from a first audio segment at beginning of the target audio data; in which, the second audio segment is later than the first audio segment or has an overlapping portion with the first audio segment.


