Speech Processing Method for Adaptive Wake-Up Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing wake-up word models fail to accurately detect wake-up words due to variations in speech rates and speaking habits, leading to instances where the wake-up word is not spoken within the training window, resulting in device failure.
Innovation Solution
A speech processing method that obtains initial speech information exceeding a preset duration threshold, identifies similar speech segments, deletes similar frames to create secondary speech information within the threshold, and analyzes this information to determine user intent, allowing for flexible speech input without requiring users to adjust their speaking speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a fixed training window length is used for wake-up word detection, then the model structure is simple and training is efficient, but users with slow speech rates or long pauses cannot complete their wake-up words within the window, causing detection failure
Solution Approach 1:
The patent applies dynamics by making the analysis window length adaptive rather than fixed. The system dynamically adjusts the analysis window length based on the actual speech duration, allowing the window to expand or contract to accommodate different speech rates and pauses while maintaining efficient detection
2Adaptability or versatility
If the analysis time window is extended to accommodate slow speakers, then adaptability to different speech rates improves, but the complexity of the analysis system increases and processing efficiency decreases
Solution Approach 1:
The patent applies parameter changes by adjusting the analysis window length parameter based on speech characteristics. Instead of creating a complex system with multiple fixed windows, the solution changes the single parameter of window length to adapt to different speech durations, maintaining system simplicity while improving adaptability
3Adaptability or versatility
If the analysis time window is extended to accommodate slow speakers, then adaptability to different speech rates improves, but processing time and computational resources increase
Solution Approach 1:
The system dynamically determines the analysis window length based on actual speech duration, processing only the necessary time span. This dynamic approach avoids the fixed overhead of extended windows while ensuring slow speakers have sufficient time to complete their wake-up words, optimizing both adaptability and processing efficiency
Data Source
AI summary
A speech processing method includes obtaining first speech information from a user, determining one or more similar speech segments in the first speech information and deleting one or more similar frames each of the one or more similar speech segments to obtain second speech information, and analyzing the second speech information to determine a user intent corresponding to the first speech information. A duration of the first speech information exceeds a preset analysis duration threshold, and a duration of the second speech information does not exceed the preset analysis duration threshold.


