Keyword Detection Speed Normalization for Speech Variations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Smart speech devices face challenges in detecting keywords due to variations in speech speed, leading to a low keyword detection rate, as users may speak keywords faster or slower than the predetermined speed, making it difficult for devices to accurately recognize the keywords.
Innovation Solution
A keyword detection method and apparatus that enhance the speech signal to a target speed and perform speed adjustment on the enhanced signal to improve speech recognition quality, thereby enhancing the detection rate of keywords in both fast and slow speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the smart speech device uses a predetermined keyword detection method at normal speech speed, then the detection process is simple, but the keyword detection rate is low when users speak at varying speeds
Solution Approach 1:
The patent applies preliminary action by performing speech enhancement and speed normalization on the input speech signal before keyword detection. The system pre-processes the speech to adjust its speed to a standard range and enhances its quality, ensuring that the keyword detection algorithm receives optimized input regardless of the original speech speed variations.
Solution Approach 2:
The patent implements parameter changes by dynamically adjusting the speech signal parameters including speed normalization to convert varying speech speeds into a standard speed range, and quality enhancement parameters to improve signal-to-noise ratio. These parameter transformations enable the detection system to maintain high accuracy across different speaking rates.
2Measurement precision
If the device processes speech at the original varying speed, then the processing time is shorter, but the speech recognition quality is poor leading to missed keywords
Solution Approach 1:
The system performs preliminary speed normalization and quality enhancement on the speech signal before keyword detection. By pre-adjusting the speech to standard speed and enhancing its quality, the system ensures accurate keyword recognition without requiring excessive processing time during the detection phase.
Solution Approach 2:
The patent replaces direct keyword matching on raw speech with a processed speech analysis approach. Instead of mechanically comparing keywords against variable-speed speech, the system substitutes this with speech enhancement and speed normalization processes that transform the input into an optimized form suitable for accurate detection.
3Productivity
If the device accelerates speech processing to reduce detection time, then the response speed improves, but the detection accuracy decreases for fast speech
Solution Approach 1:
The system performs preliminary speed normalization that converts fast speech into a standard speed range before detection. This pre-processing ensures that even fast speech is transformed into an optimal detection format, maintaining high accuracy while enabling efficient processing through standardized signal characteristics.
Solution Approach 2:
The patent implements dynamic speed normalization that adaptively adjusts the speech signal based on its original speed characteristics. The system dynamically transforms varying speech speeds into a unified standard range, enabling consistent detection accuracy across different speaking rates while optimizing processing efficiency.
Data Source
Figure 1~2
Figure 3~4
Figure 5~6
AI summary
A keyword detection method and a keyword detection device, wherein same can perform speed changing processing on an enhanced signal, and can improve the detection rate of a keyword in a high-speed voice or slow-speed voice. The keyword detection method comprises: acquiring an enhanced voice signal of a voice signal to be detected, wherein the enhanced voice signal corresponds to a target voice speed (101); performing speed changing processing on the enhanced voice signal to obtain a first speed-changed voice signal, wherein the first speed-changed voice signal corresponds to a first voice speed, and the first voice speed is inconsistent with the target voice speed (102); acquiring a first voice feature signal according to the first speed-changed voice signal (103); acquiring a keyword detection result corresponding to the first voice feature signal, wherein the keyword detection result is used for indicating whether there is a target keyword in the voice signal to be detected (104); and if it is determined according to the keyword detection result that there is a target keyword, executing an operation corresponding to the target keyword (105).