AI Speech Recognition Non-Utterance Interval Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems require multiple activations via a wakeup word for each speech command, leading to inefficiencies and increased user interaction, as they fail to detect and respond to subsequent utterances within a non-utterance interval without reactivation.
Innovation Solution
An artificial intelligence apparatus that detects non-utterance intervals and determines the presence of subsequent utterances based on characteristics of initial utterances, allowing it to maintain speech recognition activation and perform commands without the need for repeated wakeup words.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If the speech recognition function is deactivated after performing a command, then the system saves energy and reduces false activations, but the user must speak the wakeup word again for subsequent commands
Solution Approach 1:
The system performs preliminary detection of the non-utterance interval characteristics during the first utterance, predicting whether a second utterance will occur before the speech recognition function is deactivated. This allows the system to maintain activation only when necessary, avoiding both premature deactivation and unnecessary prolonged activation.
Solution Approach 2:
The system uses feedback from the detected non-utterance interval characteristics to dynamically adjust the speech recognition activation state. By analyzing characteristics such as silence duration and speech patterns, the system determines whether to remain active for potential second utterances, creating a closed-loop control system that optimizes both energy consumption and responsiveness.
2Productivity
If the speech recognition function remains activated to detect second utterances, then user interaction efficiency improves, but the system consumes more energy and may experience false activations
Solution Approach 1:
The system performs preliminary analysis of the non-utterance interval characteristics during the first utterance to predict the likelihood of a second utterance. Based on this preliminary assessment, it decides whether to maintain activation, thereby avoiding unnecessary energy consumption while still being prepared for potential additional commands.
Solution Approach 2:
The system applies partial action by maintaining speech recognition activation only for the duration necessary to detect potential second utterances, rather than keeping it continuously active. This partial maintenance of activation balances energy consumption with the need to capture additional user commands efficiently.
3Ease of operation
If the system detects non-utterance intervals and predicts second utterances, then it can maintain activation appropriately, but the detection and analysis process adds computational complexity
Solution Approach 1:
The system extracts only the essential characteristics from the non-utterance interval (such as silence duration and speech pattern features) that are necessary for predicting second utterances, rather than analyzing all possible speech features. This extraction approach reduces computational complexity while maintaining prediction accuracy.
Solution Approach 2:
The system applies different analysis depths to different parts of the speech signal. During non-utterance intervals, it focuses computational resources on detecting specific characteristics relevant to second utterance prediction, rather than performing comprehensive speech analysis, thereby optimizing the balance between detection accuracy and computational complexity.
Data Source
AI summary
Disclosed herein is an artificial intelligence apparatus including an input interface configured to receive speech data, and a processor configured to detect a non-utterance interval included in the speech data and determine presence/absence of a second utterance after the non-utterance interval according to characteristics of a first utterance before the non-utterance interval, when the non-utterance interval exceeds a set time.


