Voice Command Confidence Scoring for Real-Time Action Execution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice-activated devices experience delays in processing voice commands due to the need to receive and analyze complete voice commands before taking action, and often struggle to differentiate between words with similar pronunciation or meanings, leading to ineffective automation.
Innovation Solution
A method and system that assigns confidence scores to intended words in voice commands, altering them based on subsequent words, and performs actions when pre-determined confidence scores are reached, allowing for real-time processing and accurate interpretation of streaming voice commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the system waits to receive the complete voice command before processing, then the accuracy of action identification is improved, but the response time increases
Solution Approach 1:
The system performs preliminary processing by assigning initial confidence scores to action words as they are recognized in the voice command stream, before the complete command is received. This allows the system to start processing early and reduce overall response time while maintaining accuracy through subsequent confidence score updates.
Solution Approach 2:
The confidence scores of action words are dynamically updated in real-time as the voice command is being spoken. The system alters confidence scores based on subsequent words that follow, allowing the action identification to evolve and improve accuracy progressively without waiting for the complete command.
2Loss of time
If the system processes voice commands in real-time word-by-word, then the response time is reduced, but the accuracy of differentiating between similar words decreases
Solution Approach 1:
The system uses feedback from subsequent words in the voice command to update and refine the confidence scores of previously identified action words. As more words are processed, the confidence scores are adjusted, allowing the system to maintain high word differentiation accuracy even when processing in real-time.
Solution Approach 2:
The confidence scores are not static but dynamically adjusted based on the context provided by subsequent words. This dynamic updating process allows the system to resolve ambiguities between similar words progressively as the voice command unfolds, maintaining accuracy while enabling real-time processing.
Data Source
AI summary
A method for voice recognition to perform an action. The method includes receiving a voice command, identifying a first action intended word from the voice command, assigning a confidence score to the first action intended word, altering the confidence score of the first action intended word in a temporal manner, based on confidence scores of second action intended words following the first action intended word in the voice command, identifying the action when the confidence scores of the first and second action intended words reach a pre-determined confidence score associated therewith, and performing the identified action. Disclosed also is a system for voice recognition to perform an action. The system includes one or more voice-controlled devices, and a voice managing server communicably coupled to the one or more voice-controlled devices. The voice managing server for voice recognition to perform an action using the aforementioned method.


