Voice Command Confidence Scoring for Real-Time Action Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice-activated devices experience delays in processing voice commands due to the need to receive and analyze complete voice commands before taking action, and often struggle to differentiate between words with similar pronunciation or meanings, leading to ineffective automation.

Innovation Solution

A method and system that assigns confidence scores to intended words in voice commands, altering them based on subsequent words, and performs actions when pre-determined confidence scores are reached, allowing for real-time processing and accurate interpretation of streaming voice commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the system waits to receive the complete voice command before processing, then the accuracy of action identification is improved, but the response time increases

Engineering Contradiction:
Improveaction identification accuracyVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing by assigning initial confidence scores to action words as they are recognized in the voice command stream, before the complete command is received. This allows the system to start processing early and reduce overall response time while maintaining accuracy through subsequent confidence score updates.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The confidence scores of action words are dynamically updated in real-time as the voice command is being spoken. The system alters confidence scores based on subsequent words that follow, allowing the action identification to evolve and improve accuracy progressively without waiting for the complete command.

Inventive Principle:
Principle #15Dynamics

2Loss of time

If the system processes voice commands in real-time word-by-word, then the response time is reduced, but the accuracy of differentiating between similar words decreases

Engineering Contradiction:
Improveprocessing delayVSAvoidword differentiation accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system uses feedback from subsequent words in the voice command to update and refine the confidence scores of previously identified action words. As more words are processed, the confidence scores are adjusted, allowing the system to maintain high word differentiation accuracy even when processing in real-time.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The confidence scores are not static but dynamically adjusted based on the context provided by subsequent words. This dynamic updating process allows the system to resolve ambiguities between similar words progressively as the voice command unfolds, maintaining accuracy while enabling real-time processing.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11232793B1Methods, systems and voice managing servers for voice recognition to perform action
Publication Date: 2022.01.25 ROBLOX CORP
  • US11232793B1 patent drawing
  • US11232793B1 patent drawing
  • US11232793B1 patent drawing

AI summary

A method for voice recognition to perform an action. The method includes receiving a voice command, identifying a first action intended word from the voice command, assigning a confidence score to the first action intended word, altering the confidence score of the first action intended word in a temporal manner, based on confidence scores of second action intended words following the first action intended word in the voice command, identifying the action when the confidence scores of the first and second action intended words reach a pre-determined confidence score associated therewith, and performing the identified action. Disclosed also is a system for voice recognition to perform an action. The system includes one or more voice-controlled devices, and a voice managing server communicably coupled to the one or more voice-controlled devices. The voice managing server for voice recognition to perform an action using the aforementioned method.