Voice Command Validation Using Speaker and Context Signatures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled devices struggle to accurately differentiate between valid and invalid voice commands, particularly in multi-user environments, leading to potential unauthorized actions or missed commands due to misidentification of the user uttering the command.

Innovation Solution

The use of speaker-identification information and additional characteristics, such as voice signatures, grammar, location, and context, to determine whether a user has issued a valid voice command, ensuring that only authorized users can perform actions on voice-controlled devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If speaker identification alone is used to determine valid voice commands, then device complexity is reduced, but measurement precision of user identity deteriorates in multi-user environments

Engineering Contradiction:
Improvecomplexity of voice command validation systemVSAvoidaccuracy of user identity identification
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple identification characteristics (speaker identification, grammar analysis, location data, and contextual information) into a unified validation system. This merging approach maintains relatively simple individual components while achieving high overall precision through synergistic combination of multiple data sources for comprehensive user authentication.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The voice command validation system is designed to handle multiple identification functions simultaneously - it can recognize speaker identity, analyze grammatical patterns, determine spatial location, and interpret contextual meaning all within a single integrated framework, making the system universally applicable to various voice command scenarios without requiring separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple identification characteristics are used to validate voice commands, then measurement precision of user identity improves, but device complexity increases

Engineering Contradiction:
Improveaccuracy of user identity identificationVSAvoidcomplexity of voice command validation system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The validation system is segmented into distinct functional modules, each handling a specific identification characteristic (speaker identification module, grammar analysis module, location determination module, and contextual analysis module). This segmentation allows each component to be optimized independently while working together to achieve comprehensive user authentication without overwhelming complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that receives multiple identification characteristics from different sources and synthesizes them into a unified validation decision. This intermediary function acts as a mediator between the various data sources and the final command execution, coordinating the multiple characteristics to reach accurate user identity determination while managing system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If voice commands are accepted without strict user verification, then ease of operation improves, but reliability of command execution deteriorates due to unauthorized actions

Engineering Contradiction:
Improveconvenience of voice command usageVSAvoidaccuracy of command execution
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system performs preliminary user identification and validation actions before executing any voice command. By establishing user identity and authorization status in advance through multiple characteristics (speaker recognition, grammar patterns, location, context), the system ensures that only authorized users can execute commands, thereby maintaining both ease of operation for verified users and reliability by preventing unauthorized actions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The validation system continuously monitors and compares multiple identification characteristics against established user profiles and contextual expectations. This feedback mechanism allows the system to dynamically adjust its validation decisions based on real-time observations, ensuring that commands are executed reliably only when the user identity and context align with expected patterns, while maintaining smooth operation for authenticated users.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9460715B2Identification using audio signatures and additional characteristics
Publication Date: 2016.10.04 AMAZON TECH INC
  • US9460715B2 patent drawing
  • US9460715B2 patent drawing
  • US9460715B2 patent drawing

AI summary

Techniques for using both speaker-identification information and other characteristics associated with received voice commands to determine how and whether to respond to the received voice commands. A user may interact with a device through speech by providing voice commands. After beginning an interaction with the user, the device may detect subsequent speech, which may originate from the user, from another user, or from another source. The device may then use speaker-identification information and other characteristics associated with the speech to attempt to determine whether or not the user interacting with the device uttered the speech. The device may then interpret the speech as a valid voice command and may perform a corresponding operation in response to determining that the user did indeed utter the speech. If the device determines that the user did not utter the speech, however, then the device may refrain from taking action on the speech.