Voice Interface Entity Recognition and Action Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users of voice interfaces on computing devices often face difficulties remembering the available actions, their functions, and when to use specific voice commands due to the increasing complexity of voice interfaces.

Innovation Solution

A method where a computing device receives a spoken utterance, performs speech recognition to identify an entity, and then indicates available actions relevant to that entity, allowing the user to select and initiate the appropriate action through further speech input or manual interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the voice interface supports more actions and functions, then the versatility and capability of the system is improved, but the complexity of the interface increases making it harder for users to remember and use

Engineering Contradiction:
Improvevoice interface capabilityVSAvoidinterface complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary action by automatically identifying the entity from the user's spoken input and pre-filtering the available actions to only those relevant to the identified entity. This eliminates the need for users to remember all possible actions, as the system proactively presents only the applicable subset based on contextual understanding.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system employs feedback by listening to the user's spoken input, identifying the intended entity, and then responding with a tailored list of relevant actions. This closed-loop feedback mechanism allows the interface to adapt to user intent dynamically, presenting information in a context-relevant manner that reduces cognitive load.

Inventive Principle:
Principle #23Feedback

2Loss of information

If the system presents all available actions to the user, then the user has complete information, but the user experiences confusion and difficulty in selecting the appropriate action

Engineering Contradiction:
Improveinformation completenessVSAvoiduser experience
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system extracts and presents only the relevant subset of actions associated with the identified entity, rather than displaying all available actions. This extraction principle filters out irrelevant information while preserving the complete set of entity-specific actions, thereby maintaining information completeness within the relevant context while improving ease of operation.

Inventive Principle:
Principle #2Taking out (Extraction)

3Extent of automation

If the voice interface requires users to remember action names and functions, then the system can operate with simple recognition, but the ease of use deteriorates as users struggle to recall commands

Engineering Contradiction:
Improvespeech recognition simplicityVSAvoidcommand recall difficulty
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The system practices self-service by automatically determining which actions are relevant based on the entity identified from user input. Instead of requiring users to manually recall or search for appropriate commands, the system autonomously curates and presents the relevant action set, thereby maintaining simple speech recognition while dramatically improving ease of operation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8706505B1Voice application finding and user invoking applications related to a single entity
Publication Date: 2014.04.22 GOOGLE LLC
  • US8706505B1 patent drawing
  • US8706505B1 patent drawing
  • US8706505B1 patent drawing

AI summary

A computing device is configured to initiate actions in response to speech input that includes a name or other indication of an entity, in a first spoken utterance, followed by an action, in a second spoken utterance. The computing device receives the first spoken utterance, identifies an entity based on the first spoke utterance, and indicates a plurality of available actions based on the identified entity. The computing device then receives the second spoken utterance and identifies a selection of at least one of the available actions based on the second spoken utterance. The computing device then initiates the at least one selected action.