Speech Command Generation for Contextual Web Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in allowing users to easily identify and execute commands, as they often present lengthy lists of commands that are not contextually relevant, making it difficult for users to navigate applications and websites efficiently.
Innovation Solution
A system and method that converts specific content from web pages and applications into speech commands, allowing users to interact with visual elements by speaking commands directly, eliminating the need for lists by modifying visual elements into spoken formats and utilizing a Grammar Builder and Speech Enablement Module to generate speech grammars.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If speech recognition systems display a list of available commands to users, then users can see what commands are available, but the list becomes lengthly and users spend more time finding the desired command
Solution Approach 1:
The patent extracts only the contextually relevant commands from the full command set and displays them alongside the corresponding visual elements on the web page. Instead of showing all available commands in a separate list, the system identifies and extracts only those commands that are applicable to the current context, thereby reducing the displayed information to what is actually needed at that moment.
Solution Approach 2:
The patent applies local quality by making command availability context-dependent. Different portions of the interface display different commands based on the current context - commands are localized to specific visual elements or regions of the web page rather than being uniformly displayed everywhere. This ensures that users see only the commands relevant to their current position or action.
2Ease of manufacture
If speech recognition systems display commands in separate lists, then commands are organized, but commands are shown separated from the context in which they would be applied
Solution Approach 1:
The patent merges the command display with the visual content by integrating command text directly onto or near the relevant visual elements on the web page. Instead of separating commands into independent lists, the system combines the visual element and its associated speech command into a unified interface, allowing users to see both the visual target and the spoken command in the same context.
Solution Approach 2:
The patent adds a new dimension to the interface by overlaying speech command text on top of or adjacent to visual elements. This creates a multi-layered interface where visual and auditory information coexist in the same spatial domain, allowing users to perceive both dimensions simultaneously rather than requiring separate lists or views.
3Adaptability or versatility
If speech recognition systems provide many commands for various tasks, then system functionality is comprehensive, but it becomes difficult for users to remember the best commands to use
Solution Approach 1:
The patent makes the command set dynamic by automatically adjusting which commands are displayed based on the current context, user position, and active visual elements. The system continuously adapts the available command list rather than presenting a static comprehensive menu, thereby maintaining versatility while reducing the apparent complexity at any given moment.
Solution Approach 2:
The system performs preliminary action by pre-identifying and pre-displaying only the commands that will be relevant to the user's current task or position. Rather than presenting all possible commands and requiring users to filter through them, the system anticipates which commands will be needed and prepares them for display in advance, reducing the cognitive load on users.
Data Source
AI summary
A construction and display of speech commands system that allows a user to simply read what is on an application that involves visual elements with which the user interacts, and in doing so, gives the appropriate commands to the speech recognition system for the task at hand. The construction and display of speech commands system may include a speech recognition system, a grammar builder module, and a speech enablement module. The construction and display of speech commands system may automatically generate a speech enabled application from generated speech grammar.


