Multi-Faceted GUI for Speech Command Routing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech-driven computer systems are limited in their ability to process complex commands that require multiple applications, as they typically need sequential execution and extensive training, and are not capable of handling natural language speech effectively, leading to user frustration and inefficiency.
Innovation Solution
A speech-enabled system that uses a multi-faceted graphical user interface and a natural language processor to interpret and execute multiple commands simultaneously, allowing users to control multiple applications with a single spoken command by parsing and routing voice inputs to the appropriate applications, while also providing flexibility in input methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If sequential execution of commands is used, then system complexity is reduced, but productivity decreases due to the need to execute multiple commands one by one
Solution Approach 1:
The system segments complex commands into individual simpler commands that can be processed sequentially by the operating system. The command parser divides the spoken command into discrete actionable units, each of which can be executed independently, thereby maintaining system simplicity while improving productivity.
Solution Approach 2:
The patent introduces a command parser and command activation statement mechanism as intermediaries between the user's spoken commands and the operating system. This intermediary layer translates complex natural language into structured command sequences, enabling the system to handle multiple commands without increasing underlying system complexity.
2Measurement precision
If extensive training is required for speech commands, then command accuracy improves, but ease of operation deteriorates due to the learning curve
Solution Approach 1:
The system performs preliminary action by establishing a command activation statement (such as saying 'OK Google') before issuing actual commands. This preliminary phrase sets the context for subsequent commands, allowing the system to recognize and execute commands accurately without requiring extensive training or complex phrase memorization by the user.
Solution Approach 2:
The command parser automatically processes and routes commands to appropriate applications without requiring user training. The system self-manages command interpretation, application selection, and execution sequencing, making the interface intuitive and easy to use while maintaining high command recognition accuracy.
3Ease of operation
If a single unified interface is used for multiple applications, then ease of operation improves, but device complexity increases due to integration requirements
Solution Approach 1:
The patent implements a universal interface through the command parser that can handle commands for multiple different applications and systems. The same parser infrastructure processes commands for web browsers, media players, messaging applications, and other systems, providing a single unified interface that simplifies user interaction while managing integration complexity internally.
Solution Approach 2:
The command parser serves as an intermediary layer that sits between the user and multiple applications. It translates unified speech commands into application-specific operations, allowing a single simple interface to control diverse applications without requiring each application to have its own unique interface, thereby reducing overall system complexity from the user's perspective.
4Ease of operation
If hands-free control is implemented, then ease of operation improves, but device complexity increases due to speech processing requirements
Solution Approach 1:
The command parser acts as an intermediary that receives raw speech input, processes it into structured commands, and routes them to appropriate applications. This intermediary layer handles the speech processing complexity centrally, allowing hands-free control to be implemented without distributing complex speech recognition requirements across multiple applications, thereby maintaining ease of operation while managing system complexity.
Data Source
AI summary
A multi faceted graphic user interface with multiple shells or layers may be provided for interaction with a user to speech enable interaction with applications and processes that do not necessarily have native support for speech input. The shells may be components of an operating system or of a parent application which supports such shells. Each shell has multiple facets for displaying applications and processes, and typically speech and other input is directed the application or process in the facet which has focus within the active shell. These multiple shells lend themselves to grouping of input or grouping of related applications and processes. For example, input from a speech recognizer, a mouse and a keyboard may each be directed at different shells; or a user may group related windows within various shells, such that all documents are displayed in one shell and all windows of an instant messaging application are displayed in another, thereby enabling better organization of work and work flow.


