Multi-Faceted GUI for Speech Command Routing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech-driven computer systems are limited in their ability to process complex commands that require multiple applications, as they typically need sequential execution and extensive training, and are not capable of handling natural language speech effectively, leading to user frustration and inefficiency.

Innovation Solution

A speech-enabled system that uses a multi-faceted graphical user interface and a natural language processor to interpret and execute multiple commands simultaneously, allowing users to control multiple applications with a single spoken command by parsing and routing voice inputs to the appropriate applications, while also providing flexibility in input methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If sequential execution of commands is used, then system complexity is reduced, but productivity decreases due to the need to execute multiple commands one by one

Engineering Contradiction:
Improvecommand execution efficiencyVSAvoidcommand processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments complex commands into individual simpler commands that can be processed sequentially by the operating system. The command parser divides the spoken command into discrete actionable units, each of which can be executed independently, thereby maintaining system simplicity while improving productivity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a command parser and command activation statement mechanism as intermediaries between the user's spoken commands and the operating system. This intermediary layer translates complex natural language into structured command sequences, enabling the system to handle multiple commands without increasing underlying system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If extensive training is required for speech commands, then command accuracy improves, but ease of operation deteriorates due to the learning curve

Engineering Contradiction:
Improvecommand recognition accuracyVSAvoiduser interface usability
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs preliminary action by establishing a command activation statement (such as saying 'OK Google') before issuing actual commands. This preliminary phrase sets the context for subsequent commands, allowing the system to recognize and execute commands accurately without requiring extensive training or complex phrase memorization by the user.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The command parser automatically processes and routes commands to appropriate applications without requiring user training. The system self-manages command interpretation, application selection, and execution sequencing, making the interface intuitive and easy to use while maintaining high command recognition accuracy.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If a single unified interface is used for multiple applications, then ease of operation improves, but device complexity increases due to integration requirements

Engineering Contradiction:
Improveinterface simplicityVSAvoidsystem integration complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a universal interface through the command parser that can handle commands for multiple different applications and systems. The same parser infrastructure processes commands for web browsers, media players, messaging applications, and other systems, providing a single unified interface that simplifies user interaction while managing integration complexity internally.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The command parser serves as an intermediary layer that sits between the user and multiple applications. It translates unified speech commands into application-specific operations, allowing a single simple interface to control diverse applications without requiring each application to have its own unique interface, thereby reducing overall system complexity from the user's perspective.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of operation

If hands-free control is implemented, then ease of operation improves, but device complexity increases due to speech processing requirements

Engineering Contradiction:
Improvehands-free control capabilityVSAvoidspeech recognition system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The command parser acts as an intermediary that receives raw speech input, processes it into structured commands, and routes them to appropriate applications. This intermediary layer handles the speech processing complexity centrally, allowing hands-free control to be implemented without distributing complex speech recognition requirements across multiple applications, thereby maintaining ease of operation while managing system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11599332B1Multiple shell multi faceted graphical user interface
Publication Date: 2023.03.07 GREAT NORTHERN RES
  • US11599332B1 patent drawing
  • US11599332B1 patent drawing
  • US11599332B1 patent drawing

AI summary

A multi faceted graphic user interface with multiple shells or layers may be provided for interaction with a user to speech enable interaction with applications and processes that do not necessarily have native support for speech input. The shells may be components of an operating system or of a parent application which supports such shells. Each shell has multiple facets for displaying applications and processes, and typically speech and other input is directed the application or process in the facet which has focus within the active shell. These multiple shells lend themselves to grouping of input or grouping of related applications and processes. For example, input from a speech recognizer, a mouse and a keyboard may each be directed at different shells; or a user may group related windows within various shells, such that all documents are displayed in one shell and all windows of an instant messaging application are displayed in another, thereby enabling better organization of work and work flow.