Distributed Multimodal Browser Voice Control via Link Messages

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current distributed multimodal browsers cannot perform speech-enabled content navigation and control due to the lack of information about the graphical user agent's interface being known a priori by the voice user agent, limiting their ability to enable voice commands for navigating and controlling content.

Innovation Solution

A distributed multimodal browser system where the graphical user agent transmits link messages specifying voice commands and events to the voice user agent, which recognizes voice utterances and returns event messages to control the browser, enabling speech-enabled content navigation and control by operatively coupling the two agents across separate computers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If the voice user agent operates on a separate voice server from the graphical user agent, then resource constraints on small devices are relieved, but the voice user agent lacks a priori knowledge of the graphical user agent's interface for speech-enabled navigation and control

Engineering Contradiction:
Improvedevice resource consumptionVSAvoidinterface information availability
Core Design Contradiction:
Use of energy by moving objectVSLoss of information

Solution Approach 1:

The system performs preliminary action by having the graphical user agent send interface description data to the voice user agent before voice commands are processed. This allows the voice user agent to have the necessary interface information available in advance, enabling speech-enabled navigation and control without requiring the voice user agent to be co-located with the graphical user agent on the device.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention introduces an intermediary mechanism where interface description data acts as a mediator between the graphical user agent and the voice user agent. This data structure conveys interface element information, relationships, and navigation capabilities from the graphical user agent to the voice user agent, enabling communication and coordination across the network without direct coupling.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the graphical user agent coordinates all user interaction among user agents, then a unified multimodal experience is provided, but speech-enabled content navigation and control cannot be achieved in distributed browsers

Engineering Contradiction:
Improvemultimodal interaction capabilityVSAvoidvoice command functionality
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system implements feedback by having the graphical user agent send interface description data to the voice user agent, which then processes voice commands and returns results. This feedback loop enables the voice user agent to understand the interface structure and provide appropriate voice-controlled navigation and control, maintaining unified multimodal interaction in distributed browser architectures.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8862475B2Speech-enabled content navigation and control of a distributed multimodal browser
Publication Date: 2014.10.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8862475B2 patent drawing
  • US8862475B2 patent drawing
  • US8862475B2 patent drawing

AI summary

Speech-enabled content navigation and control of a distributed multimodal browser is disclosed, the browser providing an execution environment for a multimodal application, the browser including a graphical user agent (‘GUA’) and a voice user agent (‘VUA’), the GUA operating on a multimodal device, the VUA operating on a voice server, that includes: transmitting, by the GUA, a link message to the VUA, the link message specifying voice commands that control the browser and an event corresponding to each voice command; receiving, by the GUA, a voice utterance from a user, the voice utterance specifying a particular voice command; transmitting, by the GUA, the voice utterance to the VUA for speech recognition by the VUA; receiving, by the GUA, an event message from the VUA, the event message specifying a particular event corresponding to the particular voice command; and controlling, by the GUA, the browser in dependence upon the particular event.