Distributed Multimodal Browser Voice Control via Link Messages
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current distributed multimodal browsers cannot perform speech-enabled content navigation and control due to the lack of information about the graphical user agent's interface being known a priori by the voice user agent, limiting their ability to enable voice commands for navigating and controlling content.
Innovation Solution
A distributed multimodal browser system where the graphical user agent transmits link messages specifying voice commands and events to the voice user agent, which recognizes voice utterances and returns event messages to control the browser, enabling speech-enabled content navigation and control by operatively coupling the two agents across separate computers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If the voice user agent operates on a separate voice server from the graphical user agent, then resource constraints on small devices are relieved, but the voice user agent lacks a priori knowledge of the graphical user agent's interface for speech-enabled navigation and control
Solution Approach 1:
The system performs preliminary action by having the graphical user agent send interface description data to the voice user agent before voice commands are processed. This allows the voice user agent to have the necessary interface information available in advance, enabling speech-enabled navigation and control without requiring the voice user agent to be co-located with the graphical user agent on the device.
Solution Approach 2:
The invention introduces an intermediary mechanism where interface description data acts as a mediator between the graphical user agent and the voice user agent. This data structure conveys interface element information, relationships, and navigation capabilities from the graphical user agent to the voice user agent, enabling communication and coordination across the network without direct coupling.
2Adaptability or versatility
If the graphical user agent coordinates all user interaction among user agents, then a unified multimodal experience is provided, but speech-enabled content navigation and control cannot be achieved in distributed browsers
Solution Approach 1:
The system implements feedback by having the graphical user agent send interface description data to the voice user agent, which then processes voice commands and returns results. This feedback loop enables the voice user agent to understand the interface structure and provide appropriate voice-controlled navigation and control, maintaining unified multimodal interaction in distributed browser architectures.
Data Source
AI summary
Speech-enabled content navigation and control of a distributed multimodal browser is disclosed, the browser providing an execution environment for a multimodal application, the browser including a graphical user agent (‘GUA’) and a voice user agent (‘VUA’), the GUA operating on a multimodal device, the VUA operating on a voice server, that includes: transmitting, by the GUA, a link message to the VUA, the link message specifying voice commands that control the browser and an event corresponding to each voice command; receiving, by the GUA, a voice utterance from a user, the voice utterance specifying a particular voice command; transmitting, by the GUA, the voice utterance to the VUA for speech recognition by the VUA; receiving, by the GUA, an event message from the VUA, the event message specifying a particular event corresponding to the particular voice command; and controlling, by the GUA, the browser in dependence upon the particular event.


