Voice Library for Dynamic Markup Grammar in Multimodal Browsers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing visual browsers lack the ability to process both graphic and voice inputs, requiring developers to either create new multimodal browsers or redesign existing ones, which is time-consuming and costly, and complicates voice-enabling applications with distinct operations for rendering content, command, and control navigation.
Innovation Solution
A voice library is provided that includes a method for dynamically generating a markup language fragment specifying command and control and content navigation grammar, allowing an interpreter to process speech inputs and generate events for an application, enabling voice-enabling of applications for command and control and content navigation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If developers create a new multimodal browser or redesign an existing visual browser to provide voice functionality, then both graphic and voice processing capabilities are achieved, but development time and cost increase significantly
Solution Approach 1:
The patent segments the voice processing functionality into a separate, independent library that can be integrated into existing visual browsers without requiring complete redesign. This modular approach allows voice capabilities to be added as a distinct component, reducing development time and complexity while maintaining both graphic and voice processing capabilities.
Solution Approach 2:
The patent creates a universal voice processing library that can be applied to multiple different visual browsers and applications. This multi-functional component provides speech-to-text conversion, command interpretation, and navigation capabilities that work across various platforms, eliminating the need for separate development efforts for each application.
2Adaptability or versatility
If visual browsers are enhanced to process both graphic and voice inputs, then multimodal interaction is enabled, but system complexity increases
Solution Approach 1:
The patent introduces an intermediary voice processing library that acts as a mediator between the user's speech input and the visual browser's core functions. This intermediary layer handles speech-to-text conversion, grammar interpretation, and command routing, thereby enabling multimodal interaction without directly complicating the browser's core architecture.
Solution Approach 2:
By dividing the system into distinct modules (visual rendering engine, voice processing library, grammar interpreter, command executor), the patent reduces overall system complexity. Each module operates independently with well-defined interfaces, making the system easier to maintain and extend while supporting both graphic and voice inputs.
3Productivity
If distinct operations for rendering content, command, and control navigation are maintained separately, then each function operates efficiently, but voice-enabling requires coordinating multiple separate systems
Solution Approach 1:
The patent merges the voice processing functions (speech recognition, grammar interpretation, command generation) into a unified library that interfaces with existing content rendering and navigation systems. This consolidation allows voice-enabling to coordinate multiple separate operations through a single integrated interface, reducing coordination complexity while preserving the efficiency of distinct functional operations.
Data Source
AI summary
A method of voice-enabling an application for command and control and content navigation can include the application dynamically generating a markup language fragment specifying a command and control and content navigation grammar for the application, instantiating an interpreter from a voice library, and providing the markup language fragment to the interpreter. The method also can include the interpreter processing a speech input using the command and control and content navigation grammar specified by the markup language fragment and providing an event to the application indicating an instruction representative of the speech input.


