Multimodal Browser Speech Search via ASR and VoiceXML
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech-enabled web content searching is limited to web pages that include voice markup, restricting access to speech-enabled functionality for most web content since many web pages do not exploit voice capabilities provided by markup languages like X+V.
Innovation Solution
Implementing speech-enabled web content searching using a multimodal browser with an automatic speech recognition (ASR) engine, which supports multiple interaction modes, including voice, and utilizes grammars to render, search, and perform actions on web content, regardless of whether the content is speech-enabled, by integrating with a VoiceXML interpreter and speech engine.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech-enabled web content searching is implemented using voice markup (X+V), then speech recognition capability is improved, but web content compatibility deteriorates because most web pages do not use voice markup
Solution Approach 1:
The patent introduces a speech engine as an intermediary component between the user and the web content. This speech engine includes an automatic speech recognition (ASR) system that can process voice input and convert it to text, and a text-to-speech (TTS) system that can convert text to voice output. The speech engine acts as a mediator that enables speech-enabled searching without requiring the web content itself to be formatted with voice markup, thus resolving the contradiction between speech recognition capability and web content compatibility
Solution Approach 2:
The patent implements a universal search system that can handle both speech-enabled and non-speech-enabled web content through a unified interface. The search system is designed to work with any web page regardless of whether it uses voice markup or not, making the speech recognition capability universally applicable across all web content. This multi-functionality allows the system to serve both speech-optimized and traditional web pages through the same mechanism
2Ease of operation
If voice markup (X+V) is used for web content, then user interaction ease is improved, but device complexity increases due to integration of speech engines and markup processing
Solution Approach 1:
The patent extracts the speech processing functionality from the web content itself and places it in a separate, standalone speech engine. By taking out the speech recognition and synthesis capabilities from the markup language and implementing them as independent components, the system reduces the complexity burden on individual web pages while maintaining enhanced user interaction capabilities. The speech engine operates as a separate service that can be applied to any web content without modifying the content structure
Solution Approach 2:
The search system is designed to automatically detect and adapt to the capabilities of the web content being accessed. When a user initiates a search, the system automatically determines whether the target web page supports voice markup and adjusts its behavior accordingly. This self-service mechanism eliminates the need for manual configuration or complex integration logic, simplifying the overall system architecture while maintaining ease of operation for users
Data Source
AI summary
Speech-enabled web content searching using a multimodal browser implemented with one or more grammars in an automatic speech recognition (‘ASR’) engine, with the multimodal browser operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal browser operatively coupled to the ASR engine, includes: rendering, by the multimodal browser, web content; searching, by the multimodal browser, the web content for a search phrase, including yielding a matched search result, the search phrase specified by a first voice utterance received from a user and a search grammar; and performing, by the multimodal browser, an action in dependence upon the matched search result, the action specified by a second voice utterance received from the user and an action grammar.


