Multimodal Browser Voice Query Handling for Mobile Web Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional user interaction techniques on mobile devices, such as stylus and handwriting recognition, are not fast or accurate for accessing Web services due to smaller displays, making it difficult for users to interact with applications and Web services effectively.
Innovation Solution
A method using a mobile device to receive speech data, convert it into a query, and send it to a network service for search results, which are then formatted and rendered on the device, including dynamically generating a voice grammar for further queries, utilizing a multimodal browser and proxy server to handle communication and formatting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional user interaction techniques (stylus and handwriting recognition) are used on mobile devices, then users can interact with applications and Web services, but the interaction is not fast or accurate due to smaller displays
Solution Approach 1:
The patent replaces mechanical interaction methods (stylus and handwriting recognition) with voice-based speech recognition technology. Users speak natural language queries instead of manually writing or pointing, eliminating the limitations of small display sizes and enabling fast, accurate interaction with Web services on mobile devices.
2Productivity
If speech recognition is used to access Web services on mobile devices, then user interaction speed and accuracy improve, but the system complexity increases due to speech processing and multimodal markup requirements
Solution Approach 1:
The patent introduces a proxy server as an intermediary between the mobile device and Web services. The proxy server handles speech recognition, query processing, and result formatting, offloading complex processing from the mobile device. It also generates multimodal markup language documents that structure results for optimal presentation on small displays, managing system complexity centrally rather than distributed across numerous mobile devices.
Solution Approach 2:
The proxy server performs multiple functions: speech recognition, query formulation, Web service communication, result formatting, and voice grammar generation. This multi-functional approach consolidates complexity into a single system rather than requiring each mobile device to independently handle all these tasks, reducing overall system complexity while maintaining high productivity.
3Ease of operation
If search results are formatted for presentation on mobile device displays, then results are optimized for small screens, but additional processing time and complexity are required
Solution Approach 1:
The proxy server performs preliminary formatting of search results into optimized multimodal markup language documents before transmitting them to mobile devices. This advance preparation ensures results are ready for immediate display on small screens without requiring additional processing time at the device level, as the formatting work is completed in advance by the proxy server.
Data Source
AI summary
A method of obtaining information using a mobile device can include receiving a request including speech data from the mobile device, and querying a network service using query information extracted from the speech data, whereby search results are received from the network service. The search results can be formatted for presentation on a display of the mobile device. The search results further can be sent, along with a voice grammar generated from the search results, to the mobile device. The mobile device then can render the search results.


