Voice Interaction Plug-in for Non-Voice Web Pages
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication units with limited input mechanisms, such as twelve-button keypads, and non-voice enabled web pages and browsers hinder widespread adoption of voice and multimodal technology, as most web pages are written in basic HTML and browsers lack the capability to interpret voice enablement markup language (XHTML+Voice XML).
Innovation Solution
A software or hardware plug-in module that modifies web page attributes to enable voice interaction by converting voice input into text using a speech recognition server, allowing users to fill form fields on non-voice enabled web pages and browsers, even if they are not originally compatible with XHTML+Voice technology.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If web pages are written in XHTML+Voice XML to enable voice interaction, then voice functionality is improved, but compatibility with existing HTML web pages deteriorates
Solution Approach 1:
The patent introduces a gateway server as an intermediary component that translates between HTML and XHTML+Voice XML formats. The gateway receives requests from voice-capable browsers, converts them to compatible formats, and manages the interaction between voice-enabled and non-voice-enabled web pages, thereby resolving the compatibility issue without requiring all browsers to support XHTML+Voice XML natively
Solution Approach 2:
The patent segments the voice enablement functionality into a separate gateway component rather than requiring it to be embedded in every browser or web page. This segmentation allows voice capabilities to be added selectively through the gateway while leaving standard HTML browsers and pages unaffected, thus improving voice functionality without universally increasing system complexity
2Adaptability or versatility
If all web pages are converted to XHTML+Voice XML format, then voice interaction capability is improved, but implementation cost and complexity deteriorates
Solution Approach 1:
The gateway server acts as a mediator that performs the conversion from HTML to XHTML+Voice XML on-demand, eliminating the need for manual conversion of all existing web pages. This intermediary approach maintains voice interaction capability while avoiding the substantial effort required to convert the entire web ecosystem
Solution Approach 2:
The gateway server performs preliminary translation of HTML content into XHTML+Voice XML format before it reaches the user's browser. This preliminary action ensures that voice interaction capabilities are available without requiring end-users or web developers to perform conversion operations, significantly reducing implementation complexity
3Ease of operation
If speech recognition is implemented on the communication unit, then voice input capability is improved, but processing load and power consumption deteriorates
Solution Approach 1:
The gateway server serves as an intermediary that handles the computationally intensive speech recognition processing remotely. By offloading this function from the communication unit to the server, the system gains voice input capability while minimizing local processing load and power consumption
Solution Approach 2:
The patent replaces the mechanical/computational speech recognition system that would reside on the communication unit with a remote server-based system. This substitution maintains the user interface capability for voice input while transferring the energy-intensive processing to a remote location with unlimited power resources
Data Source
AI summary
Systems and methods for voice interaction with non-voice enabled web pages and browsers are provided. A communication unit that does not provide for voice enabled web browsing can be provided with a hardware and/or software plug-in. The plug-in can receive non-voice enabled web pages, receive voice from a user of the communication unit and provide the voice to a speech recognition server. The plug-in receives corresponding text from the speech recognition server and provides the text to a non-voice enabled web browser to fill-in form fields with the corresponding text.


