Voice Interaction Plug-in for Non-Voice Web Pages

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing communication units with limited input mechanisms, such as twelve-button keypads, and non-voice enabled web pages and browsers hinder widespread adoption of voice and multimodal technology, as most web pages are written in basic HTML and browsers lack the capability to interpret voice enablement markup language (XHTML+Voice XML).

Innovation Solution

A software or hardware plug-in module that modifies web page attributes to enable voice interaction by converting voice input into text using a speech recognition server, allowing users to fill form fields on non-voice enabled web pages and browsers, even if they are not originally compatible with XHTML+Voice technology.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If web pages are written in XHTML+Voice XML to enable voice interaction, then voice functionality is improved, but compatibility with existing HTML web pages deteriorates

Engineering Contradiction:
Improvevoice functionalityVSAvoidbrowser capability requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a gateway server as an intermediary component that translates between HTML and XHTML+Voice XML formats. The gateway receives requests from voice-capable browsers, converts them to compatible formats, and manages the interaction between voice-enabled and non-voice-enabled web pages, thereby resolving the compatibility issue without requiring all browsers to support XHTML+Voice XML natively

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the voice enablement functionality into a separate gateway component rather than requiring it to be embedded in every browser or web page. This segmentation allows voice capabilities to be added selectively through the gateway while leaving standard HTML browsers and pages unaffected, thus improving voice functionality without universally increasing system complexity

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If all web pages are converted to XHTML+Voice XML format, then voice interaction capability is improved, but implementation cost and complexity deteriorates

Engineering Contradiction:
Improvevoice interaction capabilityVSAvoidweb page conversion effort
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The gateway server acts as a mediator that performs the conversion from HTML to XHTML+Voice XML on-demand, eliminating the need for manual conversion of all existing web pages. This intermediary approach maintains voice interaction capability while avoiding the substantial effort required to convert the entire web ecosystem

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The gateway server performs preliminary translation of HTML content into XHTML+Voice XML format before it reaches the user's browser. This preliminary action ensures that voice interaction capabilities are available without requiring end-users or web developers to perform conversion operations, significantly reducing implementation complexity

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If speech recognition is implemented on the communication unit, then voice input capability is improved, but processing load and power consumption deteriorates

Engineering Contradiction:
Improvevoice input capabilityVSAvoidpower consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The gateway server serves as an intermediary that handles the computationally intensive speech recognition processing remotely. By offloading this function from the communication unit to the server, the system gains voice input capability while minimizing local processing load and power consumption

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/computational speech recognition system that would reside on the communication unit with a remote server-based system. This substitution maintains the user interface capability for voice input while transferring the energy-intensive processing to a remote location with unlimited power resources

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8060371B1System and method for voice interaction with non-voice enabled web pages
Publication Date: 2011.11.15 NEXTEL COMMUNICATIONS INC
  • US8060371B1 patent drawing
  • US8060371B1 patent drawing
  • US8060371B1 patent drawing

AI summary

Systems and methods for voice interaction with non-voice enabled web pages and browsers are provided. A communication unit that does not provide for voice enabled web browsing can be provided with a hardware and/or software plug-in. The plug-in can receive non-voice enabled web pages, receive voice from a user of the communication unit and provide the voice to a speech recognition server. The plug-in receives corresponding text from the speech recognition server and provides the text to a non-voice enabled web browser to fill-in form fields with the corresponding text.