Voice-Enabled Webpage Handoff for Continuous Site Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice assistants struggle to effectively navigate and interact with websites, particularly those designed for graphical user interfaces, as they often fail to translate visual content to voice and have difficulty with keyboard-based interactions.
Innovation Solution
A web server processes voice commands by identifying the requested webpage and providing data for a voice-enabled webpage that invokes a voice interface, allowing continuous voice interaction with the user device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If a voice assistant retrieves a webpage focused on providing visual content and links, then the webpage can provide comprehensive visual information, but the voice assistant cannot provide the visual content or easily translate the visual content to voice
Solution Approach 1:
The patent introduces a mediator component that translates visual webpage content into voice-compatible formats. This intermediary layer converts images, videos, and visual elements into audio descriptions that the voice assistant can process and convey to the user, resolving the contradiction between preserving visual information and enabling voice interaction.
Solution Approach 2:
The patent replaces the mechanical visual display system with an acoustic output system. Instead of relying on visual content delivery, the system substitutes speech synthesis and audio output mechanisms to convey the same information, enabling voice assistants to effectively present webpage content without visual displays.
2Ease of operation
If a website is designed to be interacted with using a keyboard, mouse, or other method applicable to a graphical user interface, then the website can provide precise control and interaction, but the voice assistant has difficulty interacting with such services
Solution Approach 1:
The patent implements a universal interface layer that accepts multiple input modalities including voice commands, keyboard inputs, and mouse clicks. This multi-functional interface translates voice commands into equivalent GUI actions, allowing the same service to be accessed through different interaction methods with consistent functionality and precision.
Solution Approach 2:
The patent introduces a mediator component that translates voice commands into GUI-compatible action sequences. This intermediary converts natural language voice inputs into structured commands that the existing GUI system can process, enabling precise control through voice without modifying the underlying website architecture.
3Ease of operation
If a voice assistant uses only voice commands to navigate within a retrieved webpage, then the user can interact hands-free, but navigation would be difficult without also receiving visual information
Solution Approach 1:
The patent segments the information delivery into distinct audio components including navigation status updates, content descriptions, and confirmation messages. This segmented audio feedback provides the user with structured information about their navigation progress and current position on the webpage, improving navigation accuracy while maintaining hands-free operation.
Solution Approach 2:
The patent implements a feedback mechanism where the voice assistant provides continuous audio feedback about navigation actions, current location on the webpage, and available options. This real-time audio feedback loop allows the user to understand their navigation context and make informed decisions without visual information, improving both ease of operation and navigation accuracy.
Data Source
AI summary
Techniques for processing a voice initiated request by a web server are presented. The techniques may include receiving, by a web server, request data representing a voice command to a user device, the request data including an identification of a requested webpage; determining, by the web server, that a response to the request data will continue a voice interaction; and providing, by the web server and to the user device, data for a voice enabled webpage associated with the requested webpage, where the data for the voice enabled webpage is configured to invoke a voice interface for the user device.


