Synchronized Audiovisual Response System for Voice Queries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice-enabled devices struggle to provide an interactive and accurate response to user queries, especially when these queries relate to visual content, as they lack the ability to synchronize audio and visual information effectively, leading to incomplete or misinterpreted user requests.

Innovation Solution

A system that generates a dynamic audiovisual presentation in response to user queries, using a service provider system that processes both spoken commands and visual content data to create a synchronized audio and visual experience, allowing users to interact with item details through voice or touchscreen inputs, and enabling actions like purchasing items directly from the presentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice-enabled devices provide audio responses to user queries, then user interaction convenience is improved, but accuracy and completeness of information delivery deteriorates

Engineering Contradiction:
Improveuser interaction convenienceVSAvoidinformation completeness
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent combines audio and visual response channels into a unified presentation system. The voice-enabled device simultaneously generates audio responses and displays corresponding visual content (images, text, graphics) to deliver comprehensive information. This merging of sensory channels allows the device to maintain conversational ease while compensating for the information limitations of speech alone through synchronized visual supplementation.

Inventive Principle:
Principle #5Merging (Combining)

2Device complexity

If voice-enabled devices process only spoken commands, then device complexity is reduced, but ability to interpret visual content deteriorates

Engineering Contradiction:
Improveprocessing capabilityVSAvoidquery interpretation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements a multi-functional response system that handles both audio and visual output through a single integrated architecture. The voice-enabled device processes spoken commands and generates unified presentations that incorporate multiple media types (audio speech, displayed images, text, and graphics). This universal approach allows the device to interpret visual content accurately while maintaining manageable complexity through consolidated processing logic rather than separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If voice-enabled devices provide detailed information, then information completeness is improved, but user attention and comprehension deteriorates

Engineering Contradiction:
Improveinformation completenessVSAvoiduser comprehension
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments detailed information into distinct visual and audio components that are synchronized and presented in an organized manner. Rather than delivering a continuous stream of detailed verbal information that may overwhelm the user, the system divides content into discrete elements (visual displays with images and text, audio explanations) that can be processed separately but are coordinated together. This segmentation allows complete information delivery while improving user comprehension through structured presentation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10896457B2Synchronized audiovisual responses to user requests
Publication Date: 2021.01.19 AMAZON TECH INC
  • US10896457B2 patent drawing
  • US10896457B2 patent drawing
  • US10896457B2 patent drawing

AI summary

Systems and methods are disclosed related to presenting an interactive audiovisual presentation that provides a user with information regarding individual items matching a user's search request or other request that results in a set of items. The audiovisual content may be generated to include a summary of a subset of item attributes associated with the individual items, and may include both an audio summary and visual content that are presented in synchronization with each other based on markup information in a presentation file.