Video Item Identification via Voice Command Timestamp Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice services require users to explicitly identify items by name, brand, and details when ordering, which can be cumbersome and time-consuming, especially when the item is not accurately described or not available, leading to inefficiencies in the ordering process.

Innovation Solution

A voice command system that allows users to order items seen in videos by uttering a simple command like 'Buy it' without specifying the item, using voice technology to identify the item in real-time based on the video content being streamed, enabling users to purchase items without knowing their details.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users explicitly identify items by name, brand, and details when ordering, then the ordering accuracy is improved, but the ordering process becomes time-consuming and cumbersome

Engineering Contradiction:
Improveitem identification accuracyVSAvoidordering process time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary identification of items in the video content before the user needs to order. By pre-processing the video to identify and tag items with their details (name, brand, specifications), the system prepares the information in advance so that users can simply reference items by casual description rather than providing complete identification details during the ordering process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary system (the voice service with video analysis capabilities) that bridges the gap between the user's simple item description and the detailed item identification requirements. This intermediary automatically matches the user's casual reference to the pre-identified item database, extracting full item details without requiring the user to manually provide them.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If users provide detailed item descriptions, then the item identification accuracy is improved, but the ease of operation deteriorates

Engineering Contradiction:
Improveitem identification accuracyVSAvoidordering operation simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system enables self-service by automatically performing the item identification and detail extraction functions that would otherwise require user input. The voice service autonomously analyzes the user's simple reference, cross-references it with the pre-identified items in the video, and completes the item selection without requiring the user to manually provide detailed descriptions.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The intermediary system translates between the user's simplified item reference and the detailed item identification requirements. It acts as a mediator that handles the complexity of item detail extraction and matching, allowing users to operate the system with simple commands while maintaining high identification accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If the voice service requires specific item knowledge from users, then the ordering precision is improved, but the adaptability deteriorates

Engineering Contradiction:
Improveordering accuracyVSAvoiduser vocabulary requirement
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system achieves universality by handling multiple types of user inputs (simple descriptions, brand names, partial identifiers) through a single unified process. The pre identification technology creates a comprehensive item database that can be accessed through various reference methods, making the system adaptable to users with different levels of product knowledge while maintaining accurate item identification.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11688388B1Utterance request of items as seen within video
Publication Date: 2023.06.27 AMAZON TECH INC
  • US11688388B1 patent drawing
  • US11688388B1 patent drawing
  • US11688388B1 patent drawing

AI summary

Techniques are described for fulfilling an utterance request for an item represented within a video rendered at a client device. In some implementations, a user account associated with the request is identified, enabling a video stream transmitted in association with the user account at the time that the request was uttered to be identified. In one technique, a timestamp associated with the request is used to identify the relevant portion of the video stream. The item represented within the portion of the video stream can be identified using various techniques and/or information such as image recognition, metadata within the video, subtitles, closed captions, and/or a database mapping between the item and a video content item transmitted in the video stream.