Interactive Video Object Recognition and E-commerce Linking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video viewing experiences are passive and lack the ability to identify, interact with, or purchase visual objects within videos, or receive targeted advertisements based on these objects.

Innovation Solution

Systems and methods that encode video streams with metadata to identify and track objects, allowing users to click on them for information, purchase opportunities, or targeted advertisements, using computer vision and machine learning algorithms for object recognition and linking to ecommerce sites or advertisements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If video viewing remains traditional passive mode, then system complexity is low, but user interaction capability and information accessibility are limited

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent embeds multiple layers of interaction within the video stream: object detection layers, metadata embedding layers, user interaction layers (clicking, hovering), and information delivery layers. Each layer is nested within the video playback system, allowing complex functionality to be integrated without disrupting the core video viewing experience

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent introduces metadata as an intermediary between the video content and the user interaction system. This metadata stream carries object identification information, linking visual elements to external information sources, e-commerce databases, and advertising systems without requiring direct complex processing between video pixels and interaction handlers

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If object identification and tracking are added to video streams, then user interaction and information access improve, but processing time and computational resources increase

Engineering Contradiction:
Improveobject information accessibilityVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs object detection, identification, and metadata generation in advance during video encoding or pre-processing. Objects are identified and tracked frame-by-frame before user viewing, with results stored as metadata embedded in the video stream. This eliminates the need for real-time processing during playback, as the system only needs to retrieve and display pre-computed object information when users interact

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts object identification and tracking functionality into a separate pre-processing stage, removing the computational burden from the real-time video playback path. The metadata containing object information is extracted and embedded independently, allowing the main video stream to play back without delay while interaction data is readily available

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If metadata encoding is applied to identify objects in each frame, then object recognition accuracy improves, but data transmission bandwidth and storage requirements increase

Engineering Contradiction:
Improveobject recognition accuracyVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies metadata encoding selectively rather than uniformly across all video data. Metadata is generated and embedded only for frames or regions containing detectable objects of interest, with varying levels of detail based on object importance, user interaction frequency, and relevance to the video content. This reduces overall data volume while maintaining recognition accuracy for critical objects

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the video stream into distinct components: the primary video data and the separate metadata stream. The metadata is further segmented by object, frame, and interaction type, allowing efficient compression and selective transmission. Only necessary object information (identification, location, relevant attributes) is encoded, rather than complete frame analysis data

Inventive Principle:
Principle #1Segmentation

4Productivity

If e-commerce links and advertisements are integrated with video objects, then commercial value and user engagement improve, but system complexity and advertising relevance challenges increase

Engineering Contradiction:
Improvecommercial valueVSAvoidsystem integration complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates a universal metadata structure that serves multiple functions simultaneously: object identification for interaction, product database linking for e-commerce, categorization for advertising targeting, and content analysis for recommendations. This single metadata system supports diverse commercial applications without requiring separate processing pipelines for each function

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements feedback loops where user interactions with objects (clicks, hovers, purchases) are tracked and used to refine advertising delivery and e-commerce link effectiveness. This feedback mechanism continuously optimizes the commercial value extraction while maintaining relevance, using accumulated interaction data to improve future ad targeting and product recommendations

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10299011B2Method and system for user interaction with objects in a video linked to internet-accessible information about the objects
Publication Date: 2019.05.21 GRUSD BRANDON
  • US10299011B2 patent drawing
  • US10299011B2 patent drawing
  • US10299011B2 patent drawing

AI summary

An interactive video including frames which include objects is displayed on a client computing device. A set of the objects in the interactive video has been linked to internet-accessible information external to the video during creation of the interactive video by comparing each of the objects in the interactive video with pre-defined objects stored in a database. The object is associated with internet-accessible information associated with the pre-defined objects when the object is determined to be similar to the pre-defined object. While the interactive video is playing on the display, a selection of one of the objects shown in the interactive video is received. In response to the selection, internet-accessible information linked to the selected object is displayed, where the internet-accessible information includes at least one of a link to an online e-commerce site that sells the selected object, and an advertisement associated with the selected object.