Machine-Based Video Object Recognition with Non-Intrusive Metadata Overlays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video interfaces are obtrusive, unintuitive, and require manual annotation, limiting their interactivity and increasing costs, which hinders broad distribution and user engagement.

Innovation Solution

A machine learning engine automatically generates metadata for video content, and new interactive interfaces seamlessly relay this metadata to users without disrupting the viewing experience, using neural networks for real-time object recognition and intuitive user interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If manual annotation is used to provide detailed video metadata, then information completeness is improved, but labor cost and time consumption increase prohibitively

Engineering Contradiction:
Improvevideo metadata completenessVSAvoidannotation time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical annotation processes with automated machine learning systems. Neural networks and computer vision algorithms automatically generate metadata for video content, eliminating the need for human annotators to manually tag each frame or segment, thereby resolving the contradiction between information completeness and time consumption

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The video content itself serves as the source for generating its own metadata through automated analysis. The system processes video frames, detects objects, scenes, and actions, and generates descriptive metadata autonomously without external human intervention, allowing the content to annotate itself at scale

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If traditional interactive video interfaces are added to provide product information, then user engagement is improved, but interface complexity and disruption to viewing experience increase

Engineering Contradiction:
Improveinteractive functionalityVSAvoidinterface complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent overlays interactive product information in a spatial dimension above the video content rather than requiring separate screens or complex navigational interfaces. Metadata and product details are displayed as translucent overlays synchronized with video playback, adding interactivity without increasing interface complexity or disrupting the viewing experience

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent merges the video playback interface with the metadata display interface into a single unified presentation. Product information, scene descriptions, and interactive elements are combined with the video content in a seamless overlay, eliminating the need for separate complex interface layers and reducing overall system complexity

Inventive Principle:
Principle #5Merging (Combining)

3Speed

If real-time metadata retrieval is implemented to update interfaces frequently, then user experience smoothness is improved, but network bandwidth consumption and server load increase

Engineering Contradiction:
Improveinterface update rateVSAvoidnetwork bandwidth
Core Design Contradiction:
SpeedVSLoss of energy

Solution Approach 1:

The patent performs preliminary processing of video content into discrete segments with pre-generated metadata during offline preparation. This segmentation allows the system to retrieve only relevant metadata portions during playback, enabling frequent interface updates without requiring continuous full-video metadata transmission, thereby reducing network bandwidth consumption

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent divides video content and its associated metadata into temporal and spatial segments that can be independently retrieved and displayed. By segmenting the metadata stream according to video playback position and detecting only changed portions, the system achieves smooth real-time updates while minimizing network traffic and server processing load

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12387256B2Machine-based object recognition of video content
Publication Date: 2025.08.12 PAINTED DOG INC
  • US12387256B2 patent drawing
  • US12387256B2 patent drawing
  • US12387256B2 patent drawing

AI summary

Current interfaces for displaying information about items appearing in videos are obtrusive and counterintuitive. They also rely on annotations, or metadata tags, added by hand to the frames in the video, limiting their ability to display information about items in the videos. In contrast, examples of the systems disclosed here use neural networks to identify items appearing on- and off-screen in response to intuitive user voice queries, touchscreen taps, and/or cursor movements. These systems display information about the on- and off-screen items dynamically and unobtrusively to avoid disrupting the viewing experience.