Digital Video Tagging via Neural Feature Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional digital video tagging systems are inefficient and inaccurate in generating tags for digital videos, particularly in identifying and tagging actions, objects, and attributes across multiple frames, leading to difficulties in searching and organizing large collections of videos.

Innovation Solution

A digital video tagging system that uses machine learning to automatically identify actions, objects, and attributes by generating tagged feature vectors from frames of videos, utilizing neural networks to associate metadata with feature vectors and aggregate tags for efficient searching and retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional digital video tagging systems are used, then video tagging can be performed, but the tagging accuracy and efficiency are insufficient

Engineering Contradiction:
Improvetagging accuracyVSAvoidtagging efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The video is divided into multiple frames, and each frame is processed independently to generate feature vectors. This segmentation allows the system to capture temporal dynamics and action information from individual frames while maintaining overall video context through aggregation of frame-level features.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A neural network serves as an intermediary that transforms raw video frame data into meaningful feature vectors. The neural network processes visual information from multiple frames and generates tagged feature vectors that represent actions, objects, and attributes, bridging the gap between raw video data and searchable tags.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual video searching is performed, then users can find specific videos, but it requires viewing the entire video which is time-consuming

Engineering Contradiction:
Improvevideo retrieval accuracyVSAvoidsearch time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary tagging of videos by analyzing frames, generating feature vectors, and creating action tags before the user needs to search. This advance processing creates an indexed database of video content that enables rapid retrieval without requiring users to watch entire videos.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual video viewing and searching with an automated computer vision system that uses neural networks to analyze video content and generate searchable tags. This substitution of mechanical human search with automated algorithmic processing dramatically reduces search time while maintaining retrieval accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If simple tagging methods are used, then processing is fast, but the tags cannot accurately describe actions and attributes

Engineering Contradiction:
Improvetag description accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system transitions from simple frame-level analysis to a multi-dimensional approach by generating feature vectors that capture spatial, temporal, and semantic dimensions of video content. The tagged feature vectors incorporate information from multiple frames and dimensions, enabling accurate action and attribute description while managing complexity through structured processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11949964B2Generating action tags for digital videos
Publication Date: 2024.04.02 ADOBE INC
  • US11949964B2 patent drawing
  • US11949964B2 patent drawing
  • US11949964B2 patent drawing

AI summary

Systems, methods, and non-transitory computer-readable media are disclosed for automatic tagging of videos. In particular, in one or more embodiments, the disclosed systems generate a set of tagged feature vectors (e.g., tagged feature vectors based on action-rich digital videos) to utilize to generate tags for an input digital video. For instance, the disclosed systems can extract a set of frames for the input digital video and generate feature vectors from the set of frames. In some embodiments, the disclosed systems generate aggregated feature vectors from the feature vectors. Furthermore, the disclosed systems can utilize the feature vectors (or aggregated feature vectors) to identify similar tagged feature vectors from the set of tagged feature vectors. Additionally, the disclosed systems can generate a set of tags for the input digital videos by aggregating one or more tags corresponding to identified similar tagged feature vectors.