AI Video Tagging with Image-Audio Group Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video editing processes for broadcasts, movies, and dramas are inefficient due to the large capacity and number of videos, making it difficult to search for specific people or dialogues, and current cloud services do not adequately address this issue.

Innovation Solution

A system and method using artificial intelligence to tag videos by separating image and audio data, clustering person images and speech sections, and matching them using a speaker detection model to determine and tag groups based on probability values and thresholds, improving the efficiency of video editing and search.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If videos are stored in cloud storage to reduce delivery inefficiency, then video delivery efficiency is improved, but video editing time increases due to large capacity and number of videos

Engineering Contradiction:
Improvevideo delivery efficiencyVSAvoidvideo editing time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically tagging videos with metadata (person names, dialogue, objects) before the editing process. This pre-tagging allows editors to quickly search and locate specific content without manually reviewing entire videos, thus reducing editing time while maintaining efficient cloud-based delivery

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If more videos are produced for broadcasts and movies, then content variety is improved, but search difficulty increases for specific people or dialogues

Engineering Contradiction:
Improvecontent varietyVSAvoidsearch difficulty
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system introduces metadata tags as an intermediary between the video content and the search function. These tags (containing person names, dialogue, objects) serve as searchable indices that enable quick retrieval of specific content across large video collections, solving the search difficulty problem while maintaining content variety

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces manual search mechanisms with automated AI-based tagging and search algorithms. The speaker detection model and natural language processing automatically generate searchable metadata, eliminating the need for editors to manually search through numerous videos and enabling efficient content retrieval

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20260017947A1System and method for tagging video based on artificial intelligence
Publication Date: 2026.01.15 CLOUDIKE INC
  • US20260017947A1 patent drawing
  • US20260017947A1 patent drawing
  • US20260017947A1 patent drawing

AI summary

Tagging of people appearing in a video can be performed more efficiently by grouping people appearing in the video by utilizing image data and audio data in the video together. It is possible to solve the problems of the prior art that had difficulty in searching for people appearing in a video or to analyze and edit scenes in which these people appear, and to maximize the efficiency of video editing and searching.