AI Video Tagging with Image-Audio Group Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video editing processes for broadcasts, movies, and dramas are inefficient due to the large capacity and number of videos, making it difficult to search for specific people or dialogues, and current cloud services do not adequately address this issue.
Innovation Solution
A system and method using artificial intelligence to tag videos by separating image and audio data, clustering person images and speech sections, and matching them using a speaker detection model to determine and tag groups based on probability values and thresholds, improving the efficiency of video editing and search.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If videos are stored in cloud storage to reduce delivery inefficiency, then video delivery efficiency is improved, but video editing time increases due to large capacity and number of videos
Solution Approach 1:
The system performs preliminary actions by automatically tagging videos with metadata (person names, dialogue, objects) before the editing process. This pre-tagging allows editors to quickly search and locate specific content without manually reviewing entire videos, thus reducing editing time while maintaining efficient cloud-based delivery
2Adaptability or versatility
If more videos are produced for broadcasts and movies, then content variety is improved, but search difficulty increases for specific people or dialogues
Solution Approach 1:
The system introduces metadata tags as an intermediary between the video content and the search function. These tags (containing person names, dialogue, objects) serve as searchable indices that enable quick retrieval of specific content across large video collections, solving the search difficulty problem while maintaining content variety
Solution Approach 2:
The system replaces manual search mechanisms with automated AI-based tagging and search algorithms. The speaker detection model and natural language processing automatically generate searchable metadata, eliminating the need for editors to manually search through numerous videos and enabling efficient content retrieval
Data Source
AI summary
Tagging of people appearing in a video can be performed more efficiently by grouping people appearing in the video by utilizing image data and audio data in the video together. It is possible to solve the problems of the prior art that had difficulty in searching for people appearing in a video or to analyze and edit scenes in which these people appear, and to maximize the efficiency of video editing and searching.


