Video Clipping via Object Detection and Character Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video clipping methods rely heavily on manual labor, resulting in low efficiency and high costs, especially when dealing with vertical video clipping that requires complex operations like identifying and isolating specific characters in videos.
Innovation Solution
A method and system that perform object detection, classification, and similarity calculation using pre-trained models to automatically identify and synthesize human body region images from videos, where the similarity with a target character image exceeds a preset threshold, thereby creating a clipping video without manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual labor is used for video clipping, then the clipping process can be performed with simple tools, but the efficiency and productivity are low
Solution Approach 1:
The system performs automatic object detection, classification, and similarity calculation without requiring manual intervention. The pre-trained models automatically identify human body regions and calculate similarity to target characters, enabling the system to serve itself and eliminate the need for manual clipping operations.
Solution Approach 2:
Manual mechanical operations are replaced with automated computational processes. The patent substitutes manual video clipping operations with computer-based object detection, classification, and similarity calculation systems that automatically identify and extract target characters from videos.
2Measurement precision
If pre-trained models are used for object detection and classification, then the accuracy of character identification is improved, but the computational time and processing complexity increase
Solution Approach 1:
The system uses pre-trained models that have already been trained on large datasets before deployment. The object detection model and classification model are trained in advance, so during actual video clipping operations, they can quickly and accurately identify target characters without requiring real-time training, thus reducing processing time while maintaining high accuracy.
3Manufacturing precision
If similarity calculation is performed between human body region images and target character images, then the precision of character selection is improved, but the computational energy consumption increases
Solution Approach 1:
The system extracts only the necessary human body region images from video frames using object detection and classification, then performs similarity calculation only on these extracted regions rather than processing entire frames. This extraction approach reduces the amount of data requiring computational energy while maintaining precise character selection through targeted similarity comparison.
Data Source
AI summary
Embodiments of the present disclosure describes techniques for clipping a video. The disclosed techniques comprise obtaining a video including a plurality of frames performing object detection on each frame; identifying objects contained in each frame, wherein a region where each object is located is selected through a detection box; classifying and recognizing the objects identified in each frame using a pre-trained classification model; selecting human body region images; determining a similarity between each human body region image selected from the plurality of frames and a target character image; in response to determining that a similarity between a human body region image and the target character image is greater than a predetermined threshold, identifying the human body region image as a clipping image; and synthesizing clipping images identified in the plurality of frames in order of time to obtain a clipping video.


