Video Clipping via Object Detection and Character Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video clipping methods rely heavily on manual labor, resulting in low efficiency and high costs, especially when dealing with vertical video clipping that requires complex operations like identifying and isolating specific characters in videos.

Innovation Solution

A method and system that perform object detection, classification, and similarity calculation using pre-trained models to automatically identify and synthesize human body region images from videos, where the similarity with a target character image exceeds a preset threshold, thereby creating a clipping video without manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual labor is used for video clipping, then the clipping process can be performed with simple tools, but the efficiency and productivity are low

Engineering Contradiction:
Improvevideo clipping efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs automatic object detection, classification, and similarity calculation without requiring manual intervention. The pre-trained models automatically identify human body regions and calculate similarity to target characters, enabling the system to serve itself and eliminate the need for manual clipping operations.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical operations are replaced with automated computational processes. The patent substitutes manual video clipping operations with computer-based object detection, classification, and similarity calculation systems that automatically identify and extract target characters from videos.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If pre-trained models are used for object detection and classification, then the accuracy of character identification is improved, but the computational time and processing complexity increase

Engineering Contradiction:
Improvecharacter identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses pre-trained models that have already been trained on large datasets before deployment. The object detection model and classification model are trained in advance, so during actual video clipping operations, they can quickly and accurately identify target characters without requiring real-time training, thus reducing processing time while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If similarity calculation is performed between human body region images and target character images, then the precision of character selection is improved, but the computational energy consumption increases

Engineering Contradiction:
Improvecharacter selection precisionVSAvoidcomputational energy consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The system extracts only the necessary human body region images from video frames using object detection and classification, then performs similarity calculation only on these extracted regions rather than processing entire frames. This extraction approach reduces the amount of data requiring computational energy while maintaining precise character selection through targeted similarity comparison.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11495264B2Method and system of clipping a video, computing device, and computer storage medium
Publication Date: 2022.11.08 SHANGHAI BILIBILI TECH CO LTD
  • US11495264B2 patent drawing
  • US11495264B2 patent drawing
  • US11495264B2 patent drawing

AI summary

Embodiments of the present disclosure describes techniques for clipping a video. The disclosed techniques comprise obtaining a video including a plurality of frames performing object detection on each frame; identifying objects contained in each frame, wherein a region where each object is located is selected through a detection box; classifying and recognizing the objects identified in each frame using a pre-trained classification model; selecting human body region images; determining a similarity between each human body region image selected from the plurality of frames and a target character image; in response to determining that a similarity between a human body region image and the target character image is greater than a predetermined threshold, identifying the human body region image as a clipping image; and synthesizing clipping images identified in the plurality of frames in order of time to obtain a clipping video.