Video Clip Positioning via Candidate Sub-Clip Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video clip positioning methods based on text information are hindered by cumbersome and time-consuming data labeling, leading to inaccurate training results and reduced accuracy in video clip positioning.
Innovation Solution
A video clip positioning method that determines a candidate clip from a target video based on video frames and a target text, further divides the candidate clip into sub-clips, and selects the sub-clip with the highest matching degree to the target text as the target video clip.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data labeling is performed manually to train the video recognition model, then the model can learn boundary features of video clips, but the labeling process is cumbersome, time-consuming, and has low precision, resulting in long training periods and suboptimal positioning accuracy
Solution Approach 1:
The patent applies preliminary action by pre-defining candidate clip ranges based on text information before model training. The system first determines candidate clips that may contain the target content, then only performs fine-grained boundary detection within these candidates. This preliminary filtering reduces the search space and training complexity, enabling faster training without sacrificing positioning accuracy.
Solution Approach 2:
The patent segments the video clip positioning task into two distinct stages: coarse-grained candidate clip selection and fine-grained boundary detection. By dividing the overall task, the system can use different strategies for each stage - efficient filtering for candidates and precise measurement for boundaries - thereby reducing overall training time while maintaining high accuracy in the final positioning result.
2Measurement precision
If manual labeling is used to provide precise boundary information for training, then the video recognition model can achieve better positioning accuracy, but the labeling process becomes more complex and time-consuming
Solution Approach 1:
The system performs preliminary filtering to identify candidate clips that are likely to contain the target content based on text matching. This preliminary action reduces the number of clips that require precise boundary labeling, thereby reducing labeling complexity while still achieving high boundary precision for the final target clip through focused fine-grained detection on the reduced candidate set.
3Reliability
If the video recognition model is trained with extensively labeled data to improve positioning accuracy, then the model performance increases, but the data preparation time and computational resources increase significantly
Solution Approach 1:
The patent segments the training data into candidate clips and final target clips. Only candidate clips require relatively loose labeling for initial model training, while fine-grained boundary labeling is only needed for a smaller subset of confirmed target clips. This segmentation allows the model to achieve reliable positioning performance with less extensively labeled data, improving training efficiency without sacrificing reliability.
Solution Approach 2:
The system performs preliminary candidate identification before final target clip determination. This preliminary action allows the model to be trained on a larger set of candidate clips with less stringent labeling requirements, then refined on smaller target clip subsets. This two-stage approach improves training efficiency while maintaining high positioning reliability through the refinement process.
Data Source
AI summary
This application discloses a video clip positioning method performed at a computer device. In this application, the computer device acquires a plurality of video frame features of a target video and a text feature of a target text using a video recognition model to determine a candidate clip that can be matched with the target text. The candidate clip is finely divided based on a degree of matching between a video frame in the candidate clip and the target text to acquire a plurality of sub-clips, and a sub-clip that has the highest degree of matching with the target text is used as a target video clip. According to this application, the video recognition model does not need to learn a boundary feature of the target video clip, and during model training, or precisely label a sample video, thereby shortening a training period of the video recognition model.


