Machine Learning Highlight Extraction for Short Video Creation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for short-form video content in the single-person broadcasting market necessitates efficient methods for creating high-quality short-form video content, as single creators face inconvenience in editing and producing engaging content.
Innovation Solution
A video automatic editing system utilizing a pre-trained machine learning-based highlight extraction model to acquire input video, extract highlight frames based on calculated scores, and generate highlight videos, incorporating frame information such as expression, movement, speech, and viewer reactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If single creators manually edit video content to create short-form videos, then the quality and engagement of video content can be maintained, but the time consumption and editing complexity increase significantly
Solution Approach 1:
The system enables automatic video editing through self-service mechanisms where the editing algorithm autonomously identifies highlight frames, selects appropriate clips, and assembles short-form videos without requiring manual creator intervention for each editing decision
Solution Approach 2:
Manual editing operations are replaced by an automated machine learning-based editing system that uses deep learning models to perform frame analysis, highlight detection, and video assembly tasks that previously required human creators to perform manually
2Reliability
If all frames of filmed video are included without editing to create long-form content, then the completeness of the video content is preserved, but the video length becomes excessively long and less engaging
Solution Approach 1:
The system extracts only the essential and engaging portions of the original video content by identifying and selecting highlight frames that capture key moments, thereby removing redundant or less important segments while preserving the core value of the content
Solution Approach 2:
Instead of including all video frames, the system applies partial action by selectively including only the most important frames and clips that convey the essential content, achieving effective communication with reduced video length
3Measurement precision
If a highlight extraction model uses multiple frame information parameters (expression, movement, speech, etc.) to calculate scores, then the accuracy of highlight frame identification improves, but the computational complexity increases
Solution Approach 1:
The complex analysis task is segmented into multiple independent modules, each handling a specific type of frame information (expression analysis, movement detection, speech recognition, etc.), allowing parallel processing and reducing overall computational complexity while maintaining high accuracy
Data Source
AI summary
Disclosed are a video automatic editing method and system based on machine learning. The video automatic editing system based on machine learning includes at least one processor, and the at least one processor includes a video acquirer configured to acquire input video, a highlight frame extractor configured to extract at least one highlight frame from the input video using a highlight extraction model pre-trained through machine learning, and a highlight video generator configured to generate highlight video from the at least one extracted highlight frame.


