Machine Learning Highlight Extraction for Short Video Creation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing demand for short-form video content in the single-person broadcasting market necessitates efficient methods for creating high-quality short-form video content, as single creators face inconvenience in editing and producing engaging content.

Innovation Solution

A video automatic editing system utilizing a pre-trained machine learning-based highlight extraction model to acquire input video, extract highlight frames based on calculated scores, and generate highlight videos, incorporating frame information such as expression, movement, speech, and viewer reactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If single creators manually edit video content to create short-form videos, then the quality and engagement of video content can be maintained, but the time consumption and editing complexity increase significantly

Engineering Contradiction:
Improvevideo creation efficiencyVSAvoidediting time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system enables automatic video editing through self-service mechanisms where the editing algorithm autonomously identifies highlight frames, selects appropriate clips, and assembles short-form videos without requiring manual creator intervention for each editing decision

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual editing operations are replaced by an automated machine learning-based editing system that uses deep learning models to perform frame analysis, highlight detection, and video assembly tasks that previously required human creators to perform manually

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If all frames of filmed video are included without editing to create long-form content, then the completeness of the video content is preserved, but the video length becomes excessively long and less engaging

Engineering Contradiction:
Improvecontent completenessVSAvoidvideo length
Core Design Contradiction:
ReliabilityVSLength of moving object

Solution Approach 1:

The system extracts only the essential and engaging portions of the original video content by identifying and selecting highlight frames that capture key moments, thereby removing redundant or less important segments while preserving the core value of the content

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of including all video frames, the system applies partial action by selectively including only the most important frames and clips that convey the essential content, achieving effective communication with reduced video length

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If a highlight extraction model uses multiple frame information parameters (expression, movement, speech, etc.) to calculate scores, then the accuracy of highlight frame identification improves, but the computational complexity increases

Engineering Contradiction:
Improvehighlight frame identification accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The complex analysis task is segmented into multiple independent modules, each handling a specific type of frame information (expression analysis, movement detection, speech recognition, etc.), allowing parallel processing and reducing overall computational complexity while maintaining high accuracy

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11615814B2Video automatic editing method and system based on machine learning
Publication Date: 2023.03.28 SMSYSTEMS CO LTD
  • US11615814B2 patent drawing
  • US11615814B2 patent drawing
  • US11615814B2 patent drawing

AI summary

Disclosed are a video automatic editing method and system based on machine learning. The video automatic editing system based on machine learning includes at least one processor, and the at least one processor includes a video acquirer configured to acquire input video, a highlight frame extractor configured to extract at least one highlight frame from the input video using a highlight extraction model pre-trained through machine learning, and a highlight video generator configured to generate highlight video from the at least one extracted highlight frame.