Automated Highlight Video Generation Using Deep Learning Event Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid increase in online video content, particularly during the COVID-19 pandemic, has made it time-consuming and costly to manually create highlight videos, which require human effort to edit and condense original videos effectively.

Innovation Solution

A system and method using deep learning models and ensemble learning to automatically and precisely generate highlight videos by detecting key events in videos, such as goals in soccer games, and extracting relevant clips, thereby reducing the need for manual editing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual editing is used to create highlight videos, then the quality and precision of event detection is improved, but the time consumption and cost increase significantly

Engineering Contradiction:
Improveevent detection precisionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical editing with an automated deep learning-based system. The deep learning model processes video data automatically to detect key events, substituting human manual operations with an intelligent automated system that achieves both precision and efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically detecting events, selecting clips, and generating highlight videos without requiring human intervention. The deep learning model independently processes the entire workflow from raw video input to final highlight output, eliminating the need for manual editing.

Inventive Principle:
Principle #25Self-service

2Loss of time

If automated systems are used to generate highlight videos, then the time consumption is reduced, but the precision and accuracy of event detection deteriorates

Engineering Contradiction:
Improvetime consumptionVSAvoidevent detection precision
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent replaces manual mechanical editing with an automated deep learning-based system. The deep learning model processes video data automatically to detect key events, substituting human manual operations with an intelligent automated system that achieves both precision and efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system incorporates feedback mechanisms where the deep learning model continuously refines its event detection based on training data and performance evaluation. This feedback loop enables the automated system to improve its precision over time while maintaining high processing speed.

Inventive Principle:
Principle #23Feedback

3Extent of automation

If deep learning models are used to detect key events, then the automation level and processing speed are improved, but the system complexity increases

Engineering Contradiction:
Improveautomation levelVSAvoidsystem complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the complex video processing task into distinct modules: video input, deep learning event detection, clip selection, and highlight video generation. This modular segmentation makes the complex system more manageable and easier to implement while maintaining high automation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The deep learning model serves multiple functions within the system: it detects events, classifies them, determines timing, and identifies important moments. This multi-functionality reduces the need for separate specialized components, thereby managing system complexity while maintaining high automation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12223720B2Generating highlight video from video and text inputs
Publication Date: 2025.02.11 BAIDU USA LLC
  • US12223720B2 patent drawing
  • US12223720B2 patent drawing
  • US12223720B2 patent drawing

AI summary

Presented herein are systems, methods, and datasets for automatically and precisely generating highlight or summary videos of content. In one or more embodiments, the inputs comprise a text (e.g., an article) of the key event(s) (e.g., a goal, a player action, etc.) in an activity (e.g., a game, a concert, etc.) and a video or videos of the activity. In one or more embodiments, the output is a short video of an event or events in the text, in which the video may include commentary about the highlighted events and/or other audio (e.g., music), which may also be automatically synthesized.