AI Highlight Video Generation via Deep Learning Event Spotting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid increase in online video content, particularly during the COVID-19 pandemic, has made it time-consuming and costly for humans to manually create highlight videos, which often require precise identification of key events, a task challenging for machines.

Innovation Solution

A system utilizing state-of-the-art deep learning models and an ensemble learning module to automatically generate highlight videos by processing large-scale multimodal datasets, including cloud-sourced text data and untrimmed videos, to precisely locate and extract event clips, reducing processing requirements and enhancing event spotting performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual editing is used to create highlight videos, then video quality and precision can be maintained, but time consumption and cost increase significantly

Engineering Contradiction:
Improveevent spotting precisionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical editing with an automated deep learning-based system. The system uses convolutional neural networks and recurrent neural networks to automatically detect events, extract clips, and assemble highlight videos, eliminating the need for manual human editing while achieving comparable or superior precision in event spotting.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service automated highlight generation where the deep learning model independently processes entire video streams, identifies key events, extracts relevant clips, and assembles final highlight videos without human intervention. The model serves itself by automatically optimizing the highlight selection based on learned patterns from training data.

Inventive Principle:
Principle #25Self-service

2Productivity

If deep learning models are used to automatically generate highlight videos, then productivity increases, but device complexity and computational requirements increase

Engineering Contradiction:
Improvevideo generation efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the complex video processing task into multiple independent segments: video encoding, event detection, clip extraction, and highlight assembly. Each segment is handled by specialized neural network components, allowing modular processing that reduces overall system complexity while maintaining high productivity through parallel computation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-training deep learning models on extensive training data and pre-processing videos into manageable segments. This preparation enables the system to quickly generate highlight videos without requiring complex real-time computations during actual video processing, thus reducing operational complexity while maintaining high efficiency.

Inventive Principle:
Principle #10Preliminary action

3Loss of time

If automated systems are used to process large-scale video content, then time consumption decreases, but measurement precision and accuracy of event identification become more challenging

Engineering Contradiction:
Improveprocessing timeVSAvoidevent localization accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent introduces temporal dimension analysis by using recurrent neural networks that process video frames sequentially over time. This temporal dimension allows the system to accurately identify event boundaries and localize precise moments, achieving high measurement precision while maintaining fast processing speeds through efficient temporal convolution operations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system incorporates feedback mechanisms where the deep learning model continuously refines its event detection based on comparisons with ground truth labels from training data. This feedback loop enables the system to correct localization errors and improve precision automatically, achieving accurate event identification without sacrificing processing speed.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11769327B2Automatically and precisely generating highlight videos with artificial intelligence
Publication Date: 2023.09.26 BAIDU USA LLC
  • US11769327B2 patent drawing
  • US11769327B2 patent drawing
  • US11769327B2 patent drawing

AI summary

Presented herein are systems, methods, and datasets for automatically and precisely generating highlight or summary videos of content. For example, in one or more embodiments, videos of sporting events may be digested or condensed into highlights, which will dramatically benefit sports media, broadcasters, video creators or commentators, or other short video creators, in terms of cost reduction, fast, and mass production, and saving tedious engineering hours. Embodiment of the framework may also be used or adapted for use to better promote sports teams, players, and/or games, and produce stories to glorify the spirit of sports or its players. While presented in the context of sports, it shall be noted that the methodologies may be used for videos comprising other content and events.