Transformer Temporal Detection for Video Event Spotting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid increase in online video content, particularly during the COVID-19 pandemic, has highlighted the need for automated systems to efficiently generate highlight videos, as manual editing is time-consuming and costly, and existing technologies struggle to accurately identify key events in videos.

Innovation Solution

A system utilizing state-of-the-art deep learning models and ensemble learning to detect events of interest in videos, such as goals in soccer games, by leveraging large-scale multimodal datasets, cloud-sourced text data, and untrimmed video analysis, which includes game time anchoring, feature extraction, and temporal localization to generate precise highlight clips.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual editing is used to create highlight videos, then the quality and precision of event identification is improved, but the time consumption and cost increase significantly

Engineering Contradiction:
Improveevent identification precisionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual editing with an automated deep learning-based system that uses convolutional neural networks and recurrent neural networks to detect and segment key events in videos. The system processes video data through multiple layers of neural networks to automatically identify events of interest, eliminating the need for human editors while maintaining high precision in event detection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically processing and analyzing video content without requiring human intervention. The deep learning models autonomously detect events, determine their temporal locations, and generate highlight clips, allowing the system to serve itself rather than requiring manual editing operations.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated systems are used to generate highlight videos, then the productivity is improved, but the accuracy of event detection deteriorates

Engineering Contradiction:
Improvevideo processing efficiencyVSAvoidevent detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent employs a composite approach by combining multiple types of neural networks (convolutional neural networks for spatial feature extraction and recurrent neural networks for temporal pattern recognition) within a single automated system. This composite architecture enables the system to achieve both high productivity through automation and high accuracy through the synergistic combination of different processing techniques.

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The system segments the video processing task into multiple stages: feature extraction using convolutional neural networks, temporal modeling using recurrent neural networks, event detection, and highlight generation. This segmentation allows each component to specialize in specific aspects of the task, improving overall accuracy while maintaining automated processing efficiency.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If deep learning models are used for event detection, then the accuracy is improved, but the computational complexity and resource requirements increase

Engineering Contradiction:
Improveevent detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the complex deep learning system into modular components: a convolutional neural network module for spatial feature extraction, a recurrent neural network module for temporal analysis, and separate processing stages for event detection and highlight generation. This segmentation reduces overall system complexity by making each component more manageable and easier to implement while maintaining high detection accuracy.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12148214B2Transformer-based temporal detection in video
Publication Date: 2024.11.19 BAIDU USA LLC
  • US12148214B2 patent drawing
  • US12148214B2 patent drawing
  • US12148214B2 patent drawing

AI summary

With rapidly evolving technologies and emerging tools, sports-related videos generated online are rapidly increasing. To automate the sports video editing/highlight generation process, a key task is to precisely recognize and locate events-of-interest in videos. Embodiments herein comprise a two-stage paradigm to detect categories of events and when these events happen in videos. In one or more embodiments, multiple action recognition models extract high-level semantic features, and a transformer-based temporal detection module locates target events. These novel approaches achieved state-of-the-art performance in both action spotting and replay grounding. While presented in the context of sports, it shall be noted that the systems and methods herein may be used for videos comprising other content and events.