Simultaneous Temporal Attention and Action Prediction for Long Videos

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in automatically extracting relevant information from video content, particularly in identifying specific events and actions within long video recordings, such as sports games, due to the high bit representation and time-consuming manual analysis.

Innovation Solution

A method and system for predicting temporal attention regions and action types in video clips using machine learning, involving the generation of training data with spatial and temporal attention regions, and simultaneous training of models to identify these regions and classify associated actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis is used to identify events in video recordings, then accuracy of event detection can be high, but time consumption increases significantly

Engineering Contradiction:
Improveaccuracy of event detectionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical analysis with an automated machine learning system. The system uses trained models to automatically detect temporal attention zones and classify action types in video clips, substituting human observers with computational algorithms that process video data efficiently while maintaining high detection accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the video analysis system to automatically identify and classify events without human intervention. The trained machine learning models independently process video clips, detect attention zones, and categorize actions, making the system autonomous and eliminating the need for manual analysis.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated signal processing is used to extract information from video, then time efficiency improves, but measurement precision and reliability decrease

Engineering Contradiction:
Improvespeed of information extractionVSAvoidaccuracy of event identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by pre-training machine learning models using extensively labeled training data before deployment. The models are prepared in advance with learned patterns and features from diverse video content, enabling them to accurately identify events and actions in new video clips without requiring manual analysis during actual operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates feedback mechanisms where the machine learning models continuously improve their performance based on training data and evaluation metrics. The models are trained using labeled examples and refined through iterative optimization, allowing them to maintain high precision while operating automatically at scale.

Inventive Principle:
Principle #23Feedback

3Loss of information

If comprehensive video analysis is performed on long video clips, then complete information extraction is achieved, but computational complexity and resource requirements increase

Engineering Contradiction:
Improvecompleteness of information extractionVSAvoidcomputational complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts only the most relevant information from video clips by identifying temporal attention zones that contain meaningful events. Instead of analyzing entire long videos, the system detects and extracts specific segments containing actions of interest, significantly reducing computational requirements while maintaining completeness of relevant information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments long video clips into manageable temporal attention zones based on detected events and actions. By dividing the video content into relevant segments rather than processing the entire video continuously, the computational complexity is reduced while ensuring all important information is captured in the identified zones.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250272975A1System and method for simultaneous temporal attention zone and action type prediction and applications thereof
Publication Date: 2025.08.28 YAHOO ASSETS LLC
  • US20250272975A1 patent drawing
  • US20250272975A1 patent drawing
  • US20250272975A1 patent drawing

AI summary

The present teaching relates to predicting a temporal attention region corresponding to an event of interest in a video clip. Training data is obtained with training samples, each of which includes a historic video clip with a temporal attention region in consecutive frames to represent an event of interest captured in the temporal attention region and is used for training, via machine learning, a temporal attention zone model for predicting a temporal attention zone in a video clip representing an event of interest. The trained model is used to predict, from an input video clip, a temporal attention zone represented by consecutive frames in the input video clip that capture the event of interest.