Object Attention Maps for Precise Video Event Feature Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video analysis techniques struggle to easily detect features attributable to specific objects in videos, such as games, due to the complexity of multiple objects and events.
Innovation Solution
An image analysis system utilizing a machine learning model to generate attention maps and tokens for specific objects, combined with an event prediction mechanism, to facilitate easier detection of features and events in video content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual observation or related methods are used to detect important events in videos, then events can be detected, but it is not easy to detect features attributable to a specific object due to multiple objects in the video
Solution Approach 1:
The patent segments the video analysis task by generating separate attention maps for each object type (player, ball, goal). This segmentation allows the system to focus on object-specific features independently, resolving the difficulty of detecting features attributable to specific objects among multiple objects in the video.
Solution Approach 2:
The patent introduces attention maps as an intermediary between the input video and event detection. These attention maps highlight regions associated with specific objects, serving as a mediator that enables precise detection of object-specific features without being overwhelmed by multiple objects in the scene.
2Adaptability or versatility
If multiple objects are present in video images, then rich video content is captured, but detecting features focused on a specific object becomes more difficult
Solution Approach 1:
The system segments the analysis by creating distinct attention maps for different object types (player, ball, goal). This allows the system to maintain versatility in handling multiple objects while achieving precision in detecting features for each specific object type through dedicated attention mechanisms.
Solution Approach 2:
The patent applies local quality by generating attention maps with different characteristics for different object types. Each attention map is optimized to highlight features specific to its target object type, enabling the system to adapt to multiple objects while maintaining high detection accuracy for each specific object.
Data Source
AI summary
Disclosed herein is an image analysis system including a machine learning model configured to receive input of an image and to output an image feature quantity and a map source for image analysis, a map generation section configured to generate, on the basis of a vector corresponding to an object and the output map source, an attention map indicative of a region associated with the object in the image, and a token generation section configured to generate a token indicative of a feature associated with an event of the object on the basis of the generated attention map and the image feature quantity.


