Video Processing Combined Saliency Map Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video processing methods generate inconsistent saliency maps for identifying salient objects in video streams, as they rely on different visual characteristics and do not effectively combine spatial and spatio-temporal features, leading to varying identifications of salient objects across different types of saliency maps.
Innovation Solution
A method and system for video processing that generates a combined saliency map by weighting and combining spatial and spatio-temporal saliency maps, using pre-stored saliency maps to determine optimal weights that minimize the average difference between computed and stored saliency values, thereby ensuring consistent identification of salient objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If different types of saliency maps (spatial, spatio-temporal, ground truth) are generated independently, then each map can be computed using its specific visual characteristics, but the identified salient objects differ across maps leading to inconsistency
Solution Approach 1:
The patent combines multiple independently generated saliency maps (spatial, spatio-temporal, and ground truth maps) into a single consolidated saliency map. This merging process integrates the strengths of each map type while eliminating inconsistencies, allowing the system to maintain adaptability to different visual characteristics while achieving reliable and consistent salient object identification across all map types.
2Reliability
If multiple saliency maps are generated and combined, then consistent salient object identification can be achieved, but the processing complexity and computational requirements increase
Solution Approach 1:
The patent segments the saliency map generation process into distinct independent stages: generating spatial saliency maps, generating spatio-temporal saliency maps, and generating ground truth saliency maps. Each stage processes specific visual characteristics separately. The segmentation is followed by a consolidation stage that combines these segmented results, thereby reducing processing complexity at each stage while maintaining the reliability benefits of multiple map types.
Data Source
AI summary
Various aspects of a system and method for video processing is disclosed herein. The system comprises a video processing device that is configured to generate a spatial saliency map based on spatial information associated with a current frame of a video stream. A spatio-temporal saliency map is generated based on at least motion information associated with the current frame and a previous frame of the video stream. Based on a weighted combination of the generated spatial saliency map and the generated spatio-temporal saliency map, a combined saliency map is generated.


