Saliency Map Creation Using Human Visual System Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for creating saliency maps in images rely on non-psycho visual features, which do not effectively mimic the human visual system, limiting their ability to detect salient points efficiently.

Innovation Solution

A method that creates a temporal saliency map by decomposing images into frequential sub-bands, estimating movement, and combining this with spatial saliency maps using a model based on the human visual system, including steps like perceptual sub-band decomposition, masking, and weighting with the maximum pursuit velocity of the eye.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If non-psycho visual features are used for saliency map creation, then the computational process is simpler, but the detection accuracy of salient points deteriorates

Engineering Contradiction:
Improvecomputational simplicityVSAvoiddetection accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transforms the saliency detection approach by changing the parameter basis from non-psycho visual features to psycho-visual features that model human visual system characteristics. This includes using perceptual sub-band decomposition, masking functions, and motion estimation weighted by maximum pursuit velocity, thereby improving detection accuracy while maintaining computational feasibility through efficient algorithms.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If psycho-visual features based on human visual system are used, then the detection accuracy of salient points is improved, but the device complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by decomposing the image into perceptual sub-bands based on human visual system characteristics. This breakdown allows complex psycho-visual processing to be divided into manageable stages: frequency decomposition, masking application, motion estimation, and saliency map generation, reducing overall processing complexity while maintaining high detection accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-decomposing images into perceptual sub-bands and pre-computing masking functions before saliency detection. Motion estimation is also performed in advance with weighting by maximum pursuit velocity, preparing processed data that simplifies the final saliency map generation step and reduces real-time processing complexity.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If hierarchical decomposition into frequential sub-bands is performed, then the temporal saliency detection is improved, but the processing time increases

Engineering Contradiction:
Improvetemporal saliency detectionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent employs periodic action through its multi-resolution pyramid structure, where processing is organized into discrete levels (full resolution, half resolution, quarter resolution, eighth resolution). This hierarchical periodic processing allows temporal saliency detection to be performed systematically at different scales, improving detection precision while managing processing time through efficient multi-scale analysis.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS8416992B2Device and method for creating a saliency map of an image
Publication Date: 2013.04.09 INTERDIGITAL VC HOLDINGS INC
  • US8416992B2 patent drawing
  • US8416992B2 patent drawing
  • US8416992B2 patent drawing

AI summary

Detection of the salient points in an image enable the improvement of further steps such as coding or image indexing, watermarking, video quality estimation. The methods rely on the fact that a model is fully based on the human visual system (HVS) such as the computation of early visual features, and the methods compute a saliency map for video images taking into account motion and the velocity of the eye.