Synthetic Saliency Map Generation for Faster Pedestrian Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current pedestrian detection and localization techniques in autonomous vehicles are inefficient due to the difficulty in training and testing deep learning algorithms, which require extensive and costly sensor data, and struggle to match human-like perception of object scales and locations in a scene.
Innovation Solution
The generation and use of synthetic saliency maps, which involve creating a label image with random points within bounding boxes, applying a Gaussian blur, and using these low-resolution maps to train and test deep neural networks for object detection, reducing the need for extensive sensor data and mimicking human perception.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep learning algorithms are trained using extensive sensor data, then object detection accuracy is improved, but training time and cost increase significantly
Solution Approach 1:
The patent creates synthetic saliency maps that copy and simulate the essential characteristics of human visual fixation patterns. Instead of using extensive real sensor data, the system generates artificial saliency maps that replicate the statistical properties and spatial distributions of human gaze fixations, providing a compressed representation that captures the essential information needed for training object detection algorithms
Solution Approach 2:
The patent transforms the training data representation by changing from raw sensor images to processed saliency maps. This parameter transformation involves computing fixation-based saliency metrics that emphasize regions of interest, effectively changing the data representation to highlight only the most relevant features for object detection, thereby reducing training complexity and time
2Measurement precision
If deep learning algorithms are trained using extensive sensor data, then object detection accuracy is improved, but training cost increases significantly
Solution Approach 1:
The system creates synthetic copies of fixation patterns through algorithmic generation rather than expensive data collection. The synthetic saliency maps replicate the essential statistical properties of human visual attention without requiring actual sensor data acquisition, annotation, and storage infrastructure, significantly reducing training costs
Solution Approach 2:
The patent uses computationally inexpensive synthetic saliency maps that can be generated quickly and discarded after training, replacing the need for expensive, long-term storage and processing of extensive real sensor data. The synthetic data serves its purpose efficiently without the overhead of managing large-scale real-world datasets
3Productivity
If current pedestrian detection techniques are used, then object detection is achieved, but efficiency is reduced due to difficulty in training and testing
Solution Approach 1:
The patent extracts only the most essential visual information by generating saliency maps that highlight regions corresponding to human fixation points. This extraction process removes irrelevant background information and focuses the training data on salient features, simplifying the training and testing processes while improving detection efficiency
Solution Approach 2:
The system changes the data representation parameters from full-resolution sensor images to downsampled saliency maps with reduced spatial dimensions. This parameter transformation reduces computational complexity during training and testing while preserving the essential spatial relationships needed for pedestrian detection
4Loss of time
If synthetic saliency maps with Gaussian blur are used, then training time is reduced, but measurement precision may be affected
Solution Approach 1:
The patent applies Gaussian blur as a parameter transformation that smooths the saliency maps while preserving their essential statistical properties. This blurring operation reduces high-frequency noise and sharp transitions, creating a more robust training representation that maintains accuracy while reducing training time through faster convergence
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly reduces the time and cost of training and testing deep learning algorithms, improves object detection efficiency, and allows for more accurate pedestrian localization by mimicking human perception without exhaustive data collection.
Implementation Method 1
A blur component is configured to apply a blur to the intermediate image to generate a blurred intermediate image
Data Source
AI summary
The disclosure extends to methods, systems, and apparatuses for automated fixation generation and more particularly relates to generation of synthetic saliency maps. A method for generating saliency information includes receiving a first image and an indication of one or more sub-regions within the first image corresponding to one or more objects of interest. The method includes generating and storing a label image by creating an intermediate image having one or more random points. The random points have a first color in regions corresponding to the sub-regions and a remainder of the intermediate image having a second color. Generating and storing the label image further includes applying a Gaussian blur to the intermediate image.


