Semantic Class Localization via Activation Relevancy Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional machine learning techniques for image localization require manual input from users to generate bounding boxes, which is expensive, inefficient, and often inaccurate, limiting the accuracy and availability of localization and subsequent image functionality.
Innovation Solution
The use of machine learning techniques that learn patterns through neural networks to classify and localize semantic classes within images, employing activation relevancy maps and contrastive attention maps to identify relevant image portions without manual input, thereby improving localization accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual bounding box drawing is used for training data generation, then localization can be performed, but the process becomes expensive, inefficient, and inaccurate
Solution Approach 1:
The system uses automatically generated synthetic training data with precise bounding boxes created through algorithmic processes rather than manual drawing. The synthetic data generation pipeline automatically creates images with objects and their corresponding bounding boxes, eliminating the need for human annotators while maintaining high localization accuracy.
Solution Approach 2:
The patent creates synthetic training data by generating artificial images that copy the essential characteristics of real images but with perfectly accurate bounding boxes. These synthetic copies serve as training examples that teach the localization model without requiring manual annotation of real photographs.
2Reliability
If thousands of training images are used to train a single model, then model accuracy improves, but the complexity and resource requirements increase significantly
Solution Approach 1:
The patent transforms the training approach by changing the data parameters from real photographs requiring manual annotation to synthetic images with algorithmically generated bounding boxes. This parameter change allows for scalable data generation without proportionally increasing annotation complexity, as the synthetic data can be generated automatically in large quantities.
3Ease of manufacture
If manual bounding box generation is used, then training data can be created, but the bounding boxes include portions not containing the object resulting in training inaccuracies
Solution Approach 1:
The patent replaces the mechanical process of manual bounding box drawing with an automated computational system that generates precise bounding boxes through algorithmic methods. This substitution eliminates human error and subjectivity, producing consistently accurate bounding boxes that precisely enclose objects without including extraneous regions.
Data Source
AI summary
Semantic class localization techniques and systems are described. In one or more implementation, a technique is employed to back communicate relevancies of aggregations back through layers of a neural network. Through use of these relevancies, activation relevancy maps are created that describe relevancy of portions of the image to the classification of the image as corresponding to a semantic class. In this way, the semantic class is localized to portions of the image. This may be performed through communication of positive and not negative relevancies, use of contrastive attention maps to different between semantic classes and even within a same semantic class through use of a self-contrastive technique.


