Semantic Class Localization via Activation Relevancy Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional machine learning techniques for image localization require manual input from users to generate bounding boxes, which is expensive, inefficient, and often inaccurate, limiting the accuracy and availability of localization and subsequent image functionality.

Innovation Solution

The use of machine learning techniques that learn patterns through neural networks to classify and localize semantic classes within images, employing activation relevancy maps and contrastive attention maps to identify relevant image portions without manual input, thereby improving localization accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual bounding box drawing is used for training data generation, then localization can be performed, but the process becomes expensive, inefficient, and inaccurate

Engineering Contradiction:
Improvelocalization accuracyVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system uses automatically generated synthetic training data with precise bounding boxes created through algorithmic processes rather than manual drawing. The synthetic data generation pipeline automatically creates images with objects and their corresponding bounding boxes, eliminating the need for human annotators while maintaining high localization accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates synthetic training data by generating artificial images that copy the essential characteristics of real images but with perfectly accurate bounding boxes. These synthetic copies serve as training examples that teach the localization model without requiring manual annotation of real photographs.

Inventive Principle:
Principle #26Copying

2Reliability

If thousands of training images are used to train a single model, then model accuracy improves, but the complexity and resource requirements increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the training approach by changing the data parameters from real photographs requiring manual annotation to synthetic images with algorithmically generated bounding boxes. This parameter change allows for scalable data generation without proportionally increasing annotation complexity, as the synthetic data can be generated automatically in large quantities.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If manual bounding box generation is used, then training data can be created, but the bounding boxes include portions not containing the object resulting in training inaccuracies

Engineering Contradiction:
Improvedata generation easeVSAvoidbounding box precision
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent replaces the mechanical process of manual bounding box drawing with an automated computational system that generates precise bounding boxes through algorithmic methods. This substitution eliminates human error and subjectivity, producing consistently accurate bounding boxes that precisely enclose objects without including extraneous regions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9846840B1Semantic class localization in images
Publication Date: 2017.12.19 ADOBE INC
  • US9846840B1 patent drawing
  • US9846840B1 patent drawing
  • US9846840B1 patent drawing

AI summary

Semantic class localization techniques and systems are described. In one or more implementation, a technique is employed to back communicate relevancies of aggregations back through layers of a neural network. Through use of these relevancies, activation relevancy maps are created that describe relevancy of portions of the image to the classification of the image as corresponding to a semantic class. In this way, the semantic class is localized to portions of the image. This may be performed through communication of positive and not negative relevancies, use of contrastive attention maps to different between semantic classes and even within a same semantic class through use of a self-contrastive technique.