Co-Attention Visual Content Understanding for Symbol Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current vision technologies are inadequate in understanding deeper subjective interpretations of visual content, such as sentiment and symbolism, which are crucial for analyzing advertisements that convey messages through carefully organized objects and ad messages.

Innovation Solution

The method involves determining region proposals in images, attending to corresponding symbols, extracting and fusing appearance features using a neural network, and projecting these features into a semantic embedding space to predict descriptive messages by computing similarity measures between fused features and known descriptive messages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If current vision approaches are used for object and scene-centric interpretation, then basic visual content analysis is achieved, but deeper subjective interpretation such as sentiment and symbolism cannot be understood

Engineering Contradiction:
Improvedeeper subjective interpretation informationVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the visual content analysis into multiple processing stages: region proposal generation, symbol detection, feature extraction, and message prediction. Each stage handles specific aspects of the analysis, allowing the system to progressively build from basic object detection to deeper symbolic interpretation without overwhelming complexity in a single module.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate representations including region proposals, symbol features, and fused appearance-symbol features that act as mediators between raw image data and final symbolic interpretation. These intermediate structures enable gradual transformation from visual pixels to semantic meaning, bridging the gap between basic vision and subjective interpretation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If advertisements are analyzed using traditional object detection, then objects can be identified, but the symbolic and sentimental attributes required for decoding ad messages cannot be extracted

Engineering Contradiction:
Improvesymbolic and sentimental attributesVSAvoiddifficulty of detecting symbolic attributes
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The system merges appearance features extracted from image regions with symbol features detected in those regions to create fused representation. This combination allows the system to capture both the visual characteristics of objects and their symbolic meanings, enabling extraction of sentimental attributes that neither approach could achieve alone.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent extends the analysis from the visual dimension (appearance features) to the semantic dimension (symbol features and descriptive messages). By projecting features into a semantic embedding space and analyzing them in this additional dimensional domain, the system can detect symbolic attributes that are not apparent in the original visual space.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If region proposals are generated and symbols are attended without feature fusion, then basic visual elements can be identified, but accurate understanding of visual content requiring alignment between symbols and objects cannot be achieved

Engineering Contradiction:
Improvealignment precision between symbols and objectsVSAvoidfeature processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual or rule-based alignment mechanisms with a learned feature fusion process using neural networks. The system automatically learns how to align and combine appearance features with symbol features through training on annotated data, eliminating the need for complex hand-crafted alignment rules while achieving precise symbol-object correspondence.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If semantic embedding space is created without attention mechanisms, then basic feature representation can be achieved, but precise alignment and understanding of symbol-image correspondence cannot be obtained

Engineering Contradiction:
Improvesymbol-image correspondence precisionVSAvoidattention mechanism complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary attention operations to identify and weight important symbols and image regions before creating the final semantic embedding. By pre-attending to relevant features and computing attention weights in advance, the system prepares refined input representations that improve the precision of symbol-image correspondence in the embedding space without adding complexity to the core embedding process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11210572B2Aligning symbols and objects using co-attention for understanding visual content
Publication Date: 2021.12.28 SRI INTERNATIONAL
  • US11210572B2 patent drawing
  • US11210572B2 patent drawing
  • US11210572B2 patent drawing

AI summary

A method, apparatus and system for understanding visual content includes determining at least one region proposal for an image, attending at least one symbol of the proposed image region, attending a portion of the proposed image region using information regarding the attended symbol, extracting appearance features of the attended portion of the proposed image region, fusing the appearance features of the attended image region and features of the attended symbol, projecting the fused features into a semantic embedding space having been trained using fused attended appearance features and attended symbol features of images having known descriptive messages, computing a similarity measure between the projected, fused features and fused attended appearance features and attended symbol features embedded in the semantic embedding space having at least one associated descriptive message and predicting a descriptive message for an image associated with the projected, fused features.