Visual Attention NER for Social Media Captions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Named Entity Recognition (NER) systems face difficulties in identifying named entities from short social media posts, especially when words are intentionally misspelled or use acronyms and emojis, making it challenging to accurately recognize entities in multimodal captions.

Innovation Solution

A visual named entity system is implemented using a visual attention network that processes images and captions, generating a visual context vector integrated into a recurrent neural network (RNN) with a conditional random field layer to identify named entities, which can then select and incorporate relevant content for social media posts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current NER schemes use traditional text processing methods, then they work well with standard text, but they fail to accurately recognize named entities in social media posts with misspelled words and emojis

Engineering Contradiction:
Improvenamed entity recognition accuracyVSAvoidhandling of misspelled words and emojis
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an image as an intermediary element that bridges the gap between misspelled text and correct entity recognition. The image provides visual context that helps the system interpret ambiguous or misspelled text, acting as a mediator that resolves the contradiction between maintaining text-based NER accuracy and adapting to social media's informal language patterns

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from purely textual analysis to multimodal analysis by incorporating the visual dimension through images. This dimensionality change allows the system to recognize named entities not just through text patterns but also through visual context, thereby improving accuracy for misspelled words and emoji-containing posts

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If NER systems are trained on large well-structured datasets, then they achieve good performance on standard text, but they struggle with the informal and varied language in social media posts

Engineering Contradiction:
ImproveNER model performanceVSAvoidhandling of informal social media language
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a multimodal NER system that performs multiple functions: it processes both standard text and informal social media language, handles misspelled words, interprets emojis, and utilizes image context. This universal system maintains reliability across different text types and language styles, unlike traditional single-function NER systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If visual context is integrated into NER using attention mechanisms, then recognition accuracy improves for misspelled text, but system complexity increases

Engineering Contradiction:
Improveentity identification accuracyVSAvoidvisual attention network architecture
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by using the visual attention mechanism to pre-process and weight relevant regions of the image before the main NER processing occurs. This preliminary visual analysis prepares contextual information that simplifies subsequent text processing, improving accuracy while managing complexity through staged processing

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240354508A1Named entity recognition visual context and caption data
Publication Date: 2024.10.24 SNAP INC
  • US20240354508A1 patent drawing
  • US20240354508A1 patent drawing
  • US20240354508A1 patent drawing

AI summary

A caption of a multimodal message (e.g., social media post) can be identified as a named entity using an entity recognition system. The entity recognition system can use a visual attention based mechanism to generate a visual context representation from an image and caption. The system can use the visual context representation to identify one or more terms of the caption as a named entity.