NLP-Based Image Recognition Model for New Objects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image recognition models require extensive training data to accurately recognize objects, leading to poor performance when encountering new objects not covered in the training data, necessitating a method to improve recognition efficiency with minimal training data.

Innovation Solution

An artificial intelligence apparatus and method utilizing a natural language processing (NLP) model-based image recognition model that generates and updates recognition models using label information from user input, allowing for object recognition even with limited training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional image recognition models are used, then recognition accuracy for trained objects is improved, but the ability to recognize new objects with limited training data deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoidability to recognize new objects
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the recognition task into two distinct models: a traditional image recognition model for trained objects and a text-based recognition model for new objects. This segmentation allows each model to specialize in its strength, resolving the contradiction between accuracy for trained objects and adaptability for new objects.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces text descriptions as an intermediary between the image input and recognition output. For new objects, the system uses text-based representations to bridge the gap between visual data and object identification, enabling recognition without extensive training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If extensive training data is collected for new objects, then recognition accuracy for new objects is improved, but the time and resources required deteriorates

Engineering Contradiction:
Improverecognition accuracy for new objectsVSAvoidtraining time and resources
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-processing images into text-based representations and pre-training the text-based recognition model on general object descriptions. This preliminary preparation enables rapid recognition of new objects without requiring extensive new training data when objects are encountered.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the fundamental parameter of recognition from direct image analysis to text-based analysis. By transforming the input modality from visual to linguistic, the system achieves efficient recognition of new objects with minimal training data, dramatically reducing training time and resources.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If multiple image compositions and situations are captured for training, then recognition robustness is improved, but the complexity of data collection and processing deteriorates

Engineering Contradiction:
Improverecognition robustnessVSAvoiddata collection and processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent substitutes the mechanical system of capturing and processing multiple image variations with a linguistic system. Instead of mechanically collecting diverse images, the system uses text descriptions that inherently capture semantic information across different compositions and situations, greatly simplifying data collection and processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11200467B2Artificial intelligence apparatus and method for recognizing object included in image data
Publication Date: 2021.12.14 LG ELECTRONICS INC
  • US11200467B2 patent drawing
  • US11200467B2 patent drawing
  • US11200467B2 patent drawing

AI summary

An artificial intelligence apparatus for recognizing an object included in image data can include a camera, a communication modem, a memory configured to store an image recognition model, a natural language processing (NLP) model, and an NLP model-based image recognition model learned based on the NLP model, and a processor is configured to receive image data from the camera or the communication modem, in response to recognizing an object included in the image data using the image recognition model, generate first recognition information on the object included in the image data, and in response to the recognizing the object included in the image data using the image recognition model being unsuccessful, generate second recognition information on the object included in the image data based on recognizing the object using the NLP model-based image recognition model.