NLP-Based Image Recognition Model for New Objects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image recognition models require extensive training data to accurately recognize objects, leading to poor performance when encountering new objects not covered in the training data, necessitating a method to improve recognition efficiency with minimal training data.
Innovation Solution
An artificial intelligence apparatus and method utilizing a natural language processing (NLP) model-based image recognition model that generates and updates recognition models using label information from user input, allowing for object recognition even with limited training data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional image recognition models are used, then recognition accuracy for trained objects is improved, but the ability to recognize new objects with limited training data deteriorates
Solution Approach 1:
The patent segments the recognition task into two distinct models: a traditional image recognition model for trained objects and a text-based recognition model for new objects. This segmentation allows each model to specialize in its strength, resolving the contradiction between accuracy for trained objects and adaptability for new objects.
Solution Approach 2:
The patent introduces text descriptions as an intermediary between the image input and recognition output. For new objects, the system uses text-based representations to bridge the gap between visual data and object identification, enabling recognition without extensive training data.
2Measurement precision
If extensive training data is collected for new objects, then recognition accuracy for new objects is improved, but the time and resources required deteriorates
Solution Approach 1:
The patent performs preliminary action by pre-processing images into text-based representations and pre-training the text-based recognition model on general object descriptions. This preliminary preparation enables rapid recognition of new objects without requiring extensive new training data when objects are encountered.
Solution Approach 2:
The patent changes the fundamental parameter of recognition from direct image analysis to text-based analysis. By transforming the input modality from visual to linguistic, the system achieves efficient recognition of new objects with minimal training data, dramatically reducing training time and resources.
3Reliability
If multiple image compositions and situations are captured for training, then recognition robustness is improved, but the complexity of data collection and processing deteriorates
Solution Approach 1:
The patent substitutes the mechanical system of capturing and processing multiple image variations with a linguistic system. Instead of mechanically collecting diverse images, the system uses text descriptions that inherently capture semantic information across different compositions and situations, greatly simplifying data collection and processing.
Data Source
AI summary
An artificial intelligence apparatus for recognizing an object included in image data can include a camera, a communication modem, a memory configured to store an image recognition model, a natural language processing (NLP) model, and an NLP model-based image recognition model learned based on the NLP model, and a processor is configured to receive image data from the camera or the communication modem, in response to recognizing an object included in the image data using the image recognition model, generate first recognition information on the object included in the image data, and in response to the recognizing the object included in the image data using the image recognition model being unsuccessful, generate second recognition information on the object included in the image data based on recognizing the object using the NLP model-based image recognition model.


