Image Classification Using Text-Visual Feature Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Artificial intelligence models for image classification face significant performance degradation when classifying images from domains different from those they were trained on, making it challenging to achieve high accuracy across various domains, especially when training data is limited.
Innovation Solution
An apparatus and method that combines image classification, image-text fusion, and text description generation units using specific loss functions (Ltask, Lalign, and Lexpl) to align visual and textual features, allowing for effective classification of domain non-specific images by integrating text information into the learning process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If an artificial intelligence model is trained on images from a specific domain, then classification performance on that domain is improved, but classification performance on images from different domains deteriorates
Solution Approach 1:
The patent introduces text information as an intermediary element that bridges the gap between different image domains. By aligning visual features with textual features through contrastive learning, the model learns domain-invariant representations that can generalize across different domains while maintaining high classification accuracy on specific domains
Solution Approach 2:
The patent creates a multi-functional learning framework that simultaneously performs classification (Ltask), feature alignment (Lalign), and exploration (Lexpl) tasks. This universal approach allows the model to handle multiple objectives - maintaining domain-specific accuracy while gaining cross-domain adaptability through a single unified system
2Adaptability or versatility
If images from a variety of domains are collected to train the model, then domain adaptability is improved, but data collection complexity and cost increase
Solution Approach 1:
Instead of collecting diverse images from multiple domains, the patent uses text as an intermediary that can be generated or obtained more easily. Text descriptions serve as a bridge that connects different visual domains without requiring actual diverse training images, thereby reducing data collection complexity while maintaining adaptability
Solution Approach 2:
The patent creates textual representations (copies) of image content that can be used as proxies for actual diverse images. These text descriptions capture essential semantic information without requiring the model to see actual images from all possible domains, reducing the need for extensive data collection
3Measurement precision
If text information is integrated into the learning process, then cross-domain classification accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent segments the complex learning process into three distinct loss function components: Ltask for classification, Lalign for feature alignment, and Lexpl for exploration. This segmentation allows each component to be optimized independently while working together to achieve cross-domain accuracy, making the overall computational process more manageable and efficient
Data Source
AI summary
Disclosed is an apparatus for classifying domain non-specific images using text according to one embodiment of the present invention. According to the present invention, the learning process is performed not only using the images but also text information together, and thus even during training on images from only a few specific domains, it is possible to effectively classify the images from different domains with high accuracy by applying the human inference process.


