Image Classification Using Dynamic Text Category Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image classification methods face inaccuracies due to mismatches between extracted text information and classification tasks, lacking universality and accuracy, as they fail to consider image differences and classification label quantities.
Innovation Solution
Dynamically select a text information category based on image differences and classification label quantities, extracting text information of varying dimensions to enhance compatibility with classification tasks, using a comprehensive feature extraction model and text encoding features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If uniformly extracted text information is used for image classification, then the processing method is simple, but the classification accuracy deteriorates due to mismatch between text information and classification task
Solution Approach 1:
The patent applies dynamics by dynamically selecting different text information categories (first category, second category, or third category) based on image differences and classification label quantities. This dynamic adaptation allows the system to adjust the text extraction strategy to match the specific classification task requirements, thereby improving classification accuracy while maintaining reasonable processing complexity.
Solution Approach 2:
The patent changes parameters by varying the text information category selected based on two key parameters: image differences (similarity/dissimilarity among images) and classification label quantities. By changing the text extraction parameters according to these inputs, the system optimizes the matching between text information and classification tasks, resolving the accuracy-simplicity contradiction.
2Measurement precision
If different text information categories are dynamically selected based on image differences and classification label quantities, then the classification accuracy improves, but the device complexity increases
Solution Approach 1:
The patent segments the text information extraction process into three distinct categories (first category for basic text, second category for enhanced text, third category for task-specific text). This segmentation allows the system to select only the necessary category for each classification task, improving accuracy while managing complexity through modular, conditional execution rather than always using the most complex extraction method.
Solution Approach 2:
The patent introduces an intermediary mechanism (the text information category selection module) that mediates between the image classification task requirements and the text extraction process. This intermediary dynamically determines which text category to use based on image differences and label quantities, thereby improving classification accuracy while abstracting away the complexity from the overall system architecture.
3Adaptability or versatility
If text information extraction is adapted to different classification tasks, then the universality of the method improves, but the processing time increases
Solution Approach 1:
The patent applies preliminary action by pre-defining three text information categories with different levels of detail and processing requirements. The system performs a quick assessment of image differences and classification label quantities to determine which pre-defined category to use, avoiding the need to create custom extraction methods for each task. This improves universality while controlling processing time through预先 prepared extraction strategies.
Data Source
AI summary
Some aspects of the disclosure provide a method of image processing. For example, one or more text information categories for an image classification of a plurality of images are determined based on one or more parameters that indicate a classification task difficulty level of the image classification of the plurality of images. Respective text information of the plurality of images is extracted, text information of an image in the plurality of images is extracted according to the one or more text information categories. Respective text encoding features of the plurality of images are determined based on the respective text information of the plurality of images. Respective classification labels of the plurality of images are determined based on the respective text encoding features of the plurality of images and respective image encoding features of the plurality of images. Apparatus and non-transitory computer-readable storage medium counterpart embodiments are also contemplated.


