Cross-category image classification system and generation method based on multi-modal prompt tuning

The cross-category image classification system optimized with multimodal cues solves the problems of cue word tendencies and feature interaction differences in traditional models, and improves the accuracy and generalization ability of image classification, especially on small sample data sets.

CN120707954APending Publication Date: 2025-09-26CHANGCHUN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510822311.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional multimodal visual language models have problems in image classification, such as limited generalization ability due to the tendency of prompt words, time-consuming prompt text design, prompt length restrictions, and differences in the interaction between semantic and image modal features, resulting in insufficient classification accuracy.

Method used

A cross-category image classification system based on multimodal cue tuning is adopted, including a text encoder, an image encoder, a domain-aware feature fusion module and a two-end adaptive module. The output layer is optimized through L2 normalization and L2 regularization to improve the model generalization performance.

Benefits of technology

The classification accuracy of the model in both seen and unseen classes is significantly improved, especially its performance on small sample data sets, which enhances the generalization ability and accuracy of the classification system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005456986400000041
    Figure BDA0005456986400000041
  • Figure BDA0005456986400000042
    Figure BDA0005456986400000042
  • Figure BDA0005456986400000061
    Figure BDA0005456986400000061
Patent Text Reader

Abstract

The embodiment of the invention provides a cross-category image classification system based on multi-mode prompt tuning and a generation method. The classification system is composed of a text encoder, an image encoder, a domain perception feature fusion module, a double-end adaptive module and an output layer. Different from a traditional visual language model, the system adopts a domain perception feature fusion module and a double-end self-adaptive module for fine tuning to form a domain perception sharing prompt rich in general knowledge and category domain knowledge characteristics. The prompt text can effectively improve the generalization performance of the method on visible classes and invisible classes, and the classification system has good performance on other downstream tasks. In order to emphatically reflect the generalization characteristic, the generalization performance is proved through data set experiments in three different scene modes. In addition, the generation method of the system comprises the following steps: (1) building a system development platform; (2) dividing and reading an image data set; (3) constructing an image classification system; (4) training, verifying and testing a classification system; and (5) evaluating the classification system. Through the construction of the classification system and the proposal of the generation method, the purpose of effectively classifying small sample images can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to computer vision, multimodal visual language models, and image classification. Specifically, by fine-tuning visual representations and semantic features through a cue-tuning strategy, it aims to address the key issue of traditional visual language models, where cue-based word biases limit their generalization capabilities, thereby effectively improving classification accuracy. Background Art

[0002] Multimodal vision-language models (VLMs) achieve efficient classification performance by comparatively learning semantic alignment between visual and language modalities, building universal cross-modal representation capabilities. Hint tuning is a strategy for fine-tuning VLMs to make them more adaptable to downstream image classification tasks. Compared to traditional VLM image classification methods, VLM methods utilizing hint tuning exhibit better generalization performance for downstream tasks, providing new insights into the field of VLM.

[0003] Traditional multimodal visual-language image classification methods all use large amounts of sample data for training to improve their classification performance. However, this strategy is very limited. Most researchers find it difficult to pre-train such large models from scratch due to sample size limitations and experimental conditions. With the rapid development of deep learning and cue engineering, researchers have discovered that cue tuning is a simple and effective strategy that leverages the prior knowledge of large visual-language models to fine-tune methods to adapt to downstream tasks, thereby enabling multimodal visual-language image classification methods to exhibit good generalization performance. Existing multimodal visual-language models based on cue tuning strategies can mostly be divided into two categories: visual-language image classification based on parameter tuning and visual-language image classification based on cross-modal feature interaction. Among them, visual-language image classification methods based on parameter tuning (typically cue word optimization) are more popular because they can achieve good generalization performance with fewer parameter fine-tuning strategies.

[0004] However, while cue word optimization strategies have made remarkable progress in the field of multimodal visual language image classification, they still have certain limitations. The main problems are: (1) the cue word bias leads to the problem of generalization imbalance from seen classes to unseen classes; (2) it is very time-consuming to label carefully designed cue texts for different categories of data; (3) it is challenging to generate cue words with both generality and category characteristics; (4) because the cue length is limited, it is difficult to balance the ratio of general cue and category cue; (5) there are differences in feature interactions between semantic and image modalities; Summary of the Invention

[0005] In order to solve the problems existing in the field of multimodal visual language model image classification and further improve the classification accuracy, this embodiment proposes a cross-category image classification system and generation method based on multimodal cue tuning.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] An embodiment of the present invention provides a cross-category image classification system based on multimodal cue tuning, including:

[0008] It consists of a text encoder, an image encoder, a domain-aware feature fusion module, a dual-end adaptation module and an output layer.

[0009] The purpose of a text encoder is to generate a fixed-dimensional text feature vector (such as 512 or 768 dimensions), which is then mapped to a multimodal semantic space after L2 normalization. It is a modification of the Transformer architecture, removing the "decoder" portion of traditional language models used for text generation, retaining only the encoder structure. It consists of a positional encoding embedding module encapsulated in CLIP, a multi-head self-attention mechanism, a fully connected feedforward network, and normalization operations.

[0010] The image encoder generates a fixed-dimensional image feature vector (the same dimension as the text features), also L2-normalized. It is based on a modified Vision Transformer (ViT) architecture, dividing the image into patches and modeling global spatial relationships through a Transformer encoder. The image encoder aims to learn feature information from the input image using various model architectures. It consists of a positional encoding embedding module encapsulated in CLIP, a multi-head self-attention mechanism, a multi-layer perceptron network, and normalization operations.

[0011] The purpose of the domain-aware feature fusion module is to filter, update and fuse the feature information mapped from the text and image encoders to the multimodal semantic space, so as to ultimately retain the domain-aware shared hints with common features and category features. It is composed of the original feature vectors extracted at each level of the text and image encoders as h t-1 The input of the current processing image X is the category information of x t Input, reset gate r t and update gate z t constitute.

[0012] The dual-end adaptation module aims to prevent the generated cue features from retaining too much categorical or general information, thereby preventing this method from improving generalization performance for either the seen or unseen class. Specifically, this module further updates the domain-aware shared cues retained by the domain-aware feature fusion module, optimizing them into general cues applicable to general classes. It consists of two linear layers and two ReLU activation functions.

[0013] The output layer uses a fully connected neural network layer to model the category probability distribution, mapping high-level features to the sample label space through linear transformation. To enhance the model's generalization performance and avoid overfitting, the present invention uses L2 regularization in the output layer to optimize the system.

[0014] The present invention is implemented and comprises the following steps:

[0015] (1) Build a development platform for implementing a small sample image classification system based on multimodal cue optimization features. The hardware platform of the present invention is a server based on an i5-13600k CPU and an NVIDIA RTX4060Ti GAMING SLIM 16G. The server has 16G of video memory and 64G of internal memory. The software platform is an Ubuntu 18.04 operating system with an operating environment of CUDA 11.3.0, Pytorch 1.10.2, and Python 3.8.

[0016] (2) Image dataset division and reading. We selected the small sample image datasets SUN397, EuroSAT, and DTD (SUN397 is a scene recognition dataset, EuroSAT is a satellite image dataset, and DTD is a public texture dataset). We verified the generalization performance of our method on three different types of public datasets.

[0017] (3) Construct a cross-category image classification system based on multimodal cue tuning. The constructed classification system is as described above.

[0018] (4) Training and testing of the classification system. The present invention uses 70% of the samples for training, 10% of the samples for verification, and 20% of the samples for testing. We first divide the data in each data set into visible and invisible classes; then divide them into training sets, verification sets, and test sets according to a certain ratio; finally, train the method on the training set, repeatedly tune and update the verification set, and measure the final classification results on the test set. The size of each input image is set to 224×224, and the batch size is set to 16. During the training process, in order to improve the convergence speed and convergence ability of the model, the present invention uses a gradient descent learning rate to further optimize the system.

[0019] (5) Evaluate the image classification system. To verify the performance of the classification system, the present invention uses the classification evaluation indicators of the visible class (Base), the invisible class (New), and the harmonic mean index (H) on the test set for evaluation. The present invention is compared with some classic methods of multimodal visual language model on the same dataset. The names of various methods are shown in Table 1. The specific comparison results are listed in Table 2.

[0020] Table 1 Comparison method name

[0021]

[0022] Table 2 Evaluation of image classification results on SUN397, EuroSAT and DTD datasets

[0023] BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] Figure 1 It is an overall block diagram of the system of the present invention;

[0026] Figure 2 is a schematic diagram of a text encoder module in the system of the present invention;

[0027] Figure 3 is a schematic diagram of an image encoder module in the system of the present invention;

[0028] Figure 4 is a schematic diagram of a domain-aware feature fusion module in the system of the present invention;

[0029] Figure 5 is a flow chart of the system generation method of the present invention;

[0030] Figure 6 This is a bar chart comparing the classification results of the system of the present invention with those of other methods on the SUN397 dataset;

[0031] Figure 7 This is a bar chart comparing the classification results of the system of the present invention with those of other methods on the EuroSAT dataset;

[0032] Figure 8 It is a histogram comparing the classification results of the system of the present invention with those of other methods on the DTD dataset; DETAILED DESCRIPTION

[0033] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0034] Figure 1The overall block diagram of the system of the present invention is shown in FIG. The system comprises a text and image input layer, a text and image encoder, a feature fusion module, an adaptive module and an output layer.

[0035] Figure 2 It is a schematic diagram of the text encoder module in the system of the present invention. The text encoder module consists of a semantic embedding layer, a position encoding layer, a multi-head attention layer (Multi-HeadAttention), a normalization layer, a fully connected feedforward network layer and a feature output layer. The experiment is a structured mapping from raw text to a high-dimensional semantic tensor. Its main functions include: semantic feature extraction is to capture the global dependency between words in the text through the self-attention mechanism (Self-Attention) to generate a context-aware semantic representation; cross-modal alignment is in the pre-training stage, jointly optimized with the image encoder to ensure that the cosine similarity between text features and corresponding image features is maximized (through the contrast loss function) to achieve alignment of language and visual semantics; zero-sample inference representation can encode any category name or descriptive text into a feature vector, which is directly matched with the image features without additional training.

[0036] Figure 3 It is a schematic diagram of the image encoder module in the system of the present invention. The image encoder module consists of an image embedding layer, a position encoding layer, a multi-head self-attention mechanism layer (Multi-HeadAttention), a multi-layer perceptron (MultilayerPerceptron abbreviated as: MLP) network layer, a normalization layer and a feature output layer. Its main functions include: visual feature extraction uses the self-attention mechanism to model the global relationship between image blocks, without relying on local convolution priors, and is more suitable for capturing long-range dependencies (such as the association between objects and backgrounds); cross-modal alignment means that it shares the same semantic space with the text encoder, and maximizes the similarity of matching pairs of images and texts and minimizes the similarity of unmatched pairs through contrastive learning; generalization means that the pre-training stage is exposed to massive and diverse image and text data (such as 400 million pairs), so that the model learns robust visual representation capabilities and can be directly migrated to downstream tasks (such as classification, retrieval) without fine-tuning.

[0037] Figure 4 This is a schematic diagram of the domain-aware feature fusion module in the system of the present invention. It is composed of the original feature vectors extracted at each level of the text and image encoder as h t-1 The input of the current processing image X is the category information of x t Input, reset gate r t and update gate z t The main purpose of the domain-aware feature fusion module is to continuously integrate general knowledge with domain category prompts to generate general-aware text prompts with basic general knowledge characteristics and domain category knowledge characteristics. The function of is to generate the temporary state at the current moment, combining the historical information adjusted by the reset gate and the current input; update gate z t The function of is to control how much information in the current hidden state comes from the hidden state of the previous moment, and how much comes from the newly generated candidate hidden state; reset gate r t The function is to control the influence of the previous hidden state on the candidate hidden state, which is used to capture short-term dependencies; the final hidden state h t The function of is to fuse the hidden state of the previous moment and the candidate hidden state through the update gate and output the final state of the current moment. Their formulas are as follows:

[0038]

[0039] z t =σ(U z [h t-1 ]×[x t ])

[0040] r t =σ(U r [h t-1 ]×[x t ])

[0041]

[0042] The gate signal generates a probability vector through the activation function σ, which ranges from [0,1] and controls the degree of retention of historical information and the injection ratio of new information. r and U z are learnable parameters, x t ∈R d is the semantic embedding of the category at the current moment, h t-1 is the output at the previous moment.

[0043] Figure 5 The flowchart of the system generation method of the present invention is divided into five steps: (1) building a system development platform; (2) dividing and reading image data sets; (3) building an image classification system; (4) training and testing the classification system; and (5) evaluating the classification system.

[0044] Figure 6This is a bar chart comparing the classification results of the proposed system with other methods on the SUN397 dataset. To verify the performance of the proposed system, this example evaluated the accuracy metrics Base (visible classes), New (invisible classes), and H (harmonic mean index) on the SUN397 dataset. This example was compared with current classic multimodal visual language methods, achieving high classification performance of 82.40%, 77.80%, and 80.03% with a Shot = 16 setting. Compared to other methods, the proposed system demonstrates excellent classification performance in the field of scene recognition image classification.

[0045] Figure 7 This is a bar chart comparing the classification results of the system of the present invention with other methods on the EuroSAT dataset. In order to verify the performance of the system of the present invention, this example evaluates the accuracy indicators Base (visible class), New (invisible class) and H (harmonic mean index) on the EuroSAT dataset. This example is compared with the current classic multimodal visual language method, and achieves efficient classification performance of 82.40%, 77.80%, and 80.03% on the Shot=16 setting. The system of the present invention not only surpasses the baseline method on the visible class (Base), but also significantly increases the classification performance on the invisible class (New) and the harmonic mean index (H), reaching 76.07% (an increase of 3.3 percentage points compared to the best baseline method TCP) and 84.86% (an increase of 4.94 percentage points compared to the best baseline method TCP). This significant improvement further illustrates the performance of this classification system.

[0046] Figure 8 This is a bar chart comparing the classification results of the system of the present invention with other methods on the DTD dataset. In order to verify the performance of the system of the present invention, this example evaluates the accuracy indicators Base (visible class), New (invisible class) and H (harmonic mean index) on the DTD dataset. This example is compared with the current classic multimodal visual language method, and achieved efficient classification performance of 82.50%, 58.41%, and 68.40% in the setting of Shot=16. Compared with other methods, the system of the present invention shows good classification performance in the field of texture image classification.

Claims

1. A cross-category image classification system based on multimodal cue tuning, characterized by include: Text encoder, used to extract basic text feature information; Image encoder, used to accurately extract feature information from images; The domain-aware feature fusion module is used to filter, update, and fuse the feature information mapped from the text and image encoders to the multimodal semantic space, thereby ultimately retaining domain-aware shared cues with both common and category features. The dual-end adaptation module is used to prevent the generated cue features from retaining too much category or common information, thereby avoiding the phenomenon of single-aspect generalization performance improvement for either seen or unseen classes.

2. The system according to claim 1, wherein: The text encoder generates a fixed-dimensional text feature vector (such as 512 dimensions or 768 dimensions) through the CLIP text encoder, and maps it into the multimodal semantic space after L2 normalization.

3. The system according to claim 1, wherein: The image encoder models global spatial relationships through the Vision Transformer encoder and maps the image feature vector encoding into a multimodal semantic space.

4. The system according to claim 1, wherein: The domain-aware feature fusion module accurately filters, updates, and fuses prompts through a gating mechanism, thereby ultimately retaining domain-aware shared prompts with common features and category features, thereby achieving accurate extraction of prompt information.

5. The system according to claim 1, wherein: The dual-end adaptation module further updates the domain-aware shared hints through an adaptive mechanism, optimizing them into general hints applicable to general categories, thereby improving the accuracy of the classification results.