Image classification method and system based on visual language large model

Through the image classification method based on visual language big model, visual features and text prompt words are used to fusion, the problem of dependence on large amounts of data and poor classification effect in traditional methods is solved, and image classification with high accuracy is achieved under a small amount of data.

CN120164034APending Publication Date: 2025-06-17HUNAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510313238.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Traditional image classification methods require a large amount of data for training, and when there are fewer image categories or not in the training data, the classification effect is not obvious and the accuracy is low.

Method used

Image classification method based on visual language big model is adopted, and image feature extraction, fusion and classification are performed through an image classification network composed of VIT image feature extraction module, visual feature adapter, visual feature projection module, text prompt word encoding module, language big model and classification head.

Benefits of technology

It significantly improves the ability of the image classification network to understand complex classification tasks, reduces the number of original images required, and ensures good classification results and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164034A_ABST
    Figure CN120164034A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image classification, in particular to an image classification method and system based on a visual language large model, and the method comprises the following steps: 1, obtaining a plurality of original images, and constructing an image classification network; 2, selecting one original image from a plurality of original images, inputting the selected original image into an image classification network, and finally obtaining a category prediction result; 3, constructing a loss function by utilizing a category prediction result and a real category; 4, circulating the steps 2 and 3, minimizing the loss function until the loss function converges or the number of iterations reaches a set number of times, and updating the weight of the image classification network to obtain a trained image classification network; and 5, deploying the trained image classification network to a device end, and classifying the images by using the device end to obtain a classification result. According to the method, the problems of insufficient global information capture and low visual and language information fusion efficiency in a traditional single-mode classification method are solved, and higher classification precision and task generalization ability are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification, and particularly relates to an image classification method and system based on a vision-language large model. Background Art

[0002] Traditional image classification methods generally use methods such as support vector machines and decision trees for classification. Image classification methods based on deep learning use convolutional neural networks to classify images. These classification methods usually require collecting a large amount of data to train the model. And when there are few images in a certain category or the image category is not in the training data, it will cause problems such as unclear classification effect and low accuracy. Therefore, there is an urgent need for an image classification method that does not require a large amount of data. Summary of the Invention

[0003] The present invention provides an image classification method and system based on a vision-language large model to solve the technical problems mentioned in the background art.

[0004] To achieve the above object, the technical solution of the present invention is realized as follows:

[0005] The present invention provides an image classification method based on a vision-language large model, including the following steps:

[0006] S1. Obtain multiple original images and construct an image classification network;

[0007] The image classification network includes a VIT image feature extraction module, a vision feature adapter, a vision feature projection module, a text prompt encoding module, a language large model I, a language feature adapter, and a classification head; the VIT image feature extraction module, the vision feature adapter, and the vision feature projection module are connected in sequence, the vision feature projection module and the text prompt encoding module are respectively connected to the language large model I, and the language large model I, the language feature adapter, and the classification head are connected in sequence;

[0008] S2. Select one original image from the multiple original images and input it into the image classification network to finally obtain a class prediction result;

[0009] S3. Use the class prediction result and the true class to construct a loss function;

[0010] S4. Loop S2 and S3, minimize the loss function until the loss function converges or the number of iterations reaches the set number, and update the weights of the image classification network to obtain the trained image classification network;

[0011] S5. Deploy the trained image classification network to the device side and use the device side to classify the image to obtain a classification result.

[0012] Further, the visual feature adapter includes two sequentially connected fully connected layers and one activation layer. The two fully connected layers include a series of linear transformations, and the activation layer is a non-linear activation layer. The visual feature adapter is used to reduce the dimension, denoise, and optimize the visual features output by the VIT image feature extraction module to retain key visual features.

[0013] Further, the visual feature projection module includes a linear projection layer or a multi-layer perceptron MLP.

[0014] Further, the language large model one selects the language large model Llama2.

[0015] Further, the visual feature adapter selects the I-Adapter model.

[0016] Further, the language feature adapter selects the L-Adapter model.

[0017] Further, S1 specifically includes the following steps:

[0018] S11. Use a collection device to collect multiple original images;

[0019] S12. Classify the multiple original images to obtain the classification labels of the multiple original images;

[0020] S13. Extract the key feature information that can reflect the classification object from the multiple original images to obtain the visual feature descriptions of the multiple original images.

[0021] Further, S2 specifically includes the following steps:

[0022] S21. Select one original image from the multiple original images and input it into the VIT image feature extraction module in the image classification network to obtain visual features;

[0023] S22. Input the obtained visual features into the visual feature adapter, and the visual feature adapter adapts and optimizes the visual features to obtain enhanced features;

[0024] S23. Input the enhanced features into the visual feature projection module, and the visual feature projection module maps the enhanced features to the shared semantic space through a linear projection layer or a multi-layer perceptron MLP for alignment with the text features to obtain mapped features;

[0025] S24. Concatenate the visual feature description of the original image and the classification label of the original image through the cat operation to obtain the text prompt word of the original image;

[0026] S25. Input the text prompt of the original image into the text prompt encoding module. The text prompt encoding module extracts features from the text prompt of the original image through the text encoder in the second language large model to obtain a semantic vector;

[0027] S26. Use the cat operation to concatenate the mapping features obtained in S23 and the semantic vector obtained in S25 to obtain a visual - language fusion feature;

[0028] S27. Input the visual - language fusion feature into the first language large model. The first language large model mines the correlation between visual and language information in the visual - language fusion feature to generate a first fusion feature;

[0029] S28. Input the first fusion feature into the language feature adapter. The language feature adapter optimizes the expression ability of the first fusion feature through dimensionality reduction, feature screening, and enhancement operations to obtain an optimized feature;

[0030] S29. Input the optimized feature into the classification head for prediction of specific categories to obtain a category prediction result.

[0031] On the other hand, the present invention also provides an image classification system, including a computer device, which is programmed or configured to execute the above - mentioned image classification method.

[0032] Advantages of the present invention:

[0033] 1. The present invention discloses an image classification method based on a visual - language large model. By concatenating visual feature descriptions and classification labels, and extracting features through the text prompt encoding module, the key features of the image are transformed into semantic descriptions to obtain a semantic vector. Then, the mapping feature and the semantic vector are concatenated to form a visual - language fusion feature, which contains local, global, and semantic information, providing a more semantically deep multi - modal input for downstream tasks and significantly improving the understanding ability of the image classification network for complex classification tasks.

[0034] In addition, the number of original images required by the present invention is relatively small compared to traditional image classification methods. That is, under the premise of a small number of original images, the present invention can also ensure good classification effects and high accuracy.

[0035] 2. The present invention also designs an image classification network, which includes a VIT image feature extraction module, a visual feature adapter, a visual feature projection module, a text prompt encoding module, a language model 1, a language feature adapter, and a classification head. Through the efficient cooperation among the modules, the problems of insufficient global information capture and low efficiency of visual and language information fusion in traditional single-modal classification methods are solved, and higher classification accuracy and task generalization ability are achieved. The design of the image classification network has high task pertinence and scalability, providing a new technical route for the development and application of multi-modal models. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flowchart of the image classification method in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the understanding of the disclosure of the present invention more thorough and comprehensive.

[0038] Refer to Figure 1 , the embodiments of the present application provide an image classification method based on a vision-language large model, including the following steps:

[0039] S1. Obtain multiple original images (or called defective pictures), and construct an image classification network;

[0040] The image classification network includes a VIT image feature extraction module, a visual feature adapter, a visual feature projection module, a text prompt encoding module, a language model 1, a language feature adapter, and a classification head; the VIT image feature extraction module, the visual feature adapter, and the visual feature projection module are connected in sequence, the visual feature projection module and the text prompt encoding module are respectively connected to the language model 1, and the language model 1, the language feature adapter, and the classification head are connected in sequence;

[0041] S2. Select one original image from the multiple original images and input it into the image classification network to finally obtain a class prediction result;

[0042] S3. Use the class prediction result and the true class to construct a loss function;

[0043] S4. Loop S2 and S3, minimize the loss function until the loss function converges or the number of iterations reaches the set number, and update the weights of the image classification network to obtain a trained image classification network;

[0044] S5. Deploy the trained image classification network to the device side, and use the device side to classify the image to obtain the classification result.

[0045] The present invention discloses an image classification method based on a vision-language large model. By splicing visual features and classification labels, and extracting features through a text prompt encoding module, the key features of the image are transformed into semantic descriptions to obtain semantic vectors. Then, the mapped features are spliced with the semantic vectors to form vision-language fusion features, which contain local, global, and semantic information, providing a more semantically deep multi-modal input for downstream tasks and significantly improving the understanding ability of the image classification network for complex classification tasks.

[0046] In addition, the number of original images required by the present invention is relatively small compared to traditional image classification methods. That is, the present invention can ensure good classification effects and high accuracy even under the premise of a small number of original images.

[0047] The present invention designs an image classification network. Through the efficient cooperation between modules, it solves the problems of insufficient global information capture and low visual and language information fusion efficiency in traditional single-modal classification methods, and achieves higher classification accuracy and task generalization ability. The design of the image classification network has high task pertinence and scalability, providing a new technical route for the development and application of multi-modal models.

[0048] In some embodiments, the VIT image feature extraction module is designed based on the Transformer model architecture. It can directly process the entire original image through the self-attention mechanism Self-Attention, divide the original image into visual features of a fixed size, and establish relationships between visual features globally. This characteristic makes the VIT image feature extraction module particularly suitable for processing tasks that require global context information, such as image classification of complex patterns, feature extraction, and visual understanding in multi-modal tasks. When extracting visual features, the VIT image feature extraction module can not only retain fine-grained local information but also capture the global associations between different regions in the image, thus showing superior performance in many computer vision tasks. Therefore, the VIT image feature extraction module is used in the present invention to extract visual features. Compared with the convolutional layer CNN, the convolutional layer CNN extracts local features of the image through convolutional operations and has strong spatial perception ability. However, due to the limitation of the receptive field range, the convolutional layer CNN cannot effectively capture the global relationships of image features.

[0049] Since the features output by the VIT image feature extraction module are often high-dimensional and contain complex global information, but may not directly match the downstream tasks. Therefore, a visual feature adapter is needed to adapt and optimize the extracted visual features.

[0050] In some embodiments, the visual feature adapter includes two sequentially connected fully connected layers and one activation layer. The two fully connected layers include a series of linear transformations, and the activation layer is a non-linear activation layer. The visual feature adapter is used to reduce the dimension, denoise, and optimize the visual features output by the VIT image feature extraction module to retain key visual features. At the same time, enhance its expressive ability to make it more suitable for fusion processing with subsequent modules.

[0051] In some embodiments, the visual feature projection module includes a linear projection layer or a multi-layer perceptron MLP.

[0052] Specifically, the visual feature projection module is used to uniformly map the output of the visual feature adapter (i.e., enhanced features) into a feature representation form that can be understood by the language model. This module maps the enhanced features to a shared semantic space through a linear projection layer or a multi-layer perceptron MLP for alignment with text features. Thus, ensuring that the enhanced features can be fused with the text prompt in the same feature space.

[0053] In some embodiments, the language large model I selects the language large model Llama2.

[0054] In some embodiments, the visual feature adapter selects the I-Adapter model.

[0055] In some embodiments, the language feature adapter selects the L-Adapter model.

[0056] In some embodiments, S1 specifically includes the following steps:

[0057] S11. Use a collection device to collect multiple original images;

[0058] S12. Classify the multiple original images to obtain classification labels for the multiple original images;

[0059] S13. Extract key feature information that can reflect the classification object from the multiple original images to obtain visual feature descriptions of the multiple original images.

[0060] In some embodiments, S2 specifically includes the following steps:

[0061] S21. Select one original image from the multiple original images and input it into the VIT image feature extraction module (i.e., image encoder) in the image classification network to obtain visual features;

[0062] S22. Input the obtained visual features into the visual feature adapter, and the visual feature adapter adapts and optimizes the visual features to obtain enhanced features;

[0063] S23. Input the enhanced features into the visual feature projection module. The visual feature projection module maps the enhanced features to a shared semantic space through a linear projection layer or a multi-layer perceptron (MLP) for alignment with the text features, thereby obtaining the mapped features;

[0064] S24. Concatenate the visual feature description of the original image with the classification label (i.e., label text) of the original image through the cat operation to obtain the text prompt for the original image; the text prompt is a high-level description of the target task, such as the name of the category or other relevant semantic information;

[0065] By concatenating the visual feature description of the original image with the classification label of the original image, the prompt can be made more rich, and it can enable the large language model two to deeply understand the mapped features. This step transforms the mapped features into a more semantically meaningful description, forms a combination with the label information of specific categories, and serves as the textual expression of the visual prompt, providing a basis for subsequent vision-language fusion; preferably, the large language model two selects LLaMA;

[0066] S25. Input the text prompt of the original image into the text prompt encoding module (i.e., text encoder). The text prompt encoding module extracts features from the text prompt of the original image through the text encoder in the large language model two to generate high-dimensional semantic vectors; these semantic vectors provide context information, enabling more efficient docking of the mapped features with the semantic vectors;

[0067] S26. Use the cat operation to concatenate and align the mapped features obtained in S23 and the semantic vectors obtained in S25 to obtain the vision-language fusion features; the vision-language fusion features contain both local and global features of visual information and the semantic context of the text prompt, providing a more comprehensive information input for subsequent processing;

[0068] S27. Input the vision-language fusion features into the large language model one. The large language model one utilizes its powerful multi-layer attention mechanism and context modeling ability to further explore the correlation between visual and language information, generating semantically rich fusion features, i.e., fusion feature one; due to the large language model one's ability for large-scale pre-training, its output performs particularly powerfully in multi-modal understanding tasks;

[0069] S28. Input fusion feature one into the language feature adapter. The language feature adapter optimizes the expressive ability of fusion feature one through dimensionality reduction, feature screening, and enhancement operations to obtain the optimized features;

[0070] Specifically, the language feature adapter adapts and optimizes the output features of the first language large model (i.e., the first fused feature). Through dimensionality reduction, feature screening, and enhancement operations, the language feature adapter optimizes the expressive ability of the fused features to make them more suitable for the requirements of downstream classification tasks. The introduction of the language feature adapter can effectively improve the convergence speed and classification performance of the image classification network;

[0071] S29. Input the optimized features into the classification head to perform predictions for specific categories, and obtain the category prediction results (i.e., defect classification). The classification head is a simple fully connected layer. The output category prediction results are the final image classification results, which combine visual and language information to improve the accuracy and robustness of classification.

[0072] On the other hand, the present invention also provides an image classification system, including a computer device, which is programmed or configured to execute the above image classification method.

[0073] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Moreover, the technical solutions between various embodiments of the present invention can be combined with each other, but it must be based on the fact that those skilled in the art can implement it. When the combination of technical solutions conflicts with each other or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An image classification method based on a large visual language model, characterized in that: The steps include: S1. Obtain multiple original images and build an image classification network; The image classification network includes a VIT image feature extraction module, a visual feature adapter, a visual feature projection module, a text prompt word encoding module, a language large model 1, a language feature adapter and a classification head; the VIT image feature extraction module, the visual feature adapter and the visual feature projection module are connected in sequence, the visual feature projection module and the text prompt word encoding module are respectively connected to the language large model 1, and the language large model 1, the language feature adapter and the classification head are connected in sequence; S2, select one original image from multiple original images and input it into the image classification network to finally obtain the category prediction result; S3, construct a loss function using the category prediction results and the true category; S4, loop S2 and S3, minimize the loss function, until the loss function converges or the number of iterations reaches the set number, and update the weights of the image classification network to obtain the trained image classification network; S5. Deploy the trained image classification network to the device, use the device to classify the image, and obtain the classification result.

2. The image classification method according to claim 1, characterized in that: The visual feature adapter includes two fully connected layers and one activation layer connected in sequence, the two fully connected layers include a series of linear transformations, and the activation layer is a nonlinear activation layer. The visual feature adapter is used to reduce the dimension, denoise and optimize the visual features output by the VIT image feature extraction module to retain key visual features.

3. The image classification method according to claim 1, characterized in that: The visual feature projection module includes a linear projection layer or a multi-layer perceptron MLP.

4. The image classification method according to claim 1, characterized in that: The language macro model 1 selects the language macro model LlaMA2.

5. The image classification method according to claim 1, characterized in that: The visual feature adapter uses the I-Adapter model.

6. The image classification method according to claim 1, characterized in that: The language feature adapter uses the L-Adapter model.

7. The image classification method according to claim 1, characterized in that: The S1 specifically includes the following steps: S11, using an acquisition device to acquire multiple original images; S12, classifying the multiple original images to obtain classification labels of the multiple original images; S13. Extract key feature information that can reflect the classification object from the multiple original images to obtain visual feature descriptions of the multiple original images.

8. The image classification method according to claim 7, characterized in that: The S2 specifically includes the following steps: S21, selecting an original image from multiple original images and inputting it into a VIT image feature extraction module in the image classification network to obtain visual features; S22, inputting the obtained visual features into a visual feature adapter, and the visual feature adapter adapts and optimizes the visual features to obtain enhanced features; S23, inputting the enhanced features into a visual feature projection module, which maps the enhanced features to a shared semantic space through a linear projection layer or a multi-layer perceptron MLP so as to align them with the text features, thereby obtaining mapping features; S24, concatenating the visual feature description of the original image and the classification label of the original image through a cat operation to obtain a text prompt word of the original image; S25, inputting the text prompt words of the original image into the text prompt word encoding module, and the text prompt word encoding module extracts features of the text prompt words of the original image through the text encoder in the language large model 2 to obtain a semantic vector; S26, using the cat operation to concatenate the mapping feature obtained in S23 and the semantic vector obtained in S25 to obtain the visual-language fusion feature; S27, inputting the visual-language fusion feature into the language big model 1, the language big model 1 mines the correlation between the visual and language information in the visual-language fusion feature to generate the fusion feature 1; S28, inputting the fused feature 1 into the language feature adapter, and the language feature adapter optimizes the expression ability of the fused feature 1 through dimension reduction, feature screening and enhancement operations to obtain optimized features; S29: Input the optimized features into the classification head to make predictions for specific categories and obtain category prediction results.

9. An image classification system, comprising a computer device, characterized in that: The computer device is programmed or configured to execute the image classification method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Unmanned vehicle surrounding target identification method and system based on open target detection

    CN120673380A