Lightweight zero-sample image classification method and system based on state space model
Through the lightweight zero-sample image classification method of the state space model, through the interactive learning of visual and semantic features, the problems of poor alignment effect and high computational complexity in the existing methods are solved, and efficient classification on resource-limited devices are achieved.
Patent Information
- Application Number
- CN202510410078.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing zero-sample image classification method is independently performed in the process of feature extraction and feature alignment, making it difficult to achieve effective visual and semantic feature alignment, and the computational complexity of the Transformer-based method is limited, which limits its application on resource-limited devices.
The lightweight zero-sample image classification method based on state space model is adopted. Through the visual encoder, semantic encoder and multi-mode fusion module, Dual-Mamba2 block and Mamba2 block cascade are used to realize interactive learning of visual and semantic features, reducing computational complexity and improving classification accuracy.
While maintaining the accuracy of category classification without seeing, it effectively reduces the complexity of model calculations, making it more suitable for deployment on hardware devices, and is suitable for applications such as planetary image perception, environmental monitoring and natural disaster management.
Smart Images

Figure CN120339693A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and specifically relates to a lightweight zero-shot image classification method and system based on a state space model, which can be used for planetary image perception, environmental monitoring, urban planning, and natural disaster management. Background Art
[0002] Zero-shot learning identifies categories never seen during the training phase by leveraging existing knowledge such as attribute information, text description information, etc. Currently, zero-shot learning has been widely applied in terrain and landform scene classification, animal and plant category classification, and can also be used for planetary image perception, environmental monitoring, urban planning, and natural disaster management. Zero-shot image classification is one of the basic tasks of zero-shot learning in the field of computer vision. It refers to the identification of categories never seen during the training phase when the training set and the test set do not intersect, and can provide strong support for higher-level applications such as planetary image perception, environmental feature monitoring, urban planning, and natural disaster management.
[0003] Traditional image classification methods rely on the training set to obtain a classification model, which limits the image categories that can be recognized to those included in the training set. In practical applications, the training set usually only contains a small part of the world's open set and is difficult to cover all image categories, making such methods often face problems such as data scarcity and high annotation costs. In addition, the unseen categories in practical application scenarios are usually more valuable and require high-quality recognition rates. For example, in planetary image perception, the available image data is relatively scarce, and the zero-shot classification model can obtain the inference ability for unseen category images during the training of existing data. The perception of such new categories may bring more valuable scientific discoveries and promote the in-depth exploration of planetary-related knowledge. In environmental monitoring, natural or human-induced changes may form new land cover types, and the zero-shot learning classification model can quickly adapt to and accurately classify these new categories. Similarly, in disaster management, the model can efficiently identify unseen features or objects, thereby accelerating the response and decision-making speed for disasters. Therefore, designing a zero-shot image classification method that can identify unseen categories is of great significance and value.
[0004] Most existing zero-shot image classification methods separate feature extraction and feature alignment, with the focus on designing a network for feature alignment. The process of feature extraction is usually carried out independently, that is, an image is input into a pre-trained model to obtain visual features, and a text is input into a natural language model to obtain semantic features. This modular separation design is difficult to adjust visual and semantic features and is less likely to achieve good feature alignment. In addition, the above-mentioned feature extraction or feature alignment network is usually designed based on the Transformer backbone network, which usually has a high computational complexity due to its quadratic complexity, especially when modeling long sequence inputs. This makes it difficult to deploy zero-shot image classification methods based on Transformer on devices with limited resources, restricting their application prospects and value.
[0005] The article "MFINet: A Novel Zero-Shot Remote Sensing Scene Classification Network Based on Multimodal Feature Interaction" (《IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing》, 2024, 17: 11670-11684) proposed a zero-shot remote sensing scene classification method based on cross-modal information interaction to overcome the problem of inconsistent class structures between visual images and auxiliary semantics in zero-shot image classification. This method maps visual features and semantic features into the latent space respectively, and adds information interaction between visual and semantic features in the latent space to fully exploit the useful information of visual semantic features and improve the accuracy of zero-shot remote sensing scene classification. However, the drawback of this method is that it does not consider the computational resource limitations on actual devices, restricting the practical application of zero-shot image classification algorithms. Summary of the Invention
[0006] The purpose of this application is to overcome the above defects. This application proposes a lightweight zero-shot image classification method based on a state space model, including:
[0007] Input the unseen class test samples into the trained zero-shot image classification network model to output the classification results of unseen class images;
[0008] The zero-shot image classification network model includes a visual encoder, a semantic encoder, a multimodal fusion module, and a classification module; among them,
[0009] The visual encoder is used to input an image and output visual features;
[0010] The semantic encoder is used to input text and output semantic features;
[0011] The multi-modal fusion module is used to input visual features and semantic features and output features that fuse vision and semantics;
[0012] The classification module is used to input features that fuse vision and semantics and output the classification result of the image;
[0013] The multi-modal fusion module includes a Dual-Mamba2 block, two Mamba2 blocks, and a feed-forward neural network; the Dual-Mamba2 block is cascaded with two parallel Mamba2 blocks and then connected to the feed-forward neural network;
[0014] The Dual-Mamba2 block includes a mapping layer, a convolutional layer, an activation function, a state space model, and a normalization layer;
[0015] The working process of the Dual-Mamba2 block includes:
[0016] The input visual features are divided into a first branch and a second branch; the input semantic features are divided into a third branch and a fourth branch;
[0017] The visual features in the first branch pass through a mapping layer and an activation function to obtain a first intermediate result;
[0018] The visual features in the second branch pass through a mapping layer and an activation function to obtain a second intermediate result; the second intermediate result includes auxiliary parameters containing image features and second intermediate visual features; the second intermediate visual features of the second intermediate result are replaced with the third intermediate semantic features of the third intermediate result and then input into the state space model; the output result of the state space model is connected to the first intermediate result and then passes through a normalization layer and a mapping layer to output the semantic features of the image;
[0019] The semantic features in the third branch pass through a mapping layer and an activation function to obtain a third intermediate result; the third intermediate result includes auxiliary parameters containing text features and third intermediate semantic features; the third intermediate semantic features of the third intermediate result are replaced with the second intermediate visual features of the second intermediate result and then input into the state space model; the output result of the state space model is connected to the fourth intermediate result and then passes through a normalization layer and a mapping layer to output the visual features of the text;
[0020] The semantic features in the fourth branch pass through a mapping layer and an activation function to obtain a fourth intermediate result.
[0021] As an improvement of the above method, the visual encoder is VMamba.
[0022] As an improvement of the above method, the semantic encoder is Mamba.
[0023] As an improvement of the above method, the loss function L of the zero-shot image classification network model zsl is:
[0024] L zsl = L c + λ · L r
[0025] where L r is a class regression loss function based on the L2 norm:
[0026]
[0027] where y is the text information of the image; is the output representation of the image; ||·||2 represents the L2 norm;
[0028] L c is a class semantic vector loss function based on the cross-entropy loss:
[0029]
[0030] where is the class representation of the image; C S is the total number of classes of visible class images in the dataset;
[0031] λ is the weight parameter of the class regression loss function.
[0032] As an improvement of the above method, the training of the zero-shot image classification network model adopts an end-to-end training method, uses the gradient descent method, and cyclically updates the parameters in the zero-shot image classification network model to achieve the following:
[0033] 1) Set the initial value of the learning rate, the maximum number of iterations, and the batch size, and initialize the network training parameters;
[0034] 2) Divide the training samples into batches according to the batch size, and sequentially input them into the zero-shot image classification network model to backpropagate and update the gradient of the total loss function, and update the network training parameters in the forward propagation according to the direction of the gradient descent; perform the above calculations for each batch input into the network until the entire training set samples are input into the network for traversal to complete one training;
[0035] 3) Cycle the process of 2) until the maximum number of iterations is reached to obtain the trained zero-shot image classification model.
[0036] This application also provides a lightweight zero-shot image classification system based on the state space model, which is implemented based on the above method. The system includes:
[0037] A sample classification module, which is used to input unseen-class test samples into a trained zero-shot image classification network model and output the unseen-class image classification results;
[0038] A model training module, which is used to train the zero-shot image classification network model.
[0039] Compared with the prior art, the advantages of this application are as follows:
[0040] 1. Since the present invention constructs a zero-shot image classification network including a visual encoder, a semantic encoder, a multi-modal fusion module and a classifier, and integrates the feature extraction network, the feature alignment network and the classifier into a unified framework, the input data can directly obtain the results at the output end through the trained zero-shot image classification network, avoiding manual preprocessing and subsequent processing, and improving the zero-shot classification speed of images.
[0041] 2. The Dual-Mamba2 module designed by the present invention can effectively extract the interaction information between the two modalities by using the exchange of image and text parameters.
[0042] 3. The multi-modal fusion module designed by the present invention effectively realizes the feature interaction between the visual and semantic modalities through the collaborative cascade configuration of the Dual-Mamba2 module and the Mamba2 block.
[0043] In summary, compared with the prior art, the present invention can better utilize the shared information between images and auxiliary information, and under the zero-shot setting, it can not only keep the overall classification accuracy of unseen classes at a high level, but also effectively reduce the computational complexity of the model, reduce the computing power requirements of the model for the hosting platform, and is more conducive to the deployment of the zero-shot image classification method on hardware devices, and can provide better services for more advanced applications such as ground object monitoring and change detection. Description of the Drawings
[0044] Figure 1 Shown is a schematic flow diagram of a lightweight zero-shot image classification method based on a state space model;
[0045] Figure 2 Shown is a schematic diagram of the Dual-Mamba2 module;
[0046] Figure 3 Shown is a block diagram of a zero-shot image classification model;
[0047] Figure 4 Shown is an example diagram of the image categories of the self-built Mars scene; Detailed Embodiments
[0048] The technical solution of the present application will be described in detail below with reference to the accompanying drawings.
[0049] The present application proposes a lightweight zero-shot image classification method and system based on a state space model (State Space Model, SSM) to make full use of the extracted visual and semantic information, enhance the information interaction between visual and semantic features, reduce the computational complexity of the model, and improve the zero-shot image classification effect.
[0050] The technical idea to achieve the above purpose is: to design the visual semantic feature extraction and zero-shot classification as a unified process, respectively mine the input image features and text information through the visual and semantic encoders based on the state space model, and realize the interactive learning between visual and semantic features through the feature alignment network based on the state space model, so that the obtained visual-semantic features are more discriminative, obtain a more accurate feature expression effect, effectively reduce the computational complexity of the model, and at the same time ensure the classification accuracy.
[0051] Embodiment 1
[0052] As Figure 1 shown, the lightweight zero-shot image classification method based on the state space model provided by the present invention includes the following steps:
[0053] 1. As Figure 3 shown, adopt an end-to-end training method to design a zero-shot image classification model including four modules: a visual encoder (Visual StateSpace Model, VMamba), a semantic encoder (Mamba), a multi-modal fusion module (Dual-stream Mamba Feature Fusion Module, DM2F) based on the Dual-Mamba2 unit, and a classification module.
[0054] 1) Build a visual encoder based on the state space model.
[0055] Extract image features in a bidirectional scanning manner and build a visual encoder VMamba based on Mamba.
[0056] 2) Build a semantic encoder based on the state space model.
[0057] Extract semantic features in the text by a selective scanning mechanism and build a semantic encoder Mamba based on the state space model.
[0058] 3) Construct a Dual-Mamba2 module.
[0059] As Figure 2As shown, the Dual-Mamba2 module is a dual-branch network, where each branch consists of a cascaded input mapping layer, a one-dimensional convolutional layer, an activation function (SiLU), a state space model, a normalization layer, and an output mapping layer.
[0060] First, the image and text features are processed on two mapping layers respectively to obtain the auxiliary parameters A v , B v and C v containing image features, and the auxiliary parameters A t , B t and C t containing text features. The specific formulas are as follows:
[0061]
[0062] where represents the calculation formula of the state space dual layer.
[0063] The output contains two branches, and each branch is divided into two information flows. One information flow is processed by a convolutional layer and then enters the core state space model through an activation function. The specific formulas are as follows:
[0064]
[0065]
[0066] When entering the state space model, the input positions of X v and X t are swapped and passed to the state space model for calculation.
[0067] The other information flow only passes through the activation function. The specific formulas are as follows:
[0068]
[0069] Subsequently, the two are connected, passed through the normalization layer and the output mapping layer to obtain the final output. The specific formulas are as follows:
[0070]
[0071] Among them, there are two types of output sequences Z, namely the visual features Z v corresponding to the text and the semantic features Z t corresponding to the image; is the output mapping matrix, is the output bias matrix, and Ω(·) represents the normalization function. The normalized output further refines the fused features to ensure robust and consistent multimodal alignment.
[0072] 4) Construct the multimodal fusion module DM2F.
[0073] This application introduces a novel hierarchical structure, which cascades a Dual-Mamba2 module with two Mamba2 blocks. The specific formula is as follows:
[0074] Y v ′ = Γ(A v ′, Z v , B v ′, C v ′)
[0075] Y t ′ = Γ(A t ′, Z t , B t ′, C t ′)
[0076] Among them, Γ(·) is the calculation formula of Mamba2, A v ′, B v ′, C v ′ and A t ′, B t ′, C t ′ are auxiliary parameters of the image and text respectively, Z v and Z t are the interaction features obtained by the Dual-Mamba2 module respectively.
[0077] 2. Train the zero-shot image classification model:
[0078] 1) Construct the overall loss function L of the network model zsl = L c + λ · L r , where L c is the class semantic vector loss function based on cross-entropy loss, L r is the class regression loss function based on the L2 norm, and λ is the weight parameter of the class regression loss function;
[0079] The class semantic vector loss function L c , and the class regression loss function L r are respectively expressed as follows:
[0080]
[0081] Among them, y is the text information of the image x, is the output representation of the image x. represents the class representation of the image, C S represents the total number of classes of visible class images in the dataset; ||·||2 represents the L2 norm.
[0082] 2) Initializing the parameters of the zero-shot image classification network model includes the auxiliary parameters A1, X1, B1, and C1 in the visual encoder VMamba, the auxiliary parameters A2, X2, B2, and C2 in the semantic encoder Mamba, and the auxiliary parameters A v , X v , B v and C v , A t , X t , B t and C t ;
[0083] 3) Input the visible class training samples and class semantic information into the zero-shot image classification network model. Using the gradient descent method, cycle through and update the parameters in the zero-shot image classification network to reduce the gradient value of the total loss function until the maximum number of iterations T is reached, obtaining the trained zero-shot image classification network model;
[0084] Among them, using the gradient descent method to cycle through and update the parameters in the zero-shot image classification network is implemented as follows:
[0085] a) Set the initial learning rate α, the maximum number of iterations T, the batch size B, and initialize the network training parameters;
[0086] b) Divide the training samples into batches according to the batch size, and sequentially input them into the zero-shot image classification network for backpropagation to update the gradient of the total loss function and perform forward propagation to update the network training parameters according to the descending direction. Perform the above calculations for each batch input into the network until the entire training set of samples has been input into the network for traversal, completing one training;
[0087] c) Cycle through the process of b) until the maximum number of iterations T is reached, obtaining the trained zero-shot image classification model.
[0088] 3. Input the unseen class test samples into the trained zero-shot image classification network model and output the unseen class image classification results.
[0089] Example 2
[0090] In this example, the zero-shot Mars scene classification standard dataset ZS-Mars is used for classification. The images and corresponding ground truth labels in this dataset are as Figure 4 shown, which contains 10 types of scenes.
[0091] Referring to Figure 1 , the specific implementation steps of this example are as follows:
[0092] Step 1: Randomly divide the training and test samples.
[0093] 1.1) For the ZS-Mars dataset, divide the training (visible classes) and test (unseen classes) samples according to the ratios of 6 / 4, 7 / 3, and 8 / 2, i.e., 6 visible classes / 4 unseen classes, 7 visible classes / 3 unseen classes, 8 visible classes / 2 unseen classes;
[0094] 1.2) Encode the label of each training sample into a vector y in one-hot form i , and encode the label of each test sample into a vector q in one-hot form i .
[0095] Step 2: Build a vision encoder based on the state space model.
[0096] 2.1) Extract image features in a bidirectional scanning manner, and build a vision encoder VMamba based on Mamba. The input image size is 224×224×3;
[0097] 2.2) From the samples shown in Figure 3 , obtain the training set and test set samples and their corresponding labels according to the division in Step 1. The training set is denoted as X = {x1,...,x train}, and the label is Y = {y1,...,y train}, where x i is any training sample. For different visible / invisible class ratios, the subscript train corresponds to the product of the number of visible classes and the number of scenes in each class. The label y i corresponding to x i ∈ {1, 2,..., 6}; the test set is denoted as P = {p1,...,p test}, and the label is Q = {q1,...,q test}, where p i is any test sample. For different visible / invisible class ratios, the subscript test corresponds to the product of the number of invisible classes and the number of scenes in each class. The label q i corresponding to p i ∈ {1, 2, 3, 4};
[0098] Step 3: Build a semantic encoder based on the state space model.
[0099] 3.1) Extract semantic features in the text using a selective scanning mechanism, and build a semantic encoder Mamba based on the state space model.
[0100] From Figure 3In the sample labels shown, the corresponding labels of the training set and the test set are obtained according to the division in step 1. Among them, for different visible / invisible category ratios, the number of training samples corresponds to the product of the number of visible categories and the number of scenarios in each category, the test samples correspond to the product of the number of invisible categories and the number of scenarios in each category, and the remaining sample labels are used as the test set;
[0101] Step 4: Construct a multi-modal fusion module based on the state space model.
[0102] Refer to Figure 2 , and the specific implementation of this step is as follows:
[0103] 4.1) Set the Dual-Mamba2 module, which consists of a mapping layer, a convolutional layer, an activation function, a state space model, and a normalization layer. The specific calculation process is as follows:
[0104] First, establish the mapping layer Set four mapping modules All are composed of linear matrices, and the specific formula is as follows:
[0105]
[0106] Among them, there are two types of input sequences X, namely visual feature X v and semantic feature X t ; W v in , W t in are the corresponding input mapping matrices, and b v in , b t in are the corresponding input bias matrices.
[0107] The output contains two branches, and each branch is divided into two information flows. One information flow is processed by the convolutional layer and then enters the core state space model module through the activation function. When entering the state space model, the input positions of X v and X t are exchanged, and the specific formula is as follows:
[0108]
[0109] The other information flow only passes through the activation function, and the specific formula is as follows:
[0110]
[0111] Subsequently, the two are connected, and after passing through the normalization layer and the output mapping layer, the final output is obtained, and the specific formula is as follows:
[0112]
[0113] Among them, there are two types of output sequences Z, namely the visual features Z corresponding to the text v and the semantic features Z corresponding to the image t ; W v out and W t out are the output mapping matrices, is the output bias matrix, and Ω(·) represents the normalization function.
[0114] 4.2) Set up a multi-modal fusion module, which consists of a Dual-Mamba2 module and two parallel Mamba2 blocks.
[0115] First, establish a Dual-Mamba2 module to obtain the visual features Z corresponding to the text v and the semantic features Z corresponding to the image t , and the specific formula is as described in 4.1).
[0116] Subsequently, input the above features into the parallel Mamba2 blocks respectively. The structure of the Mamba2 block is similar to that of the Dual-Mamba2 module, including a mapping layer, a convolutional layer, an activation function, a state space model, and a normalization layer. The calculation processes of the Mamba2 blocks in parallel are the same, and the specific calculation process is as follows:
[0117] Establish a mapping layer Set up four mapping modules All are composed of linear matrices, and the specific formula is as follows:
[0118] Z′ = W in ·Z + b in
[0119] Among them, Z is the input sequence, that is, the visual feature Z v or the semantic feature Z t ; W in is the input mapping matrix, and b in is the input bias matrix.
[0120] The output is divided into two information flows. One information flow is processed by the convolutional layer and then enters the core state space model module through the activation function. The specific formula is as follows:
[0121] Z c = Conv(Z′)
[0122] Z ssm = SSM(SiLU(Z c ))
[0123] Another information flow only passes through the activation function, and the specific formula is as follows:
[0124] Z a = SiLU(Z′)
[0125] Subsequently, the two are concatenated, passed through the normalization layer and the output mapping layer to obtain the final output. The specific formula is as follows:
[0126] Z out = Z ssm + Z a
[0127] Y = W out ·Ω(Z out ) + b out
[0128] Among them, Y is the output sequence, that is, the visual feature Y v or the semantic feature Y t ; W out is the output mapping matrix, b out is the output bias matrix, and Ω(·) represents the normalization function.
[0129] Finally, the outputs of the Dual-Mamba2 module and the parallel Mamba2 block are concatenated, passed through the feed-forward neural network layer to obtain the final output. The specific formula is as follows:
[0130] Y v ′ = FNN(Z v + Y v )
[0131] Y t ′ = FNN(Z t + Y t )
[0132] Among them, FNN(·) represents the calculation formula of the feed-forward neural network layer.
[0133] Step 5: Construct a zero-shot image classification model based on the state space model.
[0134] 5.1) Cascade the above visual encoder, semantic encoder, and multi-modal fusion module in sequence to form a visual-semantic feature extraction and alignment encoder;
[0135] 5.2) Construct a classifier g η . g η includes 1 fully connected layer, and the size of the transformation matrix is used to convert the output features of the above visual-semantic feature extraction and alignment encoder into probability vectors belonging to each category for judging the category to which the sample belongs.
[0136] Step 6: Construct the loss function of the zero-shot image classification model based on the state space model.
[0137] 6.1) Construct the class regression loss function L based on the L2 norm r , which is used to constrain the image representation to ensure that it is accurately mapped to the corresponding text embedding. The specific formula is as follows:
[0138]
[0139] where x represents the visual feature representation of the image, and y represents the text information representation of the image;
[0140] 6.2) Construct the cross-entropy loss function L c , which is used to force the image to have the highest compatibility score with its corresponding class semantic vector. The specific formula is as follows:
[0141]
[0142] where represents the class representation of the image, and C S represents the total number of classes of visible class images in the dataset;
[0143] Step 7: Use the training samples and adopt the gradient descent method to train the parameters of the zero-shot image classification model.
[0144] 7.1) Set the initial value of the learning rate α, the maximum number of iterations T, and the batch size B;
[0145] 7.2) Initialize the parameters of the zero-shot image classification network model, including the auxiliary parameters A1, X1, B1, and C1 in the visual encoder VMamba, the auxiliary parameters A2, X2, B2, and C2 in the semantic encoder Mamba, and the auxiliary parameters A v , X v , B v and C v , A t , X t , B t and C t , as well as the parameter η in the classifier;
[0146] 7.3) Divide the training samples into batches according to the batch size, and sequentially input them into the unbalanced hyperspectral image classification network to backpropagate and update the gradient of the total loss function . According to the direction of descent, perform forward propagation to update the network training parameters. Perform the above calculations for each batch input to the network until the entire training set samples are input to the network for traversal, and complete one training;
[0147] 7.4) Loop the process in 7.3) until the maximum number of iterations T is reached to obtain the trained zero-shot classification model.
[0148] Step 8: Classify the test samples using the trained zero-shot image classification model.
[0149] Input each test sample into the trained zero-shot image classification model to obtain the probability p(y|x) of the label to which each test sample belongs:
[0150]
[0151] where softmax is the normalization function, L zsl is the total loss function, x is the test sample, y c is the label corresponding to the c-th class test sample, z c is the semantic feature corresponding to the c-th class test sample, C is the total number of test samples, A1, X1, B1, and C1 are auxiliary parameters in the visual encoder VMamba, A2, X2, B2, and C2 are auxiliary parameters in the semantic encoder Mamba, A v 、X v 、B v and C v 、A t 、X t 、B t and C t are auxiliary parameters of the multi-modal fusion module DM2F, and η is the trainable parameter of the classifier.
[0152] The effects of the present invention can be further illustrated by the following simulation results.
[0153] Test data:
[0154] Use ZS-Mars data, with the semantic information being the sentence vector representing the long sequence input, and the images are tested according to 3 visible / invisible ratio divisions. That is, when the visible / invisible ratio is 6 / 4, 6 classes are selected as the visible class images as the training set, and the remaining 4 classes are the unseen class images as the test set; when the visible / invisible ratio is 7 / 3, 7 classes are selected as the visible class images as the training set, and the remaining 3 classes are the unseen class images as the test set; when the visible / invisible ratio is 8 / 2, 8 classes are selected as the visible class images as the training set, and the remaining 2 classes are the unseen class images as the test set. The specific division of visible / invisible classes is shown in Table 1.
[0155] Test environment:
[0156] Adopt the Linux system, NVIDIA GeForce GTX 4090 GPU, and Pytorch deep learning framework.
[0157] Division of visible / invisible categories in Table 1
[0158]
[0159]
[0160] Simulation content:
[0161] Simulation 1: Under the above conditions, use the existing zero-shot image scene classification method MFINet to classify the ZS-Mars data, and the results are shown in Table 2.
[0162] Simulation 2: Under the above conditions, use the method of the present invention to classify the ZS-Mars data, and the results are shown in Table 2.
[0163] It can be found from the comparison in Table 2 that the method proposed by the present invention obtains higher classification accuracy with smaller computational complexity and achieves better classification performance.
[0164] Evaluation index:
[0165] Repeat the above two methods 10 times respectively using the same random numbers, calculate the overall classification accuracy OA, and the test results are shown in Table 2.
[0166] The formula for the index is as follows:
[0167]
[0168] Among them, TP is the number of samples correctly classified into this category, FN is the number of samples of this category that are wrongly classified into other categories, TN is the number of non-class samples classified into other categories, and FP is the number of other samples that are wrongly classified into this category.
[0169] Table 2 Simulation experiment results of the present invention and the comparative method
[0170]
[0171] It can be seen from Table 2 that compared with the current relatively advanced MFINet method, the present invention can achieve higher classification accuracy with less resource consumption under 3 visible / invisible ratios.
[0172] Example 3
[0173] This application also provides a lightweight zero-shot image classification system based on a state space model, which is implemented based on the above method. The system includes:
[0174] A sample classification module, configured to input an unseen category test sample into a trained zero-shot image classification network model and output an unseen category image classification result;
[0175] A model training module for training the zero-shot image classification network model.
[0176] The present application may also provide a computer device, including: at least one processor, a memory, at least one network interface, and a user interface. Each component in the device is coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus.
[0177] Among them, the user interface may include a display, a keyboard, or a pointing device. For example, a mouse, a trackball, a touchpad, or a touch screen, etc.
[0178] It can be understood that the memory in the disclosed embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memories described herein are intended to include but not be limited to these and any other suitable types of memories.
[0179] In some embodiments, the memory stores the following elements, executable modules, or data structures, or subsets thereof, or extended sets thereof: an operating system and an application program.
[0180] Among them, the operating system includes various system programs, such as the framework layer, the core library layer, the driver layer, etc., which are used to implement various basic services and handle hardware-based tasks. The application programs include various application programs, such as Media Player, Browser, etc., which are used to implement various application services. The program for implementing the method of the embodiments of the present disclosure may be included in the application programs.
[0181] In the above-mentioned embodiments, the program or instruction stored in the memory may also be called. Specifically, it may be the program or instruction stored in the application program. The processor is used for:
[0182] Executing the steps of the above method.
[0183] The above method can be applied to or implemented by the processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instruction in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed above. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Combining the steps of the above-disclosed method can be directly embodied as being completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0184] It can be understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.
[0185] For software implementation, the technology of this application can be implemented by executing the functional modules of this application (such as procedures, functions, etc.). The software code can be stored in a memory and executed by a processor. The memory can be implemented inside or outside the processor.
[0186] This application can also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, the various steps in the above method embodiments can be implemented.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and not to limit them. Although this application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of this application does not depart from the spirit and scope of the technical solutions of this application, and they should all be covered within the scope of the claims of this application.
Claims
1. A lightweight zero-shot image classification method based on a state space model, comprising: Inputting an unseen class test sample into a trained zero-shot image classification network model to output an unseen class image classification result; The zero-shot image classification network model includes a visual encoder, a semantic encoder, a multi-modal fusion module, and a classification module; wherein, The visual encoder is used to input an image and output visual features; The semantic encoder is used to input text and output semantic features; The multi-modal fusion module is used to input visual features and semantic features and output features that fuse vision and semantics; The classification module is used to input features that fuse vision and semantics and output the classification result of the image; The multi-modal fusion module includes a Dual-Mamba2 block, two Mamba2 blocks, and a feed-forward neural network; the Dual-Mamba2 block and two parallel Mamba2 blocks are cascaded and then connected to the feed-forward neural network; The Dual-Mamba2 block includes a mapping layer, a convolutional layer, an activation function, a state space model, and a normalization layer; The working process of the Dual-Mamba2 block includes: The input visual features are divided into a first branch and a second branch; the input semantic features are divided into a third branch and a fourth branch; In the first branch, the visual features pass through a mapping layer and an activation function to obtain a first intermediate result; In the second branch, the visual features pass through a mapping layer and an activation function to obtain a second intermediate result; the second intermediate result includes auxiliary parameters containing image features and second intermediate visual features; after replacing the second intermediate visual features of the second intermediate result with the third intermediate semantic features of the third intermediate result, it is input into the state space model; the output result of the state space model is connected to the first intermediate result and then passes through a normalization layer and a mapping layer to output the semantic features of the image; In the third branch, the semantic features pass through a mapping layer and an activation function to obtain a third intermediate result; the third intermediate result includes auxiliary parameters containing text features and third intermediate semantic features; after replacing the third intermediate semantic features of the third intermediate result with the second intermediate visual features of the second intermediate result, it is input into the state space model; the output result of the state space model is connected to the fourth intermediate result and then passes through a normalization layer and a mapping layer to output the visual features of the text; In the fourth branch, the semantic features pass through a mapping layer and an activation function to obtain a fourth intermediate result.
2. The lightweight zero-shot image classification method based on the state space model according to claim 1, characterized in that: The visual encoder is VMamba.
3. The lightweight zero-shot image classification method based on the state space model according to claim 1, characterized in that: The semantic encoder is Mamba.
4. The lightweight zero-shot image classification method based on the state space model according to claim 1, characterized in that: The loss function L of the zero-shot image classification network model zsl is as follows: L zsl = L c + λ·L r Among them, L r is a regression-like loss function based on the L2 norm: where y is the text information of the image; is the output representation of the image; ||·||2 represents the L2 norm; L c is a class semantic vector loss function based on cross-entropy loss: Among them, is the representation of the category to which the image belongs; C S is the total number of categories of visible-class images in the dataset; λ is the weight parameter of the class regression loss function.
5. The lightweight zero-shot image classification method based on the state space model according to claim 1, characterized in that: The training of the zero-shot image classification network model adopts an end-to-end training method. Using the gradient descent method, the parameters in the zero-shot image classification network model are cyclically updated to achieve the following: 1) Set the initial value of the learning rate, the maximum number of iterations, and the batch size, and initialize the network training parameters; 2) Divide the training samples into batches according to the batch size, and sequentially input them into the zero-shot image classification network model for backpropagation to update the gradient of the total loss function, and perform forward propagation to update the network training parameters according to the direction of gradient descent; perform the above calculations for each batch input into the network until the entire training set of samples is traversed through the network, completing one training; 3) Repeat the process in 2) until the maximum number of iterations is reached, obtaining a trained zero-shot image classification model.
6. A lightweight zero-shot image classification system based on a state space model, implemented based on any one of the methods described in claims 1-5, characterized in that, The system includes: A sample classification module, configured to input unseen class test samples into the trained zero-shot image classification network model and output unseen class image classification results; and A model training module, configured to train the zero-shot image classification network model.
Citation Information
Patent Citations
Generalized zero sample image classification method based on enhanced multi-modal alignment
CN113139591A
High-resolution remote sensing image target detection method based on multi-scale network
CN118485927A
Cross-modal multilayer fusion emotion recognition method and system
CN118861773A
Vision and language fused multi-modal large model system
CN119227744A
Zero sample image classification system and method
CN119295790A
Cited By
Pathological image classification method based on state space duality
CN119048825A
A pathological image classification method based on state-space duality
CN119048825B
Zero sample image classification method and system based on multistage modulation and dynamic fusion
CN122493099A