A combined zero-shot image classification method based on hierarchical feature fusion
By adopting a hierarchical feature fusion method in the combined zero-sample image classification task, combining cross-modal interaction and feature decoupling modules, the problem of the existing technology being difficult to adapt to specific scenarios is solved, and higher classification accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510005907.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In the combined zero-sample image classification task, it is difficult to adapt to the data distribution and task requirements of specific scenarios, resulting in unsatisfactory model performance.
Using a method based on hierarchical feature fusion, we will deeply explore and fusion the features of each level of the CLIP visual encoder, combine the cross-modal interaction module and feature decoupling module to optimize the embedded representation to improve the generalization capabilities of the model.
It significantly improves the classification accuracy and stability of the model in complex attribute-object combination scenarios, and can more accurately identify images of unseen categories, meeting the stringent performance standards for practical applications.
Smart Images

Figure CN119478551B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and artificial intelligence, and particularly relates to a combined zero-shot image classification method based on hierarchical feature fusion. Background Art
[0002] Existing combined zero-shot image classification methods at home and abroad mainly fall into four categories. The first category of methods is the model based on single-concept classifiers. This method regards attributes and objects as equally important concepts. By learning classifiers for each concept and combining attributes and objects, it aims to identify the integrated concept, that is, the new attribute-object pair. However, this method has obvious drawbacks. It treats attributes and objects in a general way, ignoring the fact that attributes are highly visually related to objects and context-dependent. Therefore, it often performs poorly. The second category of methods is the image-combination compatibility model. This method treats the attribute-object pair as a whole and directly learns the compatibility feature representation between them and the image. Due to the introduction of deep learning technology, such methods have shown some improvements in model performance. However, these methods still do not clearly distinguish between attributes and objects, and attributes and objects are still intertwined and affect each other. The third category of methods is the attribute-object explicit decoupling model. It usually uses a spatial embedding method to explicitly decouple attributes and objects. Instead of regarding attributes and objects as the same concept, it successfully decouples attributes and objects in the embedding space, while emphasizing their differences and connections. This method has brought performance improvements. However, information loss inevitably occurs during the embedding process, and some subtle but crucial features or relationships may become blurred or lost, making it difficult for the model to accurately represent all attribute-object combinations. The fourth category of methods is the CLIP-based model. CLIP was proposed by the OpenAI team in 2021, which provides a brand-new idea and paradigm for multi-modal information fusion processing.
[0003] However, although CLIP shows quite remarkable learning efficiency and generalization characteristics in the pre-training stage and can widely adapt to general multi-modal data processing requirements, its limitations will become prominent when directly deployed for specific downstream tasks. Taking the combined zero-shot image classification task as an example, this task requires the model to accurately identify image category combinations that have not appeared in the training set. Since the original CLIP model is not carefully optimized for such specific scenarios and is difficult to adapt to its special data distribution and task requirements, the model performance is not satisfactory.
[0004] To resolve this dilemma, a series of improvement solutions have emerged in the academic and industrial circles. The core strategy is to perform targeted fine-tuning on the embedding vectors generated by CLIP. In principle, fine-tuning aims to reshape the text embeddings produced by the model to fit the specific task objectives, thereby enhancing the model's performance in specific scenarios. However, existing fine-tuning methods have obvious defects. This method overly emphasizes the optimization of the output ends of the visual encoder and the text encoder, ignoring the key information contained in the feature encoding process. Each layer of the visual encoder contains rich and valuable intermediate features in the feature extraction and transformation links. These features have not been properly mined and effectively utilized, directly weakening the model's generalization ability to unknown combined scenarios. As a result, when the model encounters new image-text combinations, the classification accuracy and stability are poor, and it cannot fully meet the stringent performance standards of practical applications. Summary of the Invention
[0005] The purpose of the present invention is to propose a combined zero-shot image classification method based on hierarchical feature fusion, which fully mines the key information contained in each layer of the CLIP visual encoder in the feature extraction and transformation links, and uses it for cross-modal interaction between text features and visual features, so that the generated embedding representation better meets the requirements of the combined zero-shot classification task, in order to improve the model's performance in dealing with complex attribute-object combination scenarios.
[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A combined zero-shot image classification method based on hierarchical feature fusion includes the following steps:
[0008] Step 1. Build a combined zero-shot image classification model based on hierarchical feature fusion, including a CLIP text encoder, a CLIP visual encoder, a multi-layer visual feature fusion module, a cross-modal interaction module, a decoupling module, and a loss calculation module;
[0009] Step 2. Generate respective preliminary embedding representations, i.e., word embeddings, for attributes, objects, and combinations. The preliminary embedding representations will be updated as parameters during the backpropagation process; combine the preliminary embedding representations of attributes, objects, and combinations with the prompt prefix embedding, and obtain the text features of attributes, objects, and combinations through the pre-trained CLIP text encoder;
[0010] Step 3. Take the images in the training set as input images and input them into the CLIP visual encoder, process the features of the images layer by layer, and extract features at different levels from the CLIP visual encoder;
[0011] Step 4. Screen out the initial layer features, intermediate layer features, and final layer features from the CLIP visual encoder, calculate the weights of different-level features, i.e., the weighting coefficients, through a multi-layer visual feature fusion module, perform weighted combination on features at different levels, and construct visual fusion features;
[0012] Step 5. Input the text features obtained in Step 2 and the visual fusion features constructed in Step 4 into the cross-modal interaction module, and through the cross-attention mechanism, achieve the deep fusion of visual fusion features and text features to obtain optimized text features of attributes, objects, and combinations;
[0013] Step 6. Decouple the image features output by the last layer of the CLIP visual encoder. The decoupling module uses three trainable multi-layer perceptrons (MLPs) to extract attribute visual features, object visual features, and combination visual features from the image features output by the last layer of the CLIP visual encoder respectively;
[0014] Step 7. The loss calculation module calculates the similarity between the attribute, object, and combination visual features obtained after decoupling in Step 6 and the optimized text features of attributes, objects, and combinations obtained in Step 5 using cosine similarity respectively to obtain prediction values, and measures the difference between the prediction values and the true labels through a classification loss function to form branch losses. At the same time, a covariance loss term is introduced to optimize the decoupling effect, and the total loss is obtained through weighted summation;
[0015] Step 8. Through the backpropagation mechanism of the total loss, jointly update the weighting coefficients of the multi-layer visual feature fusion module, the parameters of the MLP for decoupling, and the initial embedding representations of attributes, objects, and combinations;
[0016] Step 9. Repeat Steps 3 to 8, and evaluate the hierarchical feature fusion-based combined zero-shot image classification model after each iteration on the validation set until the loss function converges. Obtain the trained hierarchical feature fusion-based combined zero-shot image classification model and the optimized embedding representations of attributes, objects, and combinations according to the validation set results;
[0017] Step 10. Use the trained hierarchical feature fusion-based combined zero-shot image classification model to classify the input image.
[0018] The present invention has the following advantages:
[0019] As described above, the present invention relates to a combined zero-shot image classification method based on hierarchical feature fusion. First, by deeply mining and fusing the features of each layer of the CLIP visual encoder, the multi-dimensional information of the image from local details to overall semantics is effectively integrated. The multi-level feature fusion method enables the model to more comprehensively understand the image features when facing complex image content, avoiding classification errors caused by the limitations of single-level features. When processing images containing multiple objects with complex attribute associations between objects, the fused visual features can simultaneously capture the detailed information such as the edges and textures of the objects, as well as the high-level information such as the spatial relationships and semantic categories between the objects, thus providing a solid foundation for accurate classification.
[0020] Secondly, the introduction of the cross-modal interaction module greatly enhances the correlation between text features and visual features. Using the cross-attention mechanism, the text features can be dynamically adjusted according to the information in the visual fusion features, enabling the text embedding to better adapt to the actual visual content of the image. This not only improves the accuracy of the text in describing the image semantics, but also enables the model to more accurately classify images of unseen classes according to the text prompts in the combined zero-shot classification task.
[0021] Furthermore, the feature decoupling operation improves the adaptability of the model in the zero-shot combination task. By using three trainable multi-layer perceptrons to extract attribute, object, and combined visual features respectively, the independence of these key elements in the feature representation is ensured. This helps the model to flexibly utilize the existing feature knowledge for reasoning and classification when facing new attribute-object combinations, without being interfered by the confusion or redundant information between the features.
[0022] In addition, the design of the loss calculation and optimization objective makes the training process of the model more scientific and effective. By using the classification loss function to measure the difference between the predicted value and the true label, the model can be directly guided to optimize in the direction of improving classification accuracy. At the same time, the covariance loss is introduced to optimize the decoupling effect, further enhancing the independence and distinguishability of the features, thereby improving the overall performance of the model. During the backpropagation and parameter update process of the model, the parameters of multiple key modules are jointly optimized to ensure the collaborative work of each part of the model and the improvement of the overall performance.
[0023] Finally, the strategy of repeatedly iterating and optimizing the embedding representation ensures that the model can gradually converge to the optimal state. By continuously evaluating the model performance and adjusting the parameters on the validation set, the finally obtained model has excellent generalization ability and classification accuracy in the combined zero-shot image classification task. The method of the present invention is not only applicable to the combined zero-shot image classification scenario, but also can play an important role in special fields such as rare item recognition and emerging concept image classification, providing strong support for the application of image classification technology in more complex and diverse scenarios. Brief Description of the Drawings
[0024] Figure 1 It is a flowchart of the combined zero - shot image classification method based on hierarchical feature fusion in an embodiment of the present invention;
[0025] Figure 2 It is a system block diagram of the combined zero - shot image classification model based on hierarchical feature fusion in an embodiment of the present invention;
[0026] Figure 3 It is a structural schematic diagram of the multi - layer visual feature fusion module in an embodiment of the present invention; Detailed Embodiment
[0027] The present invention will be further described in detail below in conjunction with the drawings and the detailed embodiment:
[0028] In this embodiment, a combined zero - shot image classification method based on hierarchical feature fusion is proposed. Aiming at the technical difficulties of combined zero - shot image classification, a solution based on hierarchical feature fusion is innovatively proposed. By selecting features of different depths in the CLIP model visual encoder and performing hierarchical fusion on them, multi - level and multi - scale visual information is effectively extracted. The fused features are further used for cross - modal interaction between vision and semantics, significantly improving the model's understanding ability of complex combinations. In addition, the present invention effectively separates attribute and object features by introducing a specific loss term, avoiding the problem of feature aliasing. In the combined zero - shot image classification task, the method of the present invention demonstrates excellent generalization ability and recognition effect, and can significantly improve the classification performance of the model for unknown complex combination images.
[0029] As Figure 1 shown, a combined zero - shot image classification method based on hierarchical feature fusion specifically includes the following steps:
[0030] Step 1. Build a combined zero - shot image classification model based on hierarchical feature fusion, including a CLIP text encoder, a CLIP visual encoder, a multi - layer visual feature fusion module, a cross - modal interaction module, a decoupling module, and a loss calculation module.
[0031] Step 2. Preliminary construction of the embedding representation: For attributes, objects, and combinations, generate their respective preliminary embedding representations, i.e., word embeddings. The preliminary embedding representations will be updated as parameters during the backpropagation process; Combine the preliminary embedding representations of attributes, objects, and combinations with the prompt prefix embedding, and obtain the text features of attributes, objects, and combinations through the pre - trained CLIP text encoder to capture semantic information and enhance the expression ability of the text features.
[0032] Step 2.1. For the attribute set A, object set O, and combination set C, first obtain the preliminary embedding representations of each attribute, object, and their combinations through the word embedding function φ(·) of CLIP;
[0033] where A = {a1, a2, …, a m}, a1 to a m respectively represent the 1st to the mth attributes in the attribute set A, and m is the total number of attributes in the training set; O = {o1, o2, …, o n}, o1 to o n respectively represent the 1st to the nth objects in the object set O, and n represents the total number of objects in the training set; C = {c ij =(a i , o j )∣i ∈ [1, m], j ∈ [1, n]}, a i represents the ith attribute in the attribute set A, o j represents the jth object in the object set O, and c uj represents the combination of the attribute a i and the object o j .
[0034] The word embedding vectors of each attribute, object, and their combinations are respectively represented as:
[0035]
[0036] where and respectively represent the word embeddings of the attribute, object, and their combinations, that is, the preliminary embedding representations of the attribute, object, and combination. The preliminary embedding representations of the attribute, object, and their combinations will be updated as parameters in the subsequent backpropagation, and will be iteratively updated during the training of the combined zero-shot image classification model based on hierarchical feature fusion, and do not need to be regenerated.
[0037] Step 2.2. To enhance the semantic expression ability of the embedding, introduce the prompt prefix mechanism.
[0038] Process each word in the prompt prefix using the word embedding function to obtain the set P of prompt prefix embeddings, where P = {p1, p2, …, p k}, p1 to p k respectively represent the embeddings of the 1st to the kth context words in the set P, and k represents the total number of context words in the prompt prefix;
[0039] Concatenate the word embedding of the prompt prefix, that is, the prompt prefix embedding, with the word embeddings of each attribute, object, and their combinations to respectively construct the complete input text embeddings with the prompt prefix, which are represented as:
[0040]
[0041] Among them, The word embedding representing the prompt prefix and the attribute a i The word embedding of After splicing, the constructed complete input text, The word embedding representing the prompt prefix and the object o j The word embedding of After splicing, the constructed complete input text, The word embedding representing the prompt prefix and the attribute a i And the object o j The combined word embedding of After splicing, the constructed complete input text.
[0042] Step 2.3. Use the CLIP text encoder E text (·) to encode the complete input text embedding with the prompt prefix to obtain the text feature T, and the text feature T includes the attribute text feature Object text feature And the combined text feature Are respectively represented as:
[0043]
[0044] Among them, And Are respectively the text features of the attribute, object and their combination in the encoder output space, that is Figure 2 The attribute text vector, object text vector and combined text vector shown. And Combined with the context prompt information, that is, context vocabulary, to enhance the expression of complex semantics.
[0045] Step 3. Image feature extraction and multi-level feature representation: Use the images in the training set as input images and input them into the CLIP visual encoder. Process the features of the images layer by layer through a multi-level transformer structure to capture the visual information of the images at different scales and semantic levels, and extract features at different levels from the CLIP visual encoder.
[0046] Specifically, the images in the training set are used as input images and fed into the CLIP visual encoder. The visual encoder gradually extracts the features of the images through multiple layers of transformer models. Among these layers, the low-level layers mainly focus on details such as edges and textures; while the high-level layers focus on more abstract semantic information such as object categories and scene contents. In this way, the model can comprehensively understand the input images at multiple scales and semantic levels, providing rich visual representations for subsequent feature fusion and cross-modal interaction.
[0047] Step 4. Multi-level feature fusion: Select the initial layer features, intermediate layer features, and final layer features from the CLIP visual encoder, and calculate the weights of different level features, namely the weighting coefficients, through a multi-layer visual feature fusion module, and perform weighted combination on the features of different levels to construct visual fusion features.
[0048] The multi-layer visual feature fusion module adopts a four-layer network architecture, including two fully connected layers, one Relu activation layer, and one Softmax normalization layer. It uses a multi-head attention module to perform weighted combination on features of different levels to ensure that the fused features can retain both low-level detail information and high-level abstract semantics.
[0049] Step 4.1. Based on the features of different levels extracted from the CLIP visual encoder in Step 3, define the initial layer feature F1, intermediate layer feature F2, and final layer feature F3 output by the CLIP visual encoder as follows:
[0050]
[0051] Among them, respectively represent the height, width, and number of channels of the feature map of the initial layer feature F1; respectively represent the height, width, and number of channels of the feature map of the intermediate layer feature F2; respectively represent the height, width, and number of channels of the feature map of the final layer feature F3.
[0052] Step 4.2. Reshape each feature map to a unified spatial size H f ×W f :
[0053]
[0054] Among them, represents the reshaped initial layer feature, represents the reshaped intermediate layer feature, represents the reshaped final layer feature, H f represents the height of the reshaped feature map, W fLet \(W\) denote the width of the reshaped feature map, and \(C\) denote the number of channels of the reshaped feature map.
[0055] Step 4.3. Flatten and concatenate the reshaped features of each layer into a long vector:
[0056]
[0057] where \(concat(\cdot)\) represents concatenation and \(flatten(\cdot)\) represents flattening.
[0058] The dimension of the concatenated feature vector is:
[0059] where in the superscript represents the number of concatenated features in the feature vector \(F\) concat in the middle.
[0060] Step 4.4. To learn the weighted coefficients of each feature, first map the concatenated feature vector \(F\) concat to an intermediate dimension through a fully connected layer:
[0061] \(z = W\) concat \(\cdot F\) concat \(+ b\) concat ;
[0062] where \(z\) represents the mapped feature, \(W\) concat represents the weight matrix of the mapping process, and \(b\) concat represents the bias term.
[0063] Then obtain through the ReLU activation function:
[0064] \(z\) relu \(= ReLU(z)\);
[0065] where \(ReLU(\cdot)\) represents the activation function, and \(z\) relu represents the feature after being processed by the ReLU activation function.
[0066] Use another fully connected layer to map the activated feature to three weighted coefficients:
[0067] \(\alpha\) w \(= W\) weights \(\cdot z\) relu \(+ b\) weights ;
[0068] where \(\alpha\) w represents the weighted coefficient, \(w\in[1,3]\), \(W\) weights represents the weight matrix, and \(b\) weights represents the bias term;
[0069] Next, the Softmax function is used to normalize the weighting coefficients to ensure that the sum of the weighting coefficients is 1:
[0070]
[0071] Among them, α1 is the weighting coefficient of the initial layer features, α2 is the weighting coefficient of the intermediate layer features, and α3 is the weighting coefficient of the final layer features.
[0072] Finally, according to the learned weighting coefficients α1, α2, and α3, the features of each layer are weighted and summed to obtain the visual fusion feature P:
[0073]
[0074] Through this weighted fusion method, the combined zero-shot image classification model based on hierarchical feature fusion can adaptively learn how to combine features at different levels, thereby improving the classification performance of the model for downstream tasks.
[0075] Step 5. Cross-modal interaction to optimize text features: Input the text features obtained in Step 2 and the visual fusion features constructed in Step 4 into the cross-modal interaction module. Through the cross-attention mechanism, the deep fusion of the visual fusion features and the text features is realized to generate optimized text features that can better reflect visual information, improve the accurate description of image semantics, and obtain optimized text features of attributes, objects, and combinations.
[0076] Step 5.1. According to the text feature T obtained in Step 2 and the visual fusion feature P obtained in Step 4, use the cross-attention mechanism for interaction.
[0077] First, generate query vectors, key vectors, and value vectors.
[0078] Perform a linear mapping on the original text feature T to obtain the query vector q:
[0079] q = W T ·T;
[0080] Among them, W T represents the weight matrix for performing a linear mapping on the original text feature T.
[0081] Perform a linear mapping on the visual fusion feature P to obtain the key vector k v , k v which also serves as the value vector at the same time:
[0082] k v = W I ·P;
[0083] Among them, W I represents the weight matrix for performing a linear mapping on the visual fusion feature P;
[0084] Calculate the cross-attention between the text features and the visual fusion features Attention(q, k v ):
[0085]
[0086] where softmax(·) represents the normalized exponential function; d represents the dimension of the key, and usually to avoid numerical instability, the result is scaled.
[0087] After calculating the cross-attention, the original text features T and the output of the cross-attention are fused through a residual connection, and then the cross-modal optimized text features T are obtained through normalization and multi-layer perceptron mapping opt .
[0088] Step 5.2. After obtaining the cross-modal optimized text features T opt , through the method of weighted fusion, fuse T opt with the original text features T to obtain the fused text features T fused :
[0089] T fused = λ · T opt + (1 - λ) · Tq;
[0090] where λ represents the weight parameter, and λ is an adjustable weight parameter used for the relative importance of the cross-modal optimized text features T opt and the original text features T in the fusion;
[0091] The fused text features T fused are normalized by the L2 norm to obtain the final optimization result, that is, the optimized text features T final :
[0092]
[0093] The optimized text features T final include the optimized attribute text features, optimized object text features, and optimized combined text features, that is, as Figure 2 shown in the optimized attribute text vector, optimized object text vector, and optimized combined text vector.
[0094] Step 6. Feature decoupling and multi-modal representation: Decouple the image features output by the last layer of the CLIP visual encoder. The decoupling module uses three trainable multi-layer perceptrons MLP to extract attribute visual features, object visual features, and combined visual features from the image features output by the last layer of the CLIP visual encoder respectively, and obtain as Figure 2The visual vectors of attributes, object visual vectors, and combined visual vectors shown. These decoupled features ensure the independence of attributes, objects, and their combinations to better adapt to zero-shot combination tasks.
[0095] Specifically, based on the image features output by the last layer of the CLIP visual encoder obtained in step 3, i.e., the image embedding vector v ori , through three trainable multi-layer perceptrons MLP attr (·), MLP obj (·), MLP com (·), the visual features of attributes v attr , object visual features v obj , and combined visual features v com are respectively extracted as follows:
[0096] v attr = MLP attr (v ori );
[0097] v obj = MLP obj (v ori );
[0098] v com = MLP com (v ori ).
[0099] Step 7. The loss calculation module calculates the similarity between the visual features of attributes, objects, and combinations obtained by decoupling in step 6 and the optimized text features of attributes, objects, and combinations obtained in step 5 using cosine similarity respectively, obtains the predicted values, and measures the difference between the predicted values and the true labels through the classification loss function to form the losses of each branch. At the same time, a covariance loss term is introduced to optimize the decoupling effect, and the total loss is obtained through weighted summation.
[0100] Step 7.1. According to the attribute visual feature v attr , object visual feature v obj , and combined visual feature v com obtained in step 6, and the optimized text features of attributes, objects, and combinations obtained in step 5, calculate the cosine similarity sim(v, t) between the attribute visual feature, object visual feature, combined visual feature and the optimized text features of attributes, objects, and combinations respectively:
[0101]
[0102] Among them, v and t respectively represent the visual feature and the text feature of the same branch. The visual feature v includes the attribute visual feature, the object visual feature, and the combined visual feature. The optimized text feature t includes the optimized attribute text feature, the object text feature, and the combined text feature. · represents the dot product, and |v| and |t| are the norms (i.e., the modulus lengths) of the vectors v and t respectively.
[0103] After the cosine similarity calculation in step 7.2, the predicted values of the attribute branch, the object branch, and the combined branch are obtained through normalization by the softmax function.
[0104] To measure the difference between the predicted value and the true label, a classification loss function L is introduced. class It is expressed as:
[0105]
[0106] Among them, the x-th input image is the x-th sample. represents the predicted value of the x-th sample, and y x represents the true label of the x-th sample. Through the classification loss function L class The cross-entropy losses of each branch obtained will be used as part of the total loss.
[0107] In addition, in order to optimize the effect of the decoupling module and promote better separation of attributes and objects, the present invention introduces a covariance loss. Through the covariance loss, the model can reduce the similarity between the attribute and object visual features to strengthen their decoupling, so that the attributes and objects can learn more independent representations respectively.
[0108] For the visual feature v of the x-th input image x , using the decoupled attribute visual feature v attr,x and the object visual feature v obj,x of the x-th input image obtained in step 6, the covariance loss L attr,x between v obj,x and v cov is expressed as:
[0109]
[0110] Among them, s is the number of samples corresponding to the x-th input image, is the mean of the attribute visual features in the samples corresponding to the x-th input image, is the mean of the object visual features in the samples corresponding to the x-th input image. By minimizing the covariance loss, the network will be guided to learn to make the attribute visual feature and the object visual feature as irrelevant as possible, so as to achieve a better decoupling effect.
[0111] Step 7.3. Calculate the losses of the attribute branch, object branch, and combination branch through the classification loss function. At the same time, introduce the covariance loss between attributes and objects, and obtain the total loss through weighted summation.
[0112] The total loss L total is expressed as:
[0113] L total = αL comp + βL attr + γL obj + δL cov ;
[0114] where L attr 、L obj and L comp are the losses of the attribute branch, object branch, and combination branch respectively, and L cov is the covariance loss between attributes and objects; α, β, γ, and δ are the weight coefficients corresponding to each loss term, respectively, used to balance the influence of different loss terms.
[0115] Step 8. Model backpropagation and parameter update: Through the backpropagation mechanism of the total loss, jointly update the weighted coefficients of the multi-layer visual feature fusion module, the parameters of the MLP used for decoupling, and the initial embedding representations of attributes, objects, and combinations, so as to gradually improve the representation and generalization ability of the model.
[0116] Specifically, based on the total loss obtained in Step 7, through the backpropagation mechanism, the trainable parameters of each module, including the weighted coefficients of the multi-layer visual feature fusion module, the parameters in the MLP used for decoupling, and the initial embedding representations of attributes, objects, and combinations, will be jointly updated according to the loss function to improve the representation and generalization ability of the combined zero-shot image classification model based on hierarchical feature fusion.
[0117] Backpropagation starts from the total loss. First, calculate the gradients of the output layer and propagate these gradients layer by layer forward to the weight and bias parameters of each layer. For each layer, the gradients will be used to update the weight parameters of that layer with the aim of minimizing the loss function.
[0118] Specifically, for the multi-layer visual feature fusion module, the loss will be reflected in the gradient calculation through the weighted coefficients, thereby adjusting the weighted coefficients to make the fusion of visual features more reasonable and thus improving the feature expression ability of the model.
[0119] For the MLP network in the decoupling module, by calculating the influence of the covariance loss on the MLP weights, backpropagation will adjust the parameters in the MLP to further reduce the correlation between attribute and object features, thus better achieving the decoupling of attributes and objects.
[0120] Most importantly, the initial embedded representations of attributes, objects, and combinations are also affected by the loss and are updated in reverse. Through these updates, the model gradually fine-tunes and optimizes the initial embedded representations of attributes, objects, and combinations, namely word embeddings, to better meet the requirements of the downstream compositional zero-shot image classification task.
[0121] Step 9. Iteratively optimize the embedded representation: By repeatedly executing Steps 3 to 8, the embedded representations and parameters of attributes, objects, and combinations in the compositional zero-shot image classification model based on hierarchical feature fusion are optimized in each round of training; in each iteration, the model performs forward propagation based on the training data, makes predictions, and calculates the loss, and updates the parameters of each module through the backpropagation of the loss, including the weighting coefficients of the multi-layer visual feature fusion module, the MLP weights for decoupling, and the initial embedded representations of attributes, objects, and combinations; compared with the original model, the fine-tuned model can have stronger generalization ability and higher classification accuracy in the compositional zero-shot image classification task.
[0122] Specifically, by continuously executing Steps 3 to 8 iteratively, the initial embedded representations of attributes, objects, and combinations and related parameters are continuously optimized in each round of training; in each iteration, the model performs forward propagation based on the training data, makes predictions, then calculates the total loss and updates the parameters of each module through backpropagation, including the weighting coefficients of the multi-layer visual feature fusion module, the MLP weights for decoupling, and the initial embedded representations of attributes, objects, and combinations, etc.
[0123] After each iteration, the compositional zero-shot image classification model based on hierarchical feature fusion is evaluated on the validation set to monitor its performance on unseen classes; as the loss function gradually converges, the performance of the model on the validation set continues to improve, and the initial embedded representations of attributes, objects, and combinations are continuously updated as parameters in each backpropagation, finally obtaining the optimized embedded representations of attributes, objects, and combinations, as well as the trained compositional zero-shot image classification model based on hierarchical feature fusion, namely the optimal model. Compared with the original model, the fine-tuned model has stronger generalization ability and higher classification accuracy in the compositional zero-shot image classification task.
[0124] Step 10. Use the trained compositional zero-shot image classification model based on hierarchical feature fusion to classify the input image.
[0125] In addition, to verify the effectiveness of the method proposed in the present invention, the combined zero-shot image classification model based on hierarchical feature fusion proposed in the present invention was also used for image classification on the ut-zappos dataset. After experimental verification, the visible combination accuracy, unknown combination accuracy, best harmonic mean, and AUC value of the image classification method of the present invention reached 69.99%, 76.41%, 55.98%, and 0.4511 respectively, far higher than the best levels of 66.8%, 73.8%, 54.6%, and 0.4177 of all open-source models at home and abroad. The image classification method of the present invention exhibits excellent generalization ability and recognition effect, and can significantly improve the classification performance of the model for unknown complex combined images.
[0126] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.
Claims
1. A combined zero-shot image classification method based on hierarchical feature fusion, characterized in that: The steps include: Step 1. Build a combined zero-shot image classification model based on hierarchical feature fusion, including CLIP text encoder, CLIP visual encoder, multi-layer visual feature fusion module, cross-modal interaction module, decoupling module and loss calculation module; Step 2. Generate preliminary embedding representations, i.e., word embeddings, for attributes, objects, and combinations. The preliminary embedding representations will be updated as parameters during the back-propagation process. Combine the preliminary embedding representations of attributes, objects, and combinations with the prompt prefix embeddings to obtain the text features of attributes, objects, and combinations through the pre-trained CLIP text encoder. Step 3. Pass the image in the training set as the input image into the CLIP visual encoder, process the features of the image layer by layer, and extract features of different levels from the CLIP visual encoder; Step 4. Filter out the initial layer features, intermediate layer features and final layer features from the CLIP visual encoder, and calculate the weights of features at different levels, i.e. weighted coefficients, through the multi-layer visual feature fusion module, perform weighted combination of features at different levels, and construct visual fusion features; Step 5. Input the text features obtained in step 2 and the visual fusion features constructed in step 4 into the cross-modal interaction module, and realize the deep fusion of visual fusion features and text features through the cross-attention mechanism to obtain the optimized attribute, object, and combined text features; Step 6. Decouple the image features output by the last layer of the CLIP visual encoder. The decoupling module uses three trainable multi-layer perceptrons (MLPs) to extract attribute visual features, object visual features, and combined visual features from the image features output by the last layer of the CLIP visual encoder. Step 7. The loss calculation module calculates the similarity of the visual features of the attributes, objects, and combinations obtained after decoupling in step 6 and the text features of the optimized attributes, objects, and combinations obtained in step 5 using cosine similarity to obtain the predicted value, and measures the difference between the predicted value and the true label through the classification loss function to form the loss of each branch. At the same time, the covariance loss term is introduced to optimize the decoupling effect, and the total loss is obtained by weighted summation; Step 8. Through the back-propagation mechanism of the total loss, the weight coefficients of the multi-layer visual feature fusion module, the parameters of the MLP for decoupling, and the preliminary embedding representations of attributes, objects, and combinations are jointly updated; Step 9. Repeat steps 3 to 8, and evaluate the combined zero-shot image classification model based on hierarchical feature fusion after each iteration on the validation set until the loss function converges, and obtain the trained combined zero-shot image classification model based on hierarchical feature fusion and the optimized embedding representation of attributes, objects, and combinations according to the validation set results; Step 10. Use the trained combined zero-shot image classification model based on hierarchical feature fusion to classify the input image.
2. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.
1. For the attribute set A, object set O, and combination set C, first obtain the preliminary embedding representation of each attribute, object, and combination through the word embedding function φ(·) of CLIP; the preliminary embedding representation of the attribute, object, and combination will be used as a parameter to update in the back-propagation process; Where A={a1,a2,…,a m }, a1 to a m They represent the first to the mth attributes in the attribute set A, respectively, where m is the total number of attributes in the training set; O = {o1, o2, …, o n }, o1 to o n represent the 1st to nth objects in the object set O, respectively, and n represents the total number of objects in the training set; C = {c ij =(a i ,o j )|i∈[1,m],j∈[1,n]},a i represents the i-th attribute in the attribute set A, o j represents the jth object in the object set O, c ij Represents attribute a i and object o j combination of; The word embedding vectors of each attribute, object and their combination are represented as: in, and Word embeddings representing attributes, objects, and their combinations, i.e., preliminary embedding representations of attributes, objects, and combinations; Step 2.
2. To enhance the semantic expression ability of embedding, introduce the hint prefix mechanism; Each word in the prompt prefix is processed using the word embedding function to obtain the set P of prompt prefix embeddings, where P = {p1, p2, ..., p k }, p1 to p l They represent the embeddings of the 1st to the kth context words in the set P, respectively, and k represents the total number of context words in the prompt prefix; The word embedding of the prompt prefix, i.e., the prompt prefix embedding, is concatenated with the word embedding of each attribute, object, and their combination to construct the complete input text embedding with the prompt prefix, which is expressed as: in, Word embedding and attribute a representing the prompt prefix i Word embedding The complete input text constructed after concatenation, Word embedding representing the prompt prefix and object o j Word embedding The complete input text constructed after concatenation, Word embedding and attribute a representing the prompt prefix i and object o j Combination of word embeddings The complete input text constructed after splicing; Step 2.
3. Using CLIP text encoder E text (·) Encode the complete input text embedding with the prompt prefix to obtain the text feature T, which includes the attribute text feature Object Text Features and combined text features Respectively expressed as: in, and are the text features of attributes, objects and their combination in the encoder output space, respectively.
3. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: The step 3 is specifically as follows: Pass the images in the training set as input images into the CLIP visual encoder; The CLIP visual encoder gradually extracts features of the image through multiple levels of transformer models, where low-level layers focus on details and high-level layers focus on more abstract semantic information.
4. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: In step 4, the multi-layer visual feature fusion module adopts a four-layer network architecture; The multi-layer visual feature fusion module consists of two fully connected layers, a Relu activation layer and a Softmax normalization layer.
5. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: The step 4 is specifically as follows: Step 4.
1. Define the initial layer feature F1, intermediate layer feature F2 and final layer feature F3 output by the CLIP visual encoder as follows: in, They represent the feature map height, width, and number of channels of the initial layer feature F1 respectively; Respectively represent the feature map height, width, and number of channels of the intermediate layer feature F2; They represent the feature map height, width, and number of channels of the final layer feature F3 respectively; Step 4.
2. Adjust the feature maps of F1, F2 and F3 to a uniform spatial size H by upsampling or pooling f ×W f : in, represents the initial layer features after reshaping, represents the reshaped intermediate layer features, represents the final layer feature after reshaping, H f Represents the height of the reshaped feature map, W f represents the width of the reshaped feature map, and C represents the number of channels of the reshaped feature map; Step 4.
3. Reshape the features of each layer and Flatten and concatenate into a long vector F concat : Among them, concat(·) means concatenation, and flatten(·) means flattening. The concatenated feature vector F concat The dimensions are: in, Superscript Denotes the eigenvector F concat The number of concatenated features in ; Step 4.
4. To learn the weight coefficients of features at different levels, first pass the concatenated feature vector F through a fully connected layer. concat Map to an intermediate dimension: z=W concat ·F concat +b concat ; Among them, z represents the mapped feature, W concat Represents the weight matrix of the mapping process, b concat represents the bias term; Then through the ReLU activation function we get: With relu =ReLU(z); Among them, ReLU(·) represents the activation function, z relu Represents the features after being processed by the ReLU activation function; Use another fully connected layer to transform the activated feature z relu Mapped to three weighting coefficients: α w =W weights ·z relu +b weights ; Among them, α w represents the weighting coefficient, w∈[1,3], W weights represents the weight matrix, b weights represents the bias term; Then the weighted coefficients are normalized by the Softmax function to ensure that the sum of the weighted coefficients is 1: Among them, α1 is the weight coefficient of the initial layer feature, α2 is the weight coefficient of the intermediate layer feature, and α3 is the weight coefficient of the final layer feature; Finally, the features of each layer are weighted and summed according to the learned weight coefficients α1, α2, and α3 to obtain the visual fusion feature P:
6. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: The step 5 is specifically as follows: Step 5.
1. Based on the text feature T obtained in step 2 and the visual fusion feature P obtained in step 4, a cross attention mechanism is used for interaction; First, generate the query vector, key vector, and value vector; Perform linear mapping on the original text feature T to obtain the query vector q: q=W T ·T; Among them, W T Represents the weight matrix for linear mapping of the original text feature T; Linearly map the visual fusion feature P to obtain the key vector k v , k v Also as a value vector: k v =W I ·P; Among them, W I Represents the weight matrix for linear mapping of visual fusion feature P; Calculate the cross attention Attention(q,k between text features and visual fusion features v ): Here, softmax(·) represents the normalized exponential function; d represents the dimension of the key, which is used to avoid numerical instability and scale the results; After the cross-attention calculation, the original text feature T and the output of the cross-attention are fused through the residual connection, and then the cross-modal optimized text feature T is obtained through normalization and multi-layer perceptron mapping. opt ; Step 5.
2. After obtaining the cross-modal optimized text feature T opt Then, through weighted fusion, T opt Fuse with the original text feature T to obtain the fused text feature T fused : T fused =λ·T opt +(1-λ)·Tq; Among them, λ represents the weight parameter; The fused text feature T fused Normalized by L2 norm, we get the optimized text feature T final :
7. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: The step 6 is specifically as follows: The image feature obtained in step 3, output by the last layer of the CLIP visual encoder, is the image embedding vector v ori Based on three trainable multi-layer perceptrons MLP attr (·), MLP obj (·), MLP com (·) Extract attribute visual features v attr , object visual features v obj and the combined visual feature v com : v attr =MLP attr (v ori ); v obj =MLP obj (v ori ); v com =MLP com (v ori )。 8. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: The step 7 is specifically as follows: Step 7.
1. Attribute visual features v obtained from step 6 attr , object visual features v obj and the combined visual feature v com , and the optimized attribute, object, and combined text features obtained in step 5, respectively calculate the cosine similarity sim(v,t) between the attribute visual features, object visual features, combined visual features and the optimized attribute, object, and combined text features: Wherein, v and t represent the visual features and text features of the same branch respectively. The visual features v include attribute visual features, object visual features and combined visual features. The text features t include optimized attribute text features, object text features and combined text features. · represents the dot product. |v| and |t| are the norms of vectors v and t, i.e., the modulus lengths. Step 7.
2. After the cosine similarity calculation is completed, the prediction values of the attribute branch, object branch and combination branch are obtained by normalization through the softmax function; In order to measure the difference between the predicted value and the true label, the classification loss function L is introduced class , expressed as: Among them, the xth input image is the xth sample, represents the predicted value of the xth sample, y x Represents the true label of the xth sample; through the classification loss function L class The cross entropy loss of each branch obtained will be used as part of the total loss; Introducing covariance loss, for the visual feature v of the xth input image x , using the decoupled attribute visual feature v of the x-th input image obtained in step 6 attr,x and the object visual feature v obj,x , we get v attr,x and v obj,x The covariance loss L between cov It is expressed as: Among them, s is the number of samples of the category corresponding to the x-th input image, is the mean value of the attribute visual features in the samples of the category corresponding to the x-th input image, is the mean value of the visual features of the objects in the samples of the category corresponding to the x-th input image; Step 7.
3. Calculate the losses of the attribute branch, object branch, and combination branch through the classification loss function, introduce the attribute-object covariance loss, and obtain the total loss through weighted summation; Total loss L total It is expressed as: L total =αL comp +βL attr +γL obj +δL cov ; Among them, L attr , L obj and L comp denote the losses of the attribute branch, object branch, and combination branch respectively, and L cov is the covariance loss of the attribute-object; α, β, γ, and δ correspond to the weight coefficients of each loss term, which are used to balance the impact of different loss terms.
9. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: The step 8 is specifically as follows: Based on the total loss obtained in step 7, through the back-propagation mechanism, the trainable parameters of each module, including the weight coefficients of the multi-layer visual feature fusion module, the parameters in the MLP used for decoupling, and the preliminary embedding representations of attributes, objects, and combinations, will be jointly updated according to the loss function to improve the representation and generalization capabilities of the combined zero-shot image classification model based on hierarchical feature fusion; Back propagation starts from the total loss, first calculating the gradient of the output layer, and propagating the gradient forward layer by layer to the weights and bias parameters of each layer; for each layer, the gradient is used to update the weight parameters of the layer to reduce the loss function; For the multi-layer visual feature fusion module, the loss will be reflected in the gradient calculation through the weighted coefficient, so the weighted coefficient is adjusted to improve the feature expression ability of the model; For the MLP in the decoupling module, by calculating the impact of the covariance loss on the MLP weights, back-propagation adjusts the parameters in the MLP so that the correlation between attributes and object features is further reduced to achieve the decoupling of attributes and objects; The preliminary embedding representations of attributes, objects, and combinations are also affected by the loss and are gradually updated during the back-propagation process to meet the requirements of the downstream combinatorial zero-shot image classification task.
10. The combined zero-shot image classification method based on hierarchical feature fusion according to claim 1, characterized in that: The step 9 is specifically as follows: By repeatedly iterating steps 3 to 8, the embedded representations and parameters of attributes, objects, and combinations in the combined zero-shot image classification model based on hierarchical feature fusion are optimized in each round of training; In each iteration, the model performs forward propagation based on the training data, makes predictions and calculates the loss. The parameters of each module are updated through the back propagation of the loss, including the weight coefficients of the multi-layer visual feature fusion module, the MLP weights for decoupling, and the preliminary embedding representations of attributes, objects, and combinations. After each iteration, the combined zero-shot image classification model based on hierarchical feature fusion is evaluated on the validation set to monitor its performance on unseen categories; As the loss function gradually converges, the performance of the model on the validation set continues to improve. When the loss of the model on the validation set converges, we will eventually obtain the optimized embedding representation of attributes, objects, and combinations and the trained combined zero-shot image classification model based on hierarchical feature fusion.
Citation Information
Patent Citations
Visual attribute representation learning method for combined zero-order learning
CN118196428A
Combined zero sample image classification method based on progressive mutual guidance
CN118379562A