Mask perception-based efficient open vocabulary image recognition method and system, and readable storage medium

Through weight pruning and mask perception strategies, the pre-trained model is optimized, and the calculation overhead and false positive rate of the CLIP model in open vocabulary segmentation is solved, thereby achieving efficient and accurate open vocabulary recognition.

CN120563997APending Publication Date: 2025-08-29ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510624855.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing pre-trained visual language model has problems such as large computational overhead and high false positive rate in open vocabulary segmentation tasks, especially the insensitive CLIP to mask proposals lead to high false positive rate, and the traditional model compression method relies on heuristic strategies to lead to poor migration.

Method used

The sparse image encoder is obtained through weight pruning, mask perception strategies and self-distillation losses are introduced, pre-trained weight spectrum is analyzed, the training of sufficient layers is frozen, and only the insufficient layers are updated. Combined with the feature fusion of sparse image encoder and SAM image encoder, it reduces calculation costs and improves classification accuracy.

Benefits of technology

Effectively reduce mask classification false positives, reduce calculation costs, achieve a balance between model size and performance, and improve the accuracy and efficiency of open vocabulary segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563997A_ABST
    Figure CN120563997A_ABST
Patent Text Reader

Abstract

The invention discloses a mask perception-based efficient open vocabulary image recognition method and system and a readable storage medium, and the method comprises the steps: carrying out the pruning of a pre-training model, and obtaining a backbone network of a sparse image encoder; introducing a mask perception strategy, and adding a mask proposal as an attention bias to a multi-head attention module of the backbone network; the weight quality of the sparse image encoder is evaluated, a layer with insufficient training is determined by analyzing a heavy tail behavior in a weight spectrum, only the layer with insufficient training is updated, and other layers are kept frozen; inputting the image into a sparse image encoder and an SAM image encoder to obtain two image features, and fusing the two image features to obtain a fused image feature; performing feature representation on the to-be-identified category name by using a text encoder to obtain text features; calculating the cosine similarity of the text features and the fused image features to obtain classification prediction; and obtaining a final image recognition result in combination with the mask proposal. According to the invention, misinformation of mask classification can be reduced, and the calculation cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and in particular relates to a mask-aware efficient open vocabulary image recognition method, system and readable storage medium. Background Art

[0002] Recently, pre-trained vision-language models have been widely used to solve the challenging zero-shot segmentation task. Typical solutions follow the principle of first generating mask proposals and then classifying them using CLIP. In order to maintain the zero-shot transfer capability of CLIP, past studies often chose to freeze the CLIP model during training.

[0003] For example, Chinese patent document CN118710903A discloses an open-vocabulary zero-shot semantic segmentation method for traffic roads. First, a frozen CLIP backbone network is used to generate a category-independent mask, and then a low-resolution image and the predicted mask are input into the backbone network for open-vocabulary recognition.

[0004] However, CLIP shows insensitivity to different mask proposals and tends to give similar predictions for multiple mask proposals of the same image. This insensitivity leads to a large number of false positives when classifying mask proposals. This problem mainly stems from the fact that CLIP is trained with image-level supervision.

[0005] The recent success of pre-trained baseline vision-language models has made Open Vocabulary Segmentation (OVS) possible. Despite its encouraging performance, this approach faces two major challenges that result in significant computational overhead: 1) the large size of the backbone model; and 2) the high cost of fine-tuning. These issues limit the widespread applicability and cost-effectiveness of the OVS strategy in real-world applications.

[0006] While traditional model compression and efficient fine-tuning methods can address these challenges to some extent, they often rely on heuristic strategies. This means these solutions have poor transferability and require retraining for different models, which increases the cost burden. Therefore, a method that can reduce false positives in mask classification with low computational cost is urgently needed, achieving an excellent trade-off between segmentation accuracy and computational cost. Summary of the Invention

[0007] The present invention provides a mask-aware, efficient open-vocabulary image recognition method, system, and readable storage medium, which can effectively reduce mask classification false positives, shrink the model through weight pruning, and train the selected layer with pre-trained weight analysis, effectively reducing fine-tuning costs, and have open-vocabulary generalization capabilities.

[0008] A mask-aware, efficient open vocabulary image recognition method comprising the following steps:

[0009] (1) Prune the pre-trained model to obtain the backbone network of the sparse image encoder;

[0010] (2) Introducing a mask-aware strategy, mask proposals are added as attention biases to the multi-head attention module of the backbone network;

[0011] (3) Evaluate the weight quality of the sparse image encoder and identify under-trained layers by analyzing the heavy-tail behavior in the weight spectrum. Only update the under-trained layers and keep the other layers frozen.

[0012] (4) Input the image into the sparse image encoder and the SAM image encoder respectively, obtain the two image features, and then perform feature fusion to obtain the fused image features;

[0013] (5) Use the text encoder to represent the category name to be identified and obtain text features; calculate the cosine similarity between the text features and the fused image features to obtain the classification prediction; combine the mask proposal to obtain the final image recognition result.

[0014] In step (1), the pre-trained model is CLIP, and iterative amplitude pruning is performed on the pre-trained model to delete some weights with the global minimum amplitude to obtain the backbone network of the sparse image encoder.

[0015] In step (2), a mask-aware strategy is introduced, using mask-aware loss and self-distillation loss to make the backbone network of the sparse image encoder sensitive to the masked region while avoiding destroying the pre-training knowledge.

[0016] The mask-aware loss prompts the sparse image encoder to distinguish various mask proposals by minimizing the distance between the IoU score of the mask proposal and the classification score of the encoder in the pre-trained model.

[0017] In step (2), the attention bias is expressed as follows:

[0018]

[0019] Where B is the attention bias, M is the mask proposal, is the intermediate calculation result, N is the number of mask proposals, I(N,N) represents the N-th order identity matrix, cat(·) represents row-wise concatenation, and F(·) represents the flattening operation.

[0020] In step (4), after the two image features are channel-adjusted and spatially aligned, feature fusion is performed by element-by-element addition to enhance the performance of open vocabulary semantic segmentation.

[0021] A mask-aware, efficient open-vocabulary image recognition system includes a memory and one or more processors. The memory stores executable code, and the one or more processors implement the efficient open-vocabulary image recognition method when executing the executable code.

[0022] A computer-readable storage medium stores a program, which, when executed by a processor, implements the above-mentioned efficient open vocabulary image recognition method.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1. The present invention effectively utilizes contextual information in the image encoder to reduce computational costs; it obtains a sparse subnetwork through iterative amplitude pruning to achieve a balance between model size and performance.

[0025] 2. This paper introduces a mask-aware mechanism and loss to solve the problem that CLIP is insensitive to different mask proposals and tends to produce similar predictions for various mask proposals of the same image. This insensitivity leads to a large number of false positives when classifying mask proposals.

[0026] 3. This paper studies the spectral characteristics of pre-trained weights, freezing layers with heavy-tailed distributions and updating only layers with light-tailed distributions. This analysis method not only further accelerates training by reducing computational effort but also achieves zero overhead in pre-trained models. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a flow chart of a mask-aware, efficient open vocabulary image recognition method according to an embodiment of the present invention.

[0028] Figure 2 This is a schematic diagram of the present invention applied to low-altitude target recognition. DETAILED DESCRIPTION

[0029] The present invention will be described in further detail below with reference to the accompanying drawings and examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.

[0030] The overall concept of this invention is as follows: First, it uses iterative amplitude pruning to prune the pre-trained model to obtain a sparse sub-network to process the input image. Second, considering that CLIP is insensitive to different mask proposals and tends to produce similar predictions for different mask proposals of the same image, leading to false positives, a mask-aware mechanism is introduced to ensure that CLIP responds to different mask proposals. Finally, by fusing frozen SAM image encoder features, the CLIP image encoder is supplemented with important spatial information for semantic segmentation. Finally, the quality of the pre-trained model is analyzed and layers are selectively frozen to effectively accelerate model training.

[0031] like Figure 1 As shown, a mask-aware efficient open vocabulary image recognition method includes:

[0032] S1. Obtain the backbone network of the sparse image encoder through pruning, including:

[0033] S1-1. Train an unpruned network as a pre-training model.

[0034] S1-2. Delete some weights with the global minimum magnitude, and in the pruning process, only use knowledge distillation loss and exclude any semantic supervision.

[0035] S2. Apply mask proposals as attention masks in multi-head attention and specify independent classification queries for each mask proposal to improve the sparse image encoder design.

[0036] S3. Analyze the weight spectrum of the sparse image encoder, train the layers with poor training quality, and freeze the layers with good training quality.

[0037] S4. Input the image to the sparse image encoder and SAM image encoder for processing to obtain two image features.

[0038] S4-1. Image I∈R H×W×3 Input to the SAM image encoder and obtain image features from the last three global attention blocks

[0039] S4-2. Downsample I to (where p is the downsampling rate). * Input to the sparse image encoder, image features are obtained from the encoder's L / 4, L / 2 and 3L / 4 blocks (the total number of blocks in the L image encoder)

[0040] S5. A simple addition method is used to fuse the two image encoder features.

[0041] S6. Calculate the cosine similarity of the fused image features and text features to obtain the category prediction result, and combine it with the mask proposal to obtain the final image recognition result.

[0042] Step S1 primarily improves model efficiency through model pruning. Specifically, it finds effective subnetworks that can be seamlessly transferred to different OVS frameworks. To determine the backbone pruning mask, a classic iterative amplitude pruning method is used. Pruning the model backbone is performed in two steps. First, an unpruned network is trained as a pretrained model. Then, the weights with the global minimum amplitude are removed.

[0043] Knowledge Distillation Loss It plays a key role in aligning the feature space between CLIP text and visual features, significantly affecting the open vocabulary performance of the model. Regarding step (2) to find the pruning mask during the pruning process, only the knowledge distillation loss is used and any semantic supervision is excluded. This approach enables the discovery of pruning masks on an OVS framework that can be transferred to other frameworks without any further customization, while still retaining the open domain capabilities of CLIP without overfitting to specific semantic classes.

[0044] Furthermore, step S2 uses attention masks in multi-head attention and provides flexibility for accepting different numbers of queries and features of different mask regions. Therefore, mask proposals are applied as attention masks in multi-head attention, and an independent classification query is specified for each mask proposal.

[0045] In step S2, s Mask-aware loss function is performed on Among them C s The goal is to assign high scores to high-quality proposals and low scores to C s Use the ground-truth to get the IoU score and compare it with C s Alignment, prompting CLIP to become mask-aware. Assuming there are k classes in the ground-truth, k binary maps of the ground-truth can be generated and the IoU score (S IoU ). C s The maximum value of S tends to 1, and IoU The maximum value of S is in the range of 0.75 to 0.99. This inconsistency may lead to the misalignment between the two indicators. IoU A min-max normalization technique is introduced as follows:

[0046]

[0047] At the same time, the method selects k pre-existing classes in C s and use SmoothL1Loss to combine it with Align. Therefore, It can be expressed as follows:

[0048]

[0049] Furthermore, in step S3, gradients are selectively calculated and updated for untrained layers (weights of poor quality), while trained layers (weights of good quality) are frozen and skipped, effectively reducing the computational cost during fine-tuning. The quality of pre-trained weights is evaluated by analyzing the heavy tail behavior in the weight spectrum, and under-trained and fully-trained layers are identified.

[0050] In step S4, a 12-layer Transformer is used to process the image, and the features propagated between layers are represented as F i , where i = [1,2...12]. F i Indicated as F i =[F i cls ; F i feat ],∈R (1+hw)×d ,1 represents a class embedding vector F i cls , hw represents the flat image feature In order to obtain the classification of all mask proposals at the same time, we repeat F N times in the L layer. i cls , where N is the number of mask proposals, the repeated class embedding vector is represented as The modified feature is expressed as

[0051] For step S4, F i propagation, i = [1,2…L]. CLIP’s classification relies heavily on contextual information. In the first L Transformer layers, F i The propagation of is the same as in standard CLIP. Specifically, use The cross attention of all pixels in F effectively preserves the context information. In the subsequent 12-L Transformer layer, F i* The spread of can be divided into two parts: The spread and spread.

[0052] Use of dissemination and M[p] to represent and position p in M, where p = [1,2…P] and M is the mask proposal. Calculate the multi-head attention for the position of M[p] = 1 itself. To this end, an attention bias B∈R is constructed P×(P+hw) ,as follows:

[0053]

[0054] Among them, I(N,N) represents the N-th order identity matrix and F(·) represents the flattening operation. Therefore, using multi-head attention propagation

[0055]

[0056] Que(·), Key(·), and Val(·) represent linear projections, and d is the i The hidden dimension of .

[0057] The propagation uses the standard multi-head attention mechanism:

[0058]

[0059] Therefore, for any given mask proposal M[n], the corresponding class embedding Use only Perform multi-head attention, where M[n] = 1 and The propagation of is still not disturbed by the attention mask. Compared with the frozen CLIP, it effectively utilizes the context information and reduces the computational cost.

[0060] In step S5, feature extraction and fusion are required. The feature map of the sparse image encoder lacks important spatial information for semantic segmentation. The frozen SAM image encoder is used to supplement the spatial information. Specifically, given an image I∈R H ×W×3 , which is fed into the SAM image encoder and the image features are obtained from the last three global attention blocks At the same time, I is downsampled to (where p is the downsampling rate). * Input to the image encoder, image features are obtained from the encoder's L / 4, L / 2 and 3L / 4 blocks (the total number of blocks in the L image encoder)

[0061] For the feature fusion in step S5, a simple addition method is used to fuse the two image features. First, a linear layer is used to transform F b The number of channels and F a Then, F a and F b Upsampling or downsampling is Finally, F a and F b Add element by element to get the fusion feature F i

[0062]

[0063] Furthermore, in step S6, the fused image features are aligned with the text embeddings generated by the CLIP Text Encoder in a shared semantic space for cross-modal alignment. This involves calculating the cosine similarity matrix between the regional image features and all categorical text features. After temperature scaling and softmax normalization, the category probability distribution is obtained. Finally, the final open-word semantic segmentation result is obtained by combining the mask proposals.

[0064] As an example of the application of this invention to low-altitude target recognition, the dataset used for verification is based on a carefully selected, high-quality public target recognition dataset. Using EfficientSAM, the dataset automatically converts annotations into the format required for semantic segmentation. This dataset selects high-quality images from seven public datasets, encompassing a large number of various drones and a small number of other targets. For testing, 700 images were randomly sampled from this dataset.

[0065] The original image and the image segmented using the word "drones" are as follows Figure 2 As shown in the experimental results, the proposed method has a high pixel accuracy rate; the method correctly responds to different proposals for the same image and achieves a balance between efficiency and performance, which proves the superiority of the proposed method.

[0066] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A mask-aware, efficient open vocabulary image recognition method, characterized in that: The following steps are involved: (1) Prune the pre-trained model to obtain the backbone network of the sparse image encoder; (2) Introducing a mask-aware strategy, mask proposals are added as attention biases to the multi-head attention module of the backbone network; (3) Evaluate the weight quality of the sparse image encoder and identify under-trained layers by analyzing the heavy-tail behavior in the weight spectrum. Only update the under-trained layers and keep the other layers frozen. (4) Input the image into the sparse image encoder and the SAM image encoder respectively, obtain the two image features, and then perform feature fusion to obtain the fused image features; (5) Using a text encoder to represent the category name to be identified, obtain text features; calculate the cosine similarity between the text features and the fused image features to obtain classification predictions; Combine the mask proposals to obtain the final image recognition result.

2. The mask-aware efficient open vocabulary image recognition method according to claim 1, characterized in that In step (1), the pre-trained model is CLIP, and iterative amplitude pruning is performed on the pre-trained model to delete some weights with the global minimum amplitude to obtain the backbone network of the sparse image encoder.

3. The mask-aware efficient open vocabulary image recognition method according to claim 1, characterized in that: In step (2), a mask-aware strategy is introduced, using mask-aware loss and self-distillation loss to make the backbone network of the sparse image encoder sensitive to the masked region while avoiding destroying the pre-training knowledge.

4. The mask-aware efficient open vocabulary image recognition method according to claim 3, characterized in that: The mask-aware loss prompts the sparse image encoder to distinguish various mask proposals by minimizing the distance between the IoU scores of mask proposals and the classification scores of the encoder in the pre-trained model.

5. The mask-aware efficient open vocabulary image recognition method according to claim 1, characterized in that: In step (2), the attention bias is expressed as follows: Where B is the attention bias, (x, y) represents the row and column positions, and M is the mask proposal. is the intermediate calculation result, N is the number of mask proposals, I(N,N) represents the N-th order identity matrix, cat(·) represents row-wise concatenation, and F(·) represents the flattening operation.

6. The mask-aware efficient open vocabulary image recognition method according to claim 1, characterized in that: In step (4), after the two image features are channel-adjusted and spatially aligned, feature fusion is performed by element-by-element addition to enhance the performance of open vocabulary semantic segmentation.

7. A mask-aware, efficient open vocabulary image recognition system, characterized by: The method comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, the method is used to implement the efficient open vocabulary image recognition method according to any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that A program is stored thereon, and when the program is executed by a processor, the efficient open vocabulary image recognition method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Open vocabulary traffic road zero sample semantic segmentation method

    CN118710903A