Text-guided parameter efficient fine-tuning image segmentation and counting model and counting method

Through text-guided parameters, efficiently fine-tuning the image segmentation and counting model, combined with CLIP and SAM models, the problems of high computational costs and unused text features in the prior art are solved, and efficient and accurate target object recognition and segmentation are achieved.

CN120014396APending Publication Date: 2025-05-16TIANJIN NORMAL UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510103092.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing target counting method based on pretrained visual models requires a lot of training and is computationally cost-effective, and does not fully utilize text features and ignore the targets in the image.

Method used

A text-guided parameter-efficient fine-tuning image segmentation and counting model is designed, combining the pre-trained visual language big model CLIP and segmentation model SAM, and the optimal bounding box and mask features are generated through the fusion of text features and image features to achieve the recognition and segmentation of target objects.

Benefits of technology

Reduces computational costs, reduces the need for large amounts of annotated data and model training, maintains the powerful generalization performance of pre-trained models, and improves the accuracy of segmentation counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014396A_ABST
    Figure CN120014396A_ABST
Patent Text Reader

Abstract

The invention discloses a text-guided parameter efficient fine-tuning image segmentation and counting model and counting method, and the model comprises a pre-trained visual language large model CLIP, a maximum connected region and non-maximum suppression module, and a pre-trained segmentation model SAM. Wherein the pre-trained visual language large model CLIP comprises a pre-trained CLIP image encoder and a standard text encoder; the pre-trained segmentation model SAM comprises an SAM encoder, a prompt encoder and a mask decoder, the pre-trained segmentation model SAM further integrates a lightweight adapter and a CLIP feature fusion and mask generation module, the lightweight adapter is used for adjusting the SAM encoder, and the prompt encoder is used for prompting the SAM encoder. And the CLIP feature fusion and mask generation module is used for migrating and fusing the image feature FC generated by the CLIP image encoder into a mask decoder, and guiding the mask decoder to generate a high-quality segmentation mask. The model provided by the invention has strong generalization performance and relatively high counting accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and image processing, and in particular to a text-guided parameter efficient fine-tuning image segmentation and counting model and a counting method. Background Art

[0002] Accurately counting the number of objects in an image plays a vital role in today's scene understanding field. It is widely used in real-time traffic flow analysis in traffic monitoring, lesion identification in medical imaging to assist in disease diagnosis, and wildlife population statistics in ecosystem management. Traditional counting models usually rely on closed data sets and a large amount of manual annotation, which limits their application and promotion in practical scenarios.

[0003] With the continuous advancement of deep learning technology, the introduction of class-agnostic counting has brought innovative ideas and methods to this field, significantly improving the accuracy and generalization of target counting tasks. This paradigm uses a small number of examples and compares the similar features between samples and images, allowing the counting model to dynamically adjust during the counting process and improving the model's generalization ability for different categories. GMN is a representative pioneering work that extracts image and sample features for matching and regresses the results into a density map. FamNet obtains sample features and associates them with images through ROIpooling. BMNet proposes a learnable bilinear similarity metric to represent the similarity between samples and images. However, the class-agnostic counting paradigm ignores category text information, which weakens the model's ability to distinguish subtle differences between different targets.

[0004] The emergence of large multimodal pre-trained models provides new ideas and methods for improving the accuracy and robustness of target recognition and counting. CrowdCLIP builds sorting text prompts and multimodal sorting losses to explore how to use CLIP for unsupervised crowd counting. CLIPCount enhances the model's ability to locate targets in images by designing patch-text contrast loss and hierarchical patch-text interaction modules, and realizes the interaction of visual features and semantic information across resolutions. Teaching CLIP to Count automatically creates a counterfactual prompt to encourage the model to approach the real target and stay away from counterfactual counts. Although these text-guided counting methods reduce the need for manual annotation, they require a lot of training. SAMCount converts the counting task into a segmentation task. Although its performance is far behind the latest methods, the authors believe that this is still a visual task worthy of further exploration.

[0005] Although the existing object counting based on pre-trained visual models has stronger representation capabilities, it still has the following problems:

[0006] (1) Most models require a lot of training, which increases the computational cost; (2) They do not fully explore and utilize text features, thus ignoring the objects in the image. Summary of the invention

[0007] The purpose of the present invention is to provide a text-guided parameter efficient fine-tuning image segmentation and counting model to address the technical defects in the prior art.

[0008] Another object of the present invention is to provide a segmentation and counting method for the text-guided parameter-efficient fine-tuning image segmentation and counting model.

[0009] Another object of the present invention is to provide a method for fine-tuning the parameters of the text-guided image segmentation and counting model efficiently.

[0010] The technical solution adopted to achieve the purpose of the present invention is:

[0011] A text-guided parameter-efficient fine-tuning image segmentation and counting model, including a pre-trained large visual language model CLIP, a maximum connected region and non-maximum suppression module, and a pre-trained segmentation model SAM, where:

[0012] The pre-trained visual language model CLIP includes a pre-trained CLIP image encoder and a standard text encoder. The CLIP image encoder extracts image features F from the image. C , the standard text encoder extracts text features T from the name of the target category to be counted t , redundant features T r The pre-trained visual language model CLIP uses text features T t , redundant features T r and image features F C The class-related token cls_token in the class is used to calculate the class text-related embedding E t And cosine similarity Sim c ;

[0013] The maximum connected region and non-maximum suppression module are connected to the pre-trained visual language model CLIP to obtain the cosine similarity Sim c , generate the best bounding box (x1,x2,y1,y2) and pass it to the pre-trained segmentation model SAM as a hint;

[0014] The pre-trained segmentation model SAM includes a SAM encoder, a hint encoder and a mask decoder, and the SAM encoder is connected with the hint encoder to generate an image feature F sThe hint encoder is connected to the maximum connected area and non-maximum suppression module to receive the best bounding box (x1, x2, y1, y2) and form an initialization Token. The pre-trained segmentation model SAM also integrates a lightweight adapter and a CLIP feature fusion and mask generation module. The CLIP feature fusion and mask generation module is connected to the CLIP image encoder and the standard text encoder to convert the image feature F C and class text related embedding E t The light-weight adapter is used to adjust the SAM encoder. The prompt encoder, CLIP feature fusion and mask generation module, and SAM encoder are connected to the mask decoder respectively. The mask decoder combines the mask feature mask feat, the initialization Token, and the image feature F s , Class Text Related Embedding E t Generate a segmentation mask and get the required number of target objects.

[0015] In the above technical solution, the CLIP feature fusion and mask generation module includes a cross attention module, two 1×1 convolutional layers and a ReLU, a multi-head self-attention module, a residual connection and layer normalization, an MLP layer, a secondary residual connection and layer normalization.

[0016] In the above technical solution, the SAM encoder is composed of N layers, and each layer is composed of a first layer normalization, a multi-head attention layer, a second layer normalization and an MLP layer connected in sequence.

[0017] In the above technical solution, there are two lightweight adapters, each of which is composed of a Down linear layer, a GELU, and an Up linear layer connected in sequence. The first lightweight adapter is located after the multi-head self-attention layer of the SAM encoder and is connected to the residual of the previous step. The second lightweight adapter is located in the residual path of the MLP layer of the SAM encoder.

[0018] The image segmentation and counting method based on the text-guided parameter efficient fine-tuning of the image segmentation and counting model comprises the following steps:

[0019] Step 1: The pre-trained CLIP image encoder extracts image features F from the image C , the image feature F C Contains the class-related token cls_token, and the standard text encoder extracts the text features T of the category to be counted from the name of the target category to be counted t , redundant features Tr, embedding the class-related token cls_token into the text feature T tThen, with the text feature T t Perform secondary combination to generate the image-specific class text related embedding E t , class text-related embedding E t Subtract redundant features T r The information combined with the class-related token cls_token is finally combined with the CLIP image feature F C Calculate cosine similarity Sim c ;

[0020] The CLIP feature fusion and mask generation module fuses the image features F C and class text related embedding E t Get the mask feature mask feat;

[0021] Step 2: The maximum connected area and non-maximum suppression module obtain the Sim from step 1. c Get the best bounding box from (x1, x2, y1, y2);

[0022] Step 3: The pre-trained segmentation model SAM takes the optimal bounding box (x1, x2, y1, y2) as a prompt, combines the mask feature mask feat, the initialization Token, and the image feature F s , Class Text Related Embedding E t , identify and segment the target objects in the image, obtain the segmentation image mask, and then get the number of the required target objects.

[0023] In the above technical solution, Class Text Dependent Embedding E t =concat[cls_token⊙T t , T t ],in represents the cosine similarity measure, ⊙ represents the Hadamard product, concat is a function that concatenates multiple arrays, and cls_token is a class-related token.

[0024] In the above technical solution, in step 3, the optimal bounding box (x1, x2, y1, y2) is input to the prompt encoder, and Sim is obtained under the joint action of the prompt encoder and the SAM encoder. s , Sim s Then enter the prompt encoder to get the initialization Token;

[0025] The SAM encoder extracts image features F from the image to be recognized s ;

[0026] The CLIP feature fusion and mask generation module fuses the image features F Cand class text related embedding E t Get the mask feature mask feat;

[0027] Initialize Token and image features F s , Class Text Related Embedding E t And the mask feature maskfeat are input into the mask decoder for recognition and segmentation processing to obtain the number of target objects.

[0028] In the above technical solution, maskfeat = LayNorm (MLP (X3) + X3)

[0029] X3=LayNorm(E t +mh_attn(cross_attn(E t , F C , F C ), Conv(F C ), Conv(F C )))

[0030] Among them, MLP is a multi-layer perceptron, LayNorm is layer normalization, mh_attn is multi-head self-attention, cross_attn is cross attention, Conv means one 1×1 convolution, one ReLU and two 1×1 convolutions, and X3 is the intermediate value.

[0031] Another aspect of the present invention also includes a text-guided parameter-efficient fine-tuning method for fine-tuning an image segmentation and counting model, comprising the following steps:

[0032] Step 1: Build a dataset and adjust the size of each image in the dataset so that it can be used as the input of the CLIP image encoder and the SAM encoder.

[0033] Step 2: Freeze the pre-trained visual language model CLIP and segmentation model SAM;

[0034] Step 3, input the preprocessed image and the category name of the target to be counted into the text-guided parameter efficient fine-tuning image segmentation and counting model to obtain the segmentation image mask and the number of targets to be counted;

[0035] Step 4: Calculate the loss function and update the parameters of the lightweight adapter and CLIP feature fusion and mask generation module; repeat steps 3 and 4 above until the training is completed.

[0036] In the above technical solution, the calculation formula of the loss function is:

[0037] L=L2+L t_ce +L dice

[0038] Where L2 is the MSE loss, L t_ce is the loss for mask classification, L dice is the loss used for segmentation.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. The text-guided parameter-efficient fine-tuning image segmentation and counting method provided by the present invention designs a unified recognition and segmentation architecture of CLIP and SAM, converts the counting problem into a text-guided segmentation task, and eliminates the need for a large amount of labeled data and model training. The present invention extracts the best bounding box from CLIP according to the name of the target category to be counted, and no manual labeling of samples is required. Fully exploring and utilizing the knowledge of pre-trained models can reduce computational costs while maintaining their strong zero-sample reasoning capabilities.

[0041] 2. The text-guided parameter-efficient fine-tuning image segmentation and counting method provided by the present invention integrates a lightweight adapter in the SAM image encoder in a parameter-efficient fine-tuning manner, and adds a CLIP feature fusion and mask generation module in the segmentation model SAM. The former promotes feature extraction of the target quantity, and the latter improves the quality of the mask, while maintaining the powerful generalization performance of the pre-trained model with only a small number of parameters updated.

[0042] 3. The text-guided parameter-efficient fine-tuning image segmentation and counting method provided by the present invention provides three multi-source knowledge fusion strategies, including prompt priors, text-related embedding, and visual feature fusion of CLIP and SAM. These feature fusions can flexibly obtain context dependencies and improve the accuracy of segmentation and counting. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of the overall implementation method of an embodiment of the present invention;

[0044] Figure 2 A network architecture diagram provided for an embodiment of the present invention;

[0045] Figure 3 A schematic diagram of a lightweight adapter provided by an embodiment of the present invention;

[0046] Figure 4 A schematic diagram of a CLIP feature fusion and mask generation module provided in an embodiment of the present invention;

[0047] Figure 5 This is a comparison chart of the effects of predicted segmentation and real images of the text-guided parameter efficient fine-tuning image segmentation and counting method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The present invention is further described in detail below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] Example 1

[0050] like Figure 1-Figure 2 As shown, a text-guided parameter-efficient fine-tuning image segmentation and counting model includes a pre-trained visual language model CLIP, a maximum connected region and non-maximum suppression module, and a pre-trained segmentation model SAM, where:

[0051] The pre-trained visual language model CLIP includes a pre-trained CLIP image encoder and a standard text encoder. The CLIP image encoder extracts image features F from the image. C , the standard text encoder extracts text features T from the name of the target category to be counted t , redundant features T r The pre-trained visual language model CLIP uses text features T t , redundant features T r and image features F C The class-related token cls_token in the class is used to calculate the class text-related embedding E t And cosine similarity Sim c ;

[0052] The maximum connected region and non-maximum suppression module are connected to the pre-trained visual language model CLIP to obtain the cosine similarity Sim c , generate the best bounding box (x1, x2, y1, y2) and pass it to the pre-trained segmentation model SAM as a hint;

[0053] The pre-trained segmentation model SAM includes a SAM encoder, a hint encoder and a mask decoder, and the SAM encoder is connected with the hint encoder to generate an image feature F s The hint encoder is connected to the maximum connected area and non-maximum suppression module to receive the best bounding box (x1, x2, y1, y2) and form an initialization Token. The pre-trained segmentation model SAM also integrates a lightweight adapter and a CLIP feature fusion and mask generation module. The CLIP feature fusion and mask generation module is connected to the CLIP image encoder and the standard text encoder to convert the image feature F C and class text related embedding E tThe light-weight adapter is used to adjust the SAM encoder. The prompt encoder, CLIP feature fusion and mask generation module, and SAM encoder are connected to the mask decoder respectively. The mask decoder combines the mask feature mask feat, the initialization Token, and the image feature F s , Class Text Related Embedding E t Generate a segmentation mask and get the required number of target objects.

[0054] Preferably, the SAM encoder is composed of N layers, each of which is composed of a first layer normalization, a multi-head attention layer, a second layer normalization and an MLP layer connected in sequence.

[0055] Preferably, the CLIP feature fusion and mask generation module is as follows: Figure 4 As shown in the figure, it includes a cross attention module, two 1×1 convolutional layers and a ReLU, a multi-head self-attention module, a residual connection and layer normalization, an MLP layer, a secondary residual connection and layer normalization.

[0056] Preferably, two lightweight adapters are provided, each of which has a structure as follows: Figure 3 As shown in the figure, it consists of a Down linear layer, a GELU, and an Up linear layer connected in sequence.

[0057] The first Down linear layer compresses the given embedding to a lower dimension, the GELU layer makes the output nonlinear instead of linear, and the final Up linear layer expands the compressed embedding back to its original dimension. The first lightweight adapter is located after the multi-head self-attention layer of the SAM encoder and is connected to the residual of the previous step. The second lightweight adapter is placed in the residual path of the MLP layer of the SAM encoder. It allows for fast and efficient adjustments to the model without significantly increasing the computational burden.

[0058] Example 2

[0059] A text-guided parameter-efficient fine-tuning image segmentation and counting method, comprising the following steps:

[0060] Step 1: The pre-trained visual language model CLIP extracts image features F from the image C , extract the text features T of the category to be counted from the name of the target category to be counted t ;

[0061] The pre-trained visual language model CLIP includes a pre-trained CLIP image encoder (VisionTransformer image encoder) and a standard text encoder. The VisionTransformer image encoder is used to extract image features F from the image. C , image feature F C Contains the class-related token cls_token, the extracted image features F C ∈R 1×1025×d , where d = 512, indicating the dimension of the CLIP embedding space;

[0062] The standard text encoder extracts the text features T of the category to be counted based on the name of the target category to be counted t ∈R 1 ×d , empty character text feature is redundant feature T r ∈R 1×d ; Embed the class-related token cls_token into the text feature T t Then, with the text feature T t Perform secondary combination to generate the image-specific class text related embedding E t , class text-related embedding E t Subtract redundant features T r The information combined with the class-related token cls_token is finally combined with the CLIP image feature F C Calculate cosine similarity Sim c , the process is expressed as:

[0063]

[0064] where the class text related embedding E t =concat[cls_token⊙T t , T t ], represents the cosine similarity measure, ⊙ represents the Hadamard product, concat is a function that concatenates multiple arrays, and cls_token is a class-related token.

[0065] The CLIP feature fusion and mask generation module fuses the image features F C and class text related embedding E t Get the mask feature mask feat.

[0066] Step 2: The maximum connected area and non-maximum suppression module obtain the Sim from step 1. c Get the best bounding box from (x1, x2, y1, y2);

[0067] The maximum connected area and non-maximum suppression module includes gray value image conversion, image labeling, maximum area selection, contour detection, bounding box calculation, non-maximum suppression, bounding box adjustment, selecting appropriate samples as prompts, and obtaining the best bounding box (x1, x2, y1, y2).

[0068] Step 3: The pre-trained segmentation model SAM takes the optimal bounding box (x1, x2, y1, y2) as a prompt, combines the mask feature mask feat, the initialization Token, and the image feature F s , Class Text Related Embedding E t , identify and segment the target objects in the image, obtain the segmentation image mask, and then get the number of the required target objects.

[0069] The pre-trained segmentation model SAM consists of a SAM encoder, a prompt encoder, and a mask decoder. The pre-trained segmentation model SAM integrates a lightweight adapter and a CLIP feature fusion and mask generation module. The lightweight adapter is used to quickly and effectively adjust the SAM encoder to improve the segmentation model SAM's understanding of quantity and unseen categories. The CLIP feature fusion and mask generation module is used to convert image features F C Transfer and fuse into the mask decoder and guide the mask decoder to generate high-quality segmentation masks.

[0070] Specifically, in the cross-attention module of the CLIP feature fusion and mask generation module, the class text related embedding E t As Q, image features F C As K and V, we get X1, and then the image feature F C After a 1×1 convolution layer, a ReLU layer, and a second 1×1 convolution layer, X2 is obtained. In the multi-head self-attention module, X1 is used as Q, X2 is used as K and V, and a residual connection and layer normalization are used to obtain X3; then the mask feature mask feat of the CLIP image feature and text feature is obtained after an MLP layer, a second residual connection, and layer normalization. The formula is:

[0071] X3=LayNorm(E t +mh_attn(cross_attn(E t , F C , F C ), Conv(F C ), Conv(F C )))

[0072] mask feat=LayNorm(MLP(X3)+X3)

[0073] Among them, MLP is a multi-layer perceptron, LayNorm is layer normalization, mh_attn is multi-head self-attention, cross_attn is cross attention, Conv represents a 1×1 convolution, a ReLU and a second 1×1 convolution, mask feat∈R 1 ×1×256 , X3 is the middle value.

[0074] The SAM encoder extracts image features F from the image to be recognized s ∈R 1×d×64×64 , d = 256, using the best bounding box (x1, x2, y1, y2) as a hint for the hint encoder, the SAM encoder and the hint encoder generate similar feature maps Sim s , SAM encoder generates image features F s The optimal bounding box (x1, x2, y1, y2) is input to the hint encoder, and the hint encoder and SAM encoder work together to get Sim s , Sim s Then enter the prompt encoder to get the initialization Token) and set Sim s Pixel values ​​greater than the set threshold are set to 1, and those less than the set threshold are set to 0, as the next point prompt, and then prompt the encoder to use the point prompt to segment all the contents of the image.

[0075] Finally, initialize Token and image features F s , Class Text Related Embedding E t The mask feat generated by the CLIP feature fusion and mask generation module is input together into the mask decoder for recognition and segmentation processing to obtain a mask feature map, which is then upsampled inside the mask decoder to generate a segmentation result map of the same size as the original image. The number of masks is the number of desired targets.

[0076] Example 3

[0077] The above text-guided parameter-efficient fine-tuning method for fine-tuning image segmentation and counting models includes the following steps:

[0078] Step 1: Build and preprocess the dataset

[0079] The dataset is collected for few-shot counting, and each image contains a varying number of target objects. The images in the training set are resized to 512×512 and 1024×1024 as inputs to the CLIP image encoder and SAM encoder, respectively;

[0080] Step 2: Freeze the pre-trained visual language model CLIP and segmentation model SAM;

[0081] Step 3: Input the preprocessed image and the name of the target category to be counted into the text-guided parameter efficient fine-tuning image segmentation and counting model to obtain the segmented image mask;

[0082] Step 4: Calculate the loss function and update the parameters of the lightweight adapter and CLIP feature fusion and mask generation module; repeat steps 3 and 4 above until the training is completed.

[0083] Specifically, in step 1, the present invention selects the FSC147 dataset which is widely used in few-sample counting tasks. It consists of 6135 images of 147 object classes, of which the training set contains 89 object categories, the validation set and the test set each contain 29 categories, and the object categories of each set do not overlap.

[0084] Specifically, in step 4, the loss function formula is

[0085] L=L2+L t_ce +L dice

[0086] Where L2 is the MSE loss, L t_ce is the loss for mask classification, L dice is the loss used for segmentation.

[0087] The present invention is trained and tested on an Nvidia RTX-3090 GPU, with the batch size set to 2, the initial learning rate set to 1e-6, the AdamW optimizer used, and β1=0.9, β2=0.999. It takes about two hours to train six epochs.

[0088] Example 4

[0089] This example uses the method proposed by the present invention and the existing method to make a quantitative comparison on the FSC147 data set, and the results are shown in Table 1. Using the mean absolute error (MAE) and the mean square error (RMSE) as evaluation indicators, it is not difficult to find that the effect of the present invention is better.

[0090] Table 1

[0091] Methods MAE RMSE Xuetal 22.09 115.17 SAM 42.48 137.50 Training-free 24.79 137.15 Ours 28.84 92.89

[0092] The qualitative results of the present invention on the FSC147 data set are as follows Figure 5 As shown, it can be seen that the segmentation counting method proposed in the present invention has a good counting effect in both sparse and dense scenes. In short, the method proposed in the present invention effectively improves the counting performance under the condition of a small amount of computing cost by using a pre-trained large model.

[0093] The above is only a preferred embodiment of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A text-guided parameter-efficient fine-tuning image segmentation and counting model, characterized in that: It includes the pre-trained visual language model CLIP, the maximum connected region and non-maximum suppression module, and the pre-trained segmentation model SAM, among which: The pre-trained visual language model CLIP includes a pre-trained CLIP image encoder and a standard text encoder. The CLIP image encoder extracts image features F from the image. C , the standard text encoder extracts text features T from the name of the target category to be counted r , redundant features T r The pre-trained visual language model CLIP uses text features T t , redundant features T r and image features F C The class-related token cls_token in the class is used to calculate the class text-related embedding E t And cosine similarity Sim c ; The maximum connected region and non-maximum suppression module are connected to the pre-trained visual language model CLIP to obtain the cosine similarity Sim c , generate the best bounding box (x1,x2,y1,y2) and pass it to the pre-trained segmentation model SAM as a hint; The pre-trained segmentation model SAM includes a SAM encoder, a hint encoder and a mask decoder, and the SAM encoder is connected with the hint encoder to generate an image feature F s The hint encoder is connected to the maximum connected area and non-maximum suppression module to receive the best bounding box (x1, x2, y1, y2) and form an initialization Token. The pre-trained segmentation model SAM also integrates a lightweight adapter and a CLIP feature fusion and mask generation module. The CLIP feature fusion and mask generation module is connected to the CLIP image encoder and the standard text encoder to convert the image feature F C and class text related embedding E t The light-weight adapter is used to adjust the SAM encoder. The prompt encoder, CLIP feature fusion and mask generation module, and SAM encoder are connected to the mask decoder respectively. The mask decoder combines the mask feature mask feat, the initialization Token, and the image feature F s , Class Text Related Embedding E t Generate a segmentation mask and get the required number of target objects.

2. The text-guided parameter-efficient fine-tuning image segmentation and counting model as claimed in claim 1, characterized in that: The CLIP feature fusion and mask generation module includes a cross-attention module, two 1×1 convolutional layers and a ReLU, a multi-head self-attention module, a residual connection and layer normalization, an MLP layer, a secondary residual connection and layer normalization.

3. The text-guided parameter-efficient fine-tuning image segmentation and counting model as claimed in claim 1, characterized in that: The SAM encoder consists of N layers, each of which is composed of a first layer normalization, a multi-head attention layer, a second layer normalization and an MLP layer connected in sequence.

4. The text-guided parameter-efficient fine-tuning image segmentation and counting model as claimed in claim 3, characterized in that: There are two lightweight adapters, each of which is composed of a Down linear layer, a GELU, and an Up linear layer connected in sequence. The first lightweight adapter is located after the multi-head self-attention layer of the SAM encoder and is connected to the residual of the previous step. The second lightweight adapter is located in the residual path of the MLP layer of the SAM encoder.

5. The image segmentation and counting method based on the text-guided parameter-efficient fine-tuning image segmentation and counting model as claimed in claim 1, characterized in that: The following steps are involved: Step 1: The pre-trained CLIP image encoder extracts image features F from the image C , the image feature F C Contains the class-related token cls_token, and the standard text encoder extracts the text features T of the category to be counted from the name of the target category to be counted t , redundant features T r , embed the class-related token cls_token into the text feature T t Then, with the text feature T t Perform secondary combination to generate the image-specific class text related embedding E t , class text-related embedding E t Subtract redundant features T r The information combined with the class-related token cls_token is finally combined with the CLIP image feature F C Calculate cosine similarity Sim c ; The CLIP feature fusion and mask generation module fuses the image features F C and class text related embedding E t Get the mask feature mask feat; Step 2: The maximum connected area and non-maximum suppression module obtain the Sim from step 1. c Get the best bounding box (x1,x2,y1,y2); Step 3: The pre-trained segmentation model SAM takes the optimal bounding box (x1, x2, y1, y2) as a prompt, combines the mask feature mask feat, the initialization Token, and the image feature F s , Class Text Related Embedding E t , identify and segment the target objects in the image, obtain the segmentation image mask, and then get the number of the required target objects.

6. The image segmentation and counting method according to claim 5, characterized in that: Class Text Dependent Embedding E t =concat[cls_token⊙T t ,T t ],in represents the cosine similarity measure, ⊙ represents the Hadamard product, concat is a function that concatenates multiple arrays, and cls_token is a class-related token.

7. The image segmentation and counting method according to claim 5, characterized in that: In step 3, the optimal bounding box (x1, x2, y1, y2) is input to the prompt encoder, and Sim is obtained under the joint action of the prompt encoder and the SAM encoder. s , Sim s Then enter the prompt encoder to get the initialization Token; The SAM encoder extracts image features F from the image to be recognized s ; The CLIP feature fusion and mask generation module fuses the image features F C and class text related embedding E t Get the mask feature mask feat; Initialize Token and image features F s , Class Text Related Embedding E t The mask feature mask feat and the mask feature mask feat are input into the mask decoder for recognition and segmentation processing to obtain the number of target objects.

8. The image segmentation and counting method according to claim 5, characterized in that: mask feat=LayNorm(MLP(X3)+X3) X3=LayNorm(E t +mh_attn(cross_attn(E t ,F C ,F C ),Conv(F C ),Conv(F C ))) Among them, MLP is a multi-layer perceptron, LayNorm is layer normalization, mh_attn is multi-head self-attention, cross_attn is cross attention, Conv means one 1×1 convolution, one ReLU and two 1×1 convolutions, and X3 is the intermediate value.

9. The text-guided parameter-efficient fine-tuning method for fine-tuning an image segmentation and counting model as claimed in claim 1, characterized in that: The following steps are involved: Step 1: Build a dataset and adjust the size of each image in the dataset so that it can be used as the input of the CLIP image encoder and the SAM encoder. Step 2: Freeze the pre-trained visual language model CLIP and segmentation model SAM; Step 3, input the preprocessed image and the category name of the target to be counted into the text-guided parameter efficient fine-tuning image segmentation and counting model to obtain the segmentation image mask and the number of targets to be counted; Step 4: Calculate the loss function and update the parameters of the lightweight adapter and CLIP feature fusion and mask generation module; Repeat steps 3 and 4 above until the training is completed.

10. The fine-tuning method according to claim 9, characterized in that: The loss function is calculated as: L=L2+L t_ce +L dice Where L2 is the MSE loss, L t_ce is the loss for mask classification, L dice is the loss used for segmentation.

Citation Information

Cited By

  • Entity segmentation method and device based on dynamic programming, equipment and medium

    CN120451561A

  • An entity segmentation method, device, equipment and medium based on dynamic programming

    CN120451561B

  • Cervical cancer MRI image automatic segmentation method based on multi-modal fusion

    CN120931927A

  • Cervical cancer MRI image automatic segmentation method based on multi-modal fusion

    CN120931927B