Fine adjustment method and device for image semantic segmentation model

By introducing cross attention adapters into the image semantic segmentation model and only fine-tuning it, the problems of low performance and large storage overhead of general models in semantic segmentation tasks are solved, and a more efficient image semantic segmentation effect is achieved.

CN120219731APending Publication Date: 2025-06-27BEIJING JINGWEI HIRAIN TECH CO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311797854.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The general model for image segmentation tasks lacks semantic information annotation in the pre-training stage, which makes it difficult to complete downstream semantic segmentation tasks. The traditional fine-tuning method has the problem of large storage overhead or low performance.

Method used

A fine-tuning method for image semantic segmentation model is proposed. By constructing a model including a cross attention adapter, an image encoder and a convolutional segmentation head, and only fine-tuning of the cross attention adapter is combined with multi-scale feature extraction to improve the segmentation effect.

Benefits of technology

It effectively reduces the overhead of fine-tuning of general models, and improves the effect of image semantic segmentation, which can better segment objects of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219731A_ABST
    Figure CN120219731A_ABST
Patent Text Reader

Abstract

The invention discloses a fine adjustment method and device for an image semantic segmentation model. When the method is executed, an image semantic segmentation model to be fine-tuned is constructed, and the image semantic segmentation model to be fine-tuned comprises a cross attention adapter, an image encoder and a convolution integral cutting head; and acquiring a preprocessed image semantic segmentation data set, then according to the image semantic segmentation data set and the loss function, performing fine tuning on a cross attention adapter in the to-be-fine-tuned image semantic segmentation model until a model convergence condition is reached, and obtaining a fine-tuned image semantic segmentation model. In the fine tuning process of the image semantic segmentation model, on one hand, the parameters of the general model are fixed, only a part of learnable parameters are introduced, and the overhead of fine tuning of the general model oriented to the image segmentation task is reduced, and on the other hand, the adapter module based on cross attention is combined. The image semantic segmentation fine tuning performance of a general model can be effectively improved, and then the image semantic segmentation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to a method and device for fine-tuning an image semantic segmentation model. Background Art

[0002] A general model refers to a model trained on a large-scale dataset. These models learn general knowledge related to downstream tasks from the large-scale data. Therefore, the parameters of the general model are used to initialize the downstream task model, and then the general model is fine-tuned using downstream data.

[0003] The general model (Segment Anything Model, SAM) for the image segmentation task uses binary segmentation data in the pre-training stage and lacks corresponding semantic information annotation, resulting in difficulty for the SAM model to complete the downstream semantic segmentation task.

[0004] Traditional general model fine-tuning techniques mainly include the method based on full-parameter fine-tuning and the method based on linear probing. The method based on full-parameter fine-tuning uses downstream task data to train the general model. When the number of parameters of the general model is large, this method will bring huge storage overhead. The method based on linear probing fixes all the parameters of the general model and relies on the performance of the general model itself. When the performance of the general model itself is poor, the performance of the downstream task of this method is often low, resulting in poor image semantic segmentation effect. Summary of the Invention

[0005] In view of this, this application provides a method and device for fine-tuning an image semantic segmentation model, aiming to reduce the overhead during the fine-tuning of the general model for the image segmentation task and improve the effect of image semantic segmentation.

[0006] In a first aspect, this application provides a method for fine-tuning an image semantic segmentation model, the method comprising:

[0007] Construct an image semantic segmentation model to be fine-tuned, the image semantic segmentation model to be fine-tuned comprising a cross-attention adapter, an image encoder, and a convolutional segmentation head; wherein, the image encoder is configured to extract features of an input image and output a feature map corresponding to the input image; the cross-attention adapter is configured to extract target multi-scale features of the input image; the convolutional segmentation head is configured to perform a segmentation operation on the feature map and output a semantic segmentation map of the input image;

[0008] Obtain a preprocessed image semantic segmentation dataset;

[0009] According to the image semantic segmentation data set and the loss function, fine-tune the cross-attention adapter in the image semantic segmentation model to be fine-tuned until the model convergence condition is reached, and obtain the fine-tuned image semantic segmentation model.

[0010] Optionally, the image encoder is a neural network based on visual attention;

[0011] The image encoder is used to extract the features of the input image and output the feature map corresponding to the input image, including:

[0012] The image encoder alternately uses the global attention mechanism and the window attention mechanism to extract the features of the input image and output the feature map corresponding to the input image; wherein, the image encoder uses relative position encoding in each layer to enhance the relative position information of the model output feature map.

[0013] Optionally, the cross-attention adapter includes a local spatial prior module, a spatial feature injection module, and a multi-scale feature extraction module;

[0014] The cross-attention adapter is used to extract the multi-scale features of the input image, including:

[0015] The local spatial prior module extracts the multi-scale local spatial features of the input image through multiple deep convolutional neural networks;

[0016] The spatial feature injection module injects the multi-scale local spatial features into the image encoder by using the cross-attention mechanism;

[0017] The multi-scale feature extraction module extracts multi-scale features from the image encoder by using the cross-attention mechanism;

[0018] The spatial feature injection module injects the multi-scale features into the image encoder.

[0019] Optionally, the local spatial prior module uses multiple convolutional neural networks with different depths to extract the multi-scale local spatial features of the input image, including:

[0020] The local spatial prior module extracts local spatial features of multiple scales from the input image through the multiple convolutional neural networks respectively;

[0021] The local spatial prior module maps the local spatial features of the multiple scales through the fully connected layers of the multiple convolutional neural networks to obtain the local spatial features of multiple scales under the same feature dimension;

[0022] The local spatial prior module splices the local spatial features at multiple scales in the same feature dimension to obtain the multi-scale local spatial features of the input image.

[0023] Optionally, the spatial feature injection module is a neural network based on the cross-attention mechanism;

[0024] The spatial feature injection module injects the multi-scale local spatial features into the image encoder by using the cross-attention mechanism, including:

[0025] The spatial feature injection module uses the cross-attention mechanism to obtain the injection features according to the multi-scale local spatial features and the output features of the i-th layer of the image encoder; where i is a positive integer;

[0026] The spatial feature injection module injects the injection features into the (i + 1)-th layer of the image encoder.

[0027] Optionally, the multi-scale feature extraction module is a neural network based on the cross-attention mechanism;

[0028] The multi-scale feature extraction module extracts multi-scale features from the image encoder by using the cross-attention mechanism, including:

[0029] The multi-scale feature extraction module receives the output features of the i-th layer of the image encoder and the spatial features of the i-th layer;

[0030] The multi-scale feature extraction module uses the cross-attention mechanism to obtain the multi-scale features of the i-th layer according to the spatial features of the i-th layer and the output features of the i-th layer of the image encoder;

[0031] The multi-scale feature extraction module uses a feed-forward neural network to perform a normalization operation on the second multi-scale features to obtain the multi-scale features of the (i + 1)-th layer.

[0032] Optionally, the convolutional segmentation head is a neural network including a convolutional layer and an upsampling layer;

[0033] The convolutional segmentation head is used to perform a segmentation operation on the feature map and output the semantic segmentation map of the input image, including:

[0034] The convolutional segmentation head receives the low-resolution feature map output by the image encoder as input, and alternately performs convolutional operations and upsampling operations on the low-resolution feature map to obtain the semantic segmentation map of the input image.

[0035] Optionally, the method further includes:

[0036] Obtaining an image to be segmented;

[0037] Input the image to be segmented into the image semantic segmentation model to obtain the segmentation result of the image to be segmented.

[0038] In a second aspect, the present application provides a fine-tuning device for an image semantic segmentation model. The device includes:

[0039] A construction module for constructing an image semantic segmentation model to be fine-tuned. The image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head. Among them, the image encoder is used to extract the features of the input image and output the feature map corresponding to the input image. The cross-attention adapter is used to extract the target multi-scale features of the input image. The convolutional segmentation head is used to perform a segmentation operation on the feature map and output the semantic segmentation map of the input image.

[0040] An acquisition module for acquiring a preprocessed image semantic segmentation data set.

[0041] A fine-tuning module for fine-tuning the cross-attention adapter in the image semantic segmentation model to be fine-tuned according to the image semantic segmentation data set and the loss function until the model convergence condition is reached, and obtaining the fine-tuned image semantic segmentation model.

[0042] Optionally, the image encoder is a neural network based on visual attention.

[0043] Specifically, the image encoder is used to alternately use the global attention mechanism and the window attention mechanism to extract the features of the input image and output the feature map corresponding to the input image. Among them, the image encoder uses relative position encoding in each layer to enhance the relative position information of the model output feature map.

[0044] In a third aspect, an embodiment of the present application provides an electronic device. The electronic device includes:

[0045] A memory for storing one or more programs.

[0046] A processor; when the one or more programs are executed by the processor, the fine-tuning method of the image semantic segmentation model according to any one of the foregoing first aspects is implemented.

[0047] In a fourth aspect, an embodiment of the present application provides a computer storage medium. A program is stored in the computer storage medium. When the program is executed by a processor, the fine-tuning method of the image semantic segmentation model according to any one of the foregoing first aspects is implemented.

[0048] The above technical solutions have the following beneficial effects:

[0049] An embodiment of the present application provides a method and device for fine-tuning an image semantic segmentation model. When executing the method, an image semantic segmentation model to be fine-tuned is constructed. The image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head. Among them, the image encoder is used to extract the features of the input image and output the feature map corresponding to the input image. The cross-attention adapter is used to extract the target multi-scale features of the input image. The convolutional segmentation head is used to perform segmentation operations on the feature map and output the semantic segmentation map of the input image. The preprocessed image semantic segmentation data set is obtained, and then the cross-attention adapter in the image semantic segmentation model to be fine-tuned is fine-tuned according to the image semantic segmentation data set and the loss function until the model convergence condition is reached, and the fine-tuned image semantic segmentation model is obtained. By the above method, in the process of fine-tuning the image semantic segmentation model, on the one hand, the parameters of the general model itself are fixed, and only a part of the learnable parameters are introduced, reducing the overhead of fine-tuning the general model for the image segmentation task. On the other hand, an adapter module based on cross-attention is combined. By obtaining the multi-scale information of the input image, the performance of fine-tuning the general model for image semantic segmentation can be effectively improved, which helps to segment objects of different sizes, and thus improves the effect of image semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] To more clearly illustrate the technical solutions in the embodiments or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.

[0051] Figure 1 FIG. 9 is a flowchart of a method for fine-tuning an image semantic segmentation model provided by an embodiment of the present application;

[0052] FIG. 2(a) is a structural diagram of a local spatial prior module provided by an embodiment of the present application;

[0053] FIG. 2(b) is a structural diagram of a spatial feature injection module provided by an embodiment of the present application;

[0054] FIG. 2(c) is a structural diagram of a multi-scale feature extraction module provided by an embodiment of the present application;

[0055] Figure 3 FIG. 22 is a schematic diagram of an image semantic segmentation model provided by an embodiment of the present application;

[0056] Figure 4 FIG. 26 is a schematic structural diagram of a device for fine-tuning an image semantic segmentation model provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0058] To facilitate a further understanding of the technical solutions provided in the present application, the background art related to the present application will be described first below.

[0059] The general model (Segment Anything Model, SAM) for image segmentation tasks uses binary segmentation data in the pre-training stage and lacks corresponding semantic information annotation, resulting in difficulty for the SAM model to complete downstream semantic segmentation tasks.

[0060] Traditional general model fine-tuning techniques mainly include the method based on full-parameter fine-tuning and the method based on linear probing. The method based on full-parameter fine-tuning uses the data of the downstream task to train all the parameters in the general model and the task module. Such methods usually have high performance, but they need to independently store complete copies of the general model and the task module for each downstream task. When the number of parameters of the general model is large, such methods will bring huge storage overhead.

[0061] The method based on linear probing fixes all the parameters of the general model and only uses the data of the downstream task to train the parameters in the task module. Since the number of parameters of the task module is often much smaller than that of the general model, such methods can save storage space. However, such methods rely heavily on the performance of the general model itself. When the performance of the general model itself is poor, the performance of the downstream task of such methods is often low.

[0062] Therefore, how to reduce the number of learnable parameters during general model fine-tuning has become one of the problems that need to be solved urgently. Considering the low performance of the traditional method based on linear probing, the linear probing method has been extended. These extension works mainly include the method based on partial fine-tuning, the method based on prompt fine-tuning, and the method based on auxiliary parameters.

[0063] The method based on partial fine-tuning achieves a better balance between the performance of the downstream task and the number of learnable parameters by fine-tuning some parameters in the general model. For example, the bias fine-tuning method migrates the general model to the downstream task by adjusting the bias terms in the network layers of the general model. However, these methods need to search for the best fine-tuning parameters in the general model, which will bring a large computational overhead when the number of parameters of the general model is high, and at the same time, the performance of such methods is low.

[0064] The method of fine-tuning based on prompts optimizes the transfer of the general model by inserting learnable prompt vectors into the general model and optimizing these prompt vectors during the fine-tuning stage. However, these methods ignore the multi-scale information of the input images being modeled, resulting in difficulty in perceiving small-sized objects in downstream dense prediction tasks. The method based on auxiliary parameters adds learnable lightweight modules to the network structure of the general model and completes the transfer of the general model by fine-tuning these modules. For example, in the method based on low-rank factorization of the parameter matrix, the fully connected layer in the general model is first factorized into low rank, and then the low-rank factorization matrix is fine-tuned to achieve efficient parameter fine-tuning of the general model.

[0065] The method of fine-tuning the model based on adapters inserts adapter modules with bottleneck structures into the original general model, fixes all the parameters of the general model during the fine-tuning stage, and only learns these lightweight adapters to complete the transfer of the general model. However, these methods also ignore the multi-scale information of the modeled images, resulting in low performance in dense prediction tasks such as semantic segmentation tasks.

[0066] To overcome the above technical problems, the embodiments of the present application provide a method for fine-tuning an image semantic segmentation model. This method can be executed by a fine-tuning device for the image semantic segmentation model, which can be implemented in software and / or hardware and is generally integrated into a server or a terminal device.

[0067] See Figure 1 , Figure 1 FIG. is a flowchart of a method for fine-tuning an image semantic segmentation model provided by an embodiment of the present application.

[0068] This method may include:

[0069] Step S101: Construct an image semantic segmentation model to be fine-tuned. The image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head; wherein, the image encoder is used to extract the features of the input image and output the feature map corresponding to the input image; the cross-attention adapter is used to extract the target multi-scale features of the input image; the convolutional segmentation head is used to perform a segmentation operation on the feature map and output the semantic segmentation map of the input image.

[0070] In this embodiment, the fine-tuning device of the image semantic segmentation model constructs an image semantic segmentation model to be fine-tuned, and the image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head.

[0071] Among them, the image encoder uses the ViT-L low-resolution variant model in the Vision Transformer (ViT) network based on the Transformer architecture, which contains 24 attention layers. Among them, the 5th, 11th, 17th, and 23rd layers use the global attention mechanism, and the remaining layers use the window attention mechanism. The image encoder is used to extract the features of the input image and output the feature map corresponding to the input image.

[0072] The cross-attention adapter is used to extract the target multi-scale features of the input image; the convolutional segmentation head is used to perform segmentation operations on the feature map and output the semantic segmentation map of the input image.

[0073] It can be understood that the image semantic segmentation model to be fine-tuned in the embodiments of the present application includes a cross-attention adapter, which gives full play to the advantages of cross-attention in modeling multi-scale features, makes up for the defect of insufficient modeling of multi-modal features, effectively improves the performance of fine-tuning the general model to the semantic segmentation task, and completes the downstream image semantic segmentation task.

[0074] Step S102: Obtain the preprocessed image semantic segmentation dataset.

[0075] In this embodiment, the fine-tuning method of the image semantic segmentation model can be applied to image semantic segmentation in the autonomous driving scenario.

[0076] Therefore, the preprocessed image semantic segmentation dataset that can be obtained is the ADE20K general scene semantic segmentation dataset and the CityScapes city street scene semantic segmentation dataset.

[0077] Among them, data preprocessing may include data augmentation of training images and label images. Specifically, image augmentation includes resizing, random cropping, random flipping, and image perturbation, etc. For example, in this example, the image size is selected to be resized to 2049×1025 pixels, a random cropped image area of 768×768 pixels, and a 50% probability of random horizontal flipping. For image perturbation augmentation, four image perturbation methods of randomly adjusting contrast, randomly adjusting brightness, randomly adjusting hue, and randomly adjusting saturation are selected. It can be understood that the same data augmentation method is adopted for the training images and label images in this embodiment to ensure the consistency of data and labels.

[0078] The data is divided according to the standards given by each dataset. For example: for the ADE20K dataset, 25,574 images are selected for training and 2,000 images are selected for validation. For the CityScapes dataset, 2,975 images are selected for training and 500 images are selected for validation.

[0079] Step S103: According to the image semantic segmentation data set and the loss function, fine-tune the cross-attention adapter in the image semantic segmentation model to be fine-tuned until the model convergence condition is reached, and obtain the fine-tuned image semantic segmentation model.

[0080] In this embodiment, during the fine-tuning process, the image encoder is initialized with a general model, and the parameters of the image encoder part are fixed. The fine-tuning device of the image semantic segmentation model calculates the corresponding loss value according to the image semantic segmentation data set and the adopted loss function, and then optimizes the model parameters of the image semantic segmentation model to be fine-tuned based on the backpropagation method, and fine-tunes the cross-attention adapter in the image semantic segmentation model to be fine-tuned until the model convergence condition is reached, and obtains the fine-tuned image semantic segmentation model.

[0081] Specifically, the loss function can be the pixel-level cross-entropy loss function L pixel , which calculates the loss between the image and the target segmentation result pixel by pixel and calculates its average value:

[0082]

[0083] Among them, and represent the label and prediction result at the pixel position corresponding to the i-th row and the j-th column. H is the width of the input image, and W is the length of the input image.

[0084] There are various situations for reaching the model convergence condition. For example, when the loss value is less than a preset value, it is considered to reach the model convergence condition. Another example is when the change amount of the model parameters between two iterations is less than or equal to the change threshold, it is considered to reach the model convergence condition.

[0085] An embodiment of the present application provides a method and device for fine-tuning an image semantic segmentation model. When executing the method, an image semantic segmentation model to be fine-tuned is constructed. The image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head. Among them, the image encoder is used to extract the features of the input image and output the feature map corresponding to the input image. The cross-attention adapter is used to extract the target multi-scale features of the input image. The convolutional segmentation head is used to perform segmentation operations on the feature map and output the semantic segmentation map of the input image. A preprocessed image semantic segmentation data set is obtained, and then, according to the image semantic segmentation data set and the loss function, the cross-attention adapter in the image semantic segmentation model to be fine-tuned is fine-tuned until the model convergence condition is reached, and a fine-tuned image semantic segmentation model is obtained. By the above method, during the fine-tuning process of the image semantic segmentation model, on the one hand, the parameters of the general model itself are fixed, and only a part of the learnable parameters are introduced, reducing the overhead during the fine-tuning of the general model for the image segmentation task. On the other hand, an adapter module based on cross-attention is combined. By obtaining the multi-scale information of the input image, the performance of the fine-tuning of the general model for image semantic segmentation can be effectively improved, which helps to segment objects of different sizes, and thus improves the effect of image semantic segmentation.

[0086] Further, in a possible implementation manner, the image encoder is a neural network based on visual attention. The image encoder is used to extract the features of the input image and output the feature map corresponding to the input image, including:

[0087] The image encoder alternately uses the global attention mechanism and the window attention mechanism to extract the features of the input image and output the feature map corresponding to the input image. Among them, the image encoder uses relative position encoding in each layer to enhance the relative position information of the model output feature map.

[0088] Specifically, for an input image with a length of W and a width of H The image encoder extracts the features of the input image and outputs a low-resolution feature map where s represents the downsampling step of the feature map resolution, D represents the feature map dimension, and R represents the real number field.

[0089] In addition, different from the standard ViT that uses the global attention mechanism in each layer and only uses absolute position encoding in the first layer, the image encoder is optimized in terms of the attention mechanism and position encoding. Considering the consistency of the semantic information in the local part of the image, using the window-based attention mechanism can better model the features of adjacent pixels. Therefore, the image encoder alternately uses the global attention mechanism and the window attention mechanism. At the same time, considering that the relative position information of the pixels in the image helps the model segment objects of different shapes, the image encoder uses relative position encoding in each layer to enhance the relative position information of the model output feature map.

[0090] In one possible implementation, the cross-attention adapter includes a local spatial prior module, a spatial feature injection module, and a multi-scale feature extraction module;

[0091] The cross-attention adapter is used to extract multi-scale features of the input image, including:

[0092] The local spatial prior module extracts multi-scale local spatial features of the input image through multiple deep convolutional neural networks;

[0093] The spatial feature injection module injects the multi-scale local spatial features into the image encoder by using the cross-attention mechanism;

[0094] The multi-scale feature extraction module extracts multi-scale features from the image encoder by using the cross-attention mechanism;

[0095] The spatial feature injection module injects the multi-scale features into the image encoder.

[0096] Specifically, the cross-attention adapter can include a standard cross-attention adapter and an enhanced cross-attention adapter. The enhanced cross-attention adapter further optimizes the multi-scale feature extraction module. By stacking multiple feature extraction modules, the multi-scale feature extraction ability of the model is enhanced. In addition, the enhanced cross-attention adapter additionally optimizes the bias term in the image encoder, further enhancing the feature extraction ability of the image encoder.

[0097] It should be noted that compared with the standard cross-attention adapter, the enhanced cross-attention adapter has more learnable parameters and better performance. When the user's storage resources are limited, the standard cross-attention adapter can be used. When the user has a high demand for performance, the enhanced cross-attention adapter can be used.

[0098] In this embodiment, the cross-attention adapter includes a local spatial prior module, a spatial feature injection module, and a multi-scale feature extraction module.

[0099] The local spatial prior module first extracts multi-scale local spatial features of the input image through multiple deep convolutional neural networks. Then, the spatial feature injection module injects the multi-scale local spatial features into the image encoder using the cross-attention mechanism. Next, the multi-scale feature extraction module extracts multi-scale features from the image encoder using the cross-attention mechanism. Finally, the spatial feature injection module injects the multi-scale features into the image encoder.

[0100] It should be noted that in the embodiments of the present application, an adapter module based on cross-attention is combined. By obtaining multi-scale information of the input image, the performance of fine-tuning the image semantic segmentation of the general model can be effectively improved, which helps to segment objects of different sizes, and thus improves the effect of image semantic segmentation.

[0101] In a possible implementation manner, the local spatial prior module uses multiple convolutional neural networks with different depths to extract multi-scale local spatial features of the input image, which may include the following steps:

[0102] The local spatial prior module extracts local spatial features of multiple scales from the input image through the multiple convolutional neural networks respectively;

[0103] The local spatial prior module maps the local spatial features of multiple scales through the fully connected layers of the multiple convolutional neural networks to obtain local spatial features of multiple scales under the same feature dimension;

[0104] The local spatial prior module splices the local spatial features of multiple scales under the same feature dimension to obtain the multi-scale local spatial features of the input image.

[0105] Specifically, as shown in Fig. 2(a), it is a structural diagram of the local spatial prior module provided by the embodiments of the present application.

[0106] The local spatial prior module includes three convolutional neural networks {f1, f2, f3} with different depths. For an input image with a length of W and a width of H Local spatial features of multiple scales are respectively extracted from the input image through multiple convolutional neural networks, such as z1, z2, and z3 in Fig. 2(a).

[0107] Among them,

[0108] Subsequently, the local spatial prior module maps the local spatial features of different dimensions to a unified feature dimension D through the fully connected layers of multiple convolutional neural networks. The local spatial prior module splices the local spatial features of multiple scales under the same feature dimension to obtain the multi-scale local spatial features of the input image.

[0109] The specific calculation process is as follows:

[0110]

[0111]

[0112] where f i (·) represents the i-th convolutional neural network, and W i represents the dimensionality mapping matrix, represents the bias vector, and is the multi-scale local spatial feature.

[0113] Furthermore, in a possible implementation, the spatial feature injection module is a neural network based on the cross-attention mechanism;

[0114] The spatial feature injection module injects the multi-scale local spatial feature into the image encoder by using the cross-attention mechanism, including:

[0115] The spatial feature injection module uses the cross-attention mechanism to obtain the injection feature according to the multi-scale local spatial feature and the output feature of the i-th layer of the image encoder; where i is a positive integer;

[0116] The spatial feature injection module injects the injection feature into the (i + 1)-th layer of the image encoder.

[0117] Specifically, as shown in FIG. 2(b), it is a structural diagram of the spatial feature injection module provided by the embodiment of the present application.

[0118] In this embodiment, the spatial feature injection module is a neural network with a cross-attention mechanism, which can inject the multi-scale local spatial feature extracted by the local spatial prior module into the image encoder, helping to segment objects of different sizes.

[0119] The spatial feature injection module receives the feature output by the i-th layer of the image encoder and the multi-scale local spatial feature as inputs, uses the output feature as the query, and the spatial feature as the key and value to calculate the attention weight. Perform an attention operation on the spatial feature based on the attention weight to obtain the injection feature

[0120] The specific calculation process is as follows:

[0121]

[0122] where Norm(·) represents the layer normalization operation, Attn(q, k, v) represents the attention operation, and γ iRepresents a zero-initialized learnable scalar used to balance the ratio of spatial features and output features. They are multi-scale local spatial features. The zero-initialization strategy ensures that the output feature distribution of the image encoder does not deviate too much due to the injection of multi-scale features.

[0123] Furthermore, in a possible implementation, the multi-scale feature extraction module is a neural network based on the cross-attention mechanism;

[0124] The multi-scale feature extraction module uses the cross-attention mechanism to extract multi-scale features from the image encoder, including:

[0125] The multi-scale feature extraction module receives the output features of the i-th layer of the image encoder and the spatial features of the i-th layer;

[0126] The multi-scale feature extraction module uses the cross-attention mechanism to obtain the multi-scale features of the i-th layer according to the spatial features of the i-th layer and the output features of the i-th layer of the image encoder;

[0127] The multi-scale feature extraction module uses a feed-forward neural network to perform a normalization operation on the second multi-scale feature to obtain the multi-scale features of the (i + 1)-th layer.

[0128] In this embodiment, as shown in Fig. 2(c), it is a structural diagram of the multi-scale feature extraction module provided by the embodiment of the present application.

[0129] The multi-scale feature extraction module is a neural network containing cross-attention and can extract multi-scale features from the image encoder.

[0130] The multi-scale feature extraction module receives the features output by the i-th layer of the image encoder and the spatial features of the i-th layer as inputs, uses the spatial features as queries, and the output features as keys and weights to calculate the attention weights. An attention operation is performed on the output features based on the attention weights to obtain the multi-scale features of the i-th layer The multi-scale features of the i-th layer are processed through normalization and a feed-forward neural network to obtain the multi-scale features of the (i + 1)-th layer as the input to the next spatial feature injection module.

[0131] The specific calculation process is as follows:

[0132]

[0133]

[0134] Among them, FFN(·) represents a feed-forward neural network, and Attn(q, k, v) represents an attention operation.

[0135] In a possible implementation, the convolutional segmentation head is a neural network including a convolutional layer and an upsampling layer.

[0136] The convolutional segmentation head is used to perform a segmentation operation on the feature map and output a semantic segmentation map of the input image, including:

[0137] The convolutional segmentation head receives the low-resolution feature map output by the image encoder as input, and alternately performs a convolutional operation and an upsampling operation on the low-resolution feature map to obtain the semantic segmentation map of the input image.

[0138] Specifically, the convolutional segmentation head is a neural network including L convolutional layers and an upsampling layer, and can receive the low-resolution feature map output by the image encoder as input. The convolutional segmentation head alternately performs convolution and upsampling operations on the low-resolution feature map and outputs a segmentation map with the same size as the original image.

[0139] The specific calculation process is as follows:

[0140] z i+1 = UpSample i (Conv i (z i ), r i ), i = 1,..., L;

[0141] p img = Conv cls (z i+1 );

[0142] Where r i represents the upsampling step size, UpSample(·) represents the upsampling operation, Conv i (·) and Conv cls (·) respectively represent the i-th convolutional layer and the 1×1 convolutional layer, and p img is the segmentation map.

[0143] In a possible implementation, the method further includes:

[0144] Obtain an image to be segmented;

[0145] Input the image to be segmented into the image semantic segmentation model to obtain a segmentation result of the image to be segmented.

[0146] Specifically, in the embodiment of the present application, after obtaining the preprocessed image to be segmented, the image to be segmented is input into the image semantic segmentation model to obtain a segmentation result of the image to be segmented.

[0147] Next, a specific embodiment will be used for illustration.

[0148] 1. Dataset preparation. Complete dataset selection, data preprocessing, and data partitioning.

[0149] 1.1. In this example, the general scene semantic segmentation dataset ADE20K and the city street scene semantic segmentation dataset CityScapes are selected to verify the efficient fine-tuning method of the general model of the invention.

[0150] 1.2. Data preprocessing includes data augmentation for training images and label images. Specifically, image augmentation includes resizing, random cropping, random flipping, and image perturbation, etc. In this example, the image size is resized to 2049×1025 pixels, a random image region of 768×768 pixels is cropped, and a 50% probability of random horizontal flipping is selected. For image perturbation augmentation, four image perturbation methods of randomly adjusting contrast, randomly adjusting brightness, randomly adjusting hue, and randomly adjusting saturation are selected. The same data augmentation method is adopted for training images and label images to ensure the consistency of data and labels.

[0151] 1.3. Data partitioning is based on the given criteria of each dataset. For example: for the ADE20K dataset, 25,574 images are selected for training and 2,000 images are selected for validation. For the CityScapes dataset, 2,975 images are selected for training and 500 images are selected for validation.

[0152] 2. Build a semantic segmentation model. As Figure 3 shown, it is a schematic diagram of an image semantic segmentation model provided by an embodiment of the present application. The image semantic segmentation model includes an image encoder, a convolutional segmentation head, and a cross-attention adapter. Among them, the image encoder uses a ViT-L network, which includes an embedding layer and 24 attention VIT layers.

[0153] In a possible implementation, the global attention mechanism is used in the 5th, 11th, 17th, and 23rd layers of the image encoder, and the window attention mechanism is used in the remaining layers. For the ADE10K dataset, the number of semantic categories C is set to 150. For the CityScapes dataset, the number of semantic categories C is set to 19. The number N of feature extraction modules of the enhanced cross-attention adapter is set to 3.

[0154] 3. Initialize the image encoder, design the model training loss, and train the semantic segmentation model. Initialize the image encoder with the ViT-L general model pre-trained using the SegmentAnything strategy, and use the pixel-level cross-entropy loss function to constrain the model training. Update and optimize the parameters of the cross-attention adapter and convolutional segmentation head using the backpropagation algorithm until the model converges. In this example, both model training and evaluation are completed on the Pytorch platform. The model is trained using two NVIDIA A800 GPUs (80GB), and the batch size is set to 2. Use the AdamW optimizer with a learning rate of 1×10 -4 to optimize the network. The model is trained for a total of 90,000 epochs, and finally report the results of the model with the best performance on the validation set.

[0155] 4. After the model training is completed, semantic segmentation testing is carried out. The preprocessed test images are input into the trained model, and the model output is the segmentation result. Among them, the model using the standard cross-attention adapter has an average intersection over union (mIoU) of 43.56% and an average pixel segmentation accuracy (mAccuracy) of 55.27% on the ADE20K dataset. The average intersection over union (mIoU) of the CityScapes dataset is 75.63%, and the average pixel segmentation accuracy (mAccuracy) is 84.12%. The intersection over union (IoU) of the drivable area segmentation task is 98.13%, and the pixel segmentation accuracy (Accuracy) is 98.97%. For other parameter efficient fine-tuning methods such as the linear probing method, the average intersection over union (mIoU) on the ADE20K dataset is 23.34%, and the average pixel segmentation accuracy (mAccuracy) is 31.10%. The average intersection over union (mIoU) of the CityScapes dataset is 44.23%, and the average pixel segmentation accuracy (mAccuracy) is 51.28%. The intersection over union (IoU) of the drivable area segmentation task is 92.71%, and the pixel segmentation accuracy (Accuracy) is 97.76%. The bias fine-tuning method has an average intersection over union (mIoU) of 74.35% and an average pixel segmentation accuracy (mAccuracy) of 82.75% on the CityScapes dataset. The intersection over union (IoU) of the drivable area segmentation task is 97.77%, and the pixel segmentation accuracy (Accuracy) is 98.80%. The results show that the method proposed in the present invention is significantly superior to the linear probing method and the bias fine-tuning method. Compared with the linear probing method, the present invention improves the mIoU index by 20.22 and 31.4% on the ADE20K and CityScapes datasets respectively. Compared with the bias fine-tuning method, the present invention improves the mIoU index and the mAccuracy index by 1.28% and 1.37% respectively on the CityScapes dataset. In addition, the number of learnable parameters of the model is 50.28M, which is 15.61% of the 322.13M learnable parameters required by the traditional full-parameter fine-tuning method. Furthermore, the model using the enhanced cross-attention adapter has an average intersection over union (mIoU) of 76.71% and an average pixel segmentation accuracy (mAccuracy) of 85.00% on the CityScapes dataset, and the intersection over union (IoU) of the drivable area segmentation task is 98.20%, and the pixel segmentation accuracy (Accuracy) is 98.94%, which can further improve the performance of the standard cross-attention adapter.

[0156] The results show that in the fine-tuning process of the image semantic segmentation model in this embodiment, on the one hand, the parameters of the general model itself are fixed, and only a part of the learnable parameters are introduced. The number of learnable parameters is relatively low, reducing the overhead during the fine-tuning of the general model for the image segmentation task. On the other hand, an adapter module based on cross-attention is incorporated, which can effectively improve the performance of the fine-tuning of the general model for image semantic segmentation, thereby enhancing the effect of image semantic segmentation.

[0157] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0158] The above are some specific implementation manners of the fine-tuning method for the image semantic segmentation model provided by the embodiments of the present application. Based on this, the present application also provides a corresponding device. The device provided by the embodiments of the present application will be introduced from the perspective of functional modularization below.

[0159] See Figure 4 , which is a schematic structural diagram of a fine-tuning device for an image semantic segmentation model provided by an embodiment of the present application. The device includes a construction module 100, an acquisition module 200, and a fine-tuning module 300.

[0160] The construction module 100 is used to construct an image semantic segmentation model to be fine-tuned. The image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head. Among them, the image encoder is used to extract the features of the input image and output the feature map corresponding to the input image. The cross-attention adapter is used to extract the target multi-scale features of the input image. The convolutional segmentation head is used to perform a segmentation operation on the feature map and output the semantic segmentation map of the input image.

[0161] The acquisition module 200 is used to acquire a preprocessed image semantic segmentation data set.

[0162] The fine-tuning module 300 is used to fine-tune the cross-attention adapter in the image semantic segmentation model to be fine-tuned according to the image semantic segmentation data set and the loss function until the model convergence condition is reached, and a fine-tuned image semantic segmentation model is obtained.

[0163] Optionally, the image encoder is a neural network based on visual attention.

[0164] The image encoder is specifically configured to alternately use the global attention mechanism and the window attention mechanism to extract the features of the input image and output the feature map corresponding to the input image; wherein, the image encoder uses relative position encoding in each layer to enhance the relative position information of the model output feature map.

[0165] As can be seen from the above technical solutions, in the embodiment of the present application, an image semantic segmentation model to be fine-tuned is constructed. The image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head; wherein, the image encoder is used to extract the features of the input image and output the feature map corresponding to the input image; the cross-attention adapter is used to extract the target multi-scale features of the input image; the convolutional segmentation head is used to perform a segmentation operation on the feature map and output the semantic segmentation map of the input image; a preprocessed image semantic segmentation data set is obtained, and then according to the image semantic segmentation data set and the loss function, the cross-attention adapter in the image semantic segmentation model to be fine-tuned is fine-tuned until the model convergence condition is reached, and a fine-tuned image semantic segmentation model is obtained.

[0166] In the above manner, during the fine-tuning process of the image semantic segmentation model, on the one hand, the parameters of the general model itself are fixed, and only a part of the learnable parameters are introduced, reducing the overhead during the fine-tuning of the general model for the image segmentation task. On the other hand, the adapter module based on cross-attention is combined. By obtaining the multi-scale information of the input image, the performance of the fine-tuning of the general model for image semantic segmentation can be effectively improved, which helps to segment objects of different sizes, and thus improves the effect of image semantic segmentation.

[0167] The embodiment of the present application also provides an electronic device, including: a memory for storing one or more programs;

[0168] a processor; when the one or more programs are executed by the processor, the fine-tuning method of the image semantic segmentation model described in the above embodiment is implemented.

[0169] The embodiment of the present application also provides a computer storage medium, in which a program is stored, and when the program is executed by a processor, the fine-tuning method of the image semantic segmentation model described in the above embodiment is implemented.

[0170] In the embodiment of the present application, the "first", "second" (if any) in the names such as "first" and "second" are only used as name identifiers and do not represent the first and second in order.

[0171] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0172] Those skilled in the art can understand that the flowchart shown in the figure is only an example in which the implementation mode of the present application can be realized, and the scope of application of the implementation mode of the present application is not limited by any aspect of the flowchart.

[0173] In several embodiments provided by the present application, it should be understood that the disclosed methods, devices and equipment can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0174] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit exists physically alone, or two or more units can be integrated in one unit.

[0175] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.

[0176] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A fine-tuning method for an image semantic segmentation model, characterized in that, The method includes: Construct an image semantic segmentation model to be fine-tuned, where the image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head; wherein, the image encoder is used to extract features of the input image and output a feature map corresponding to the input image; the cross-attention adapter is used to extract target multi-scale features of the input image; the convolutional segmentation head is used to perform a segmentation operation on the feature map and output a semantic segmentation map of the input image; Obtain a preprocessed image semantic segmentation dataset; According to the image semantic segmentation dataset and the loss function, fine-tune the cross-attention adapter in the image semantic segmentation model to be fine-tuned until the model convergence condition is reached, and obtain a fine-tuned image semantic segmentation model.

2. The method according to claim 1, wherein The image encoder is a neural network based on visual attention; The image encoder is used to extract features of the input image and output a feature map corresponding to the input image, including: The image encoder alternately uses a global attention mechanism and a window attention mechanism to extract features of the input image and output a feature map corresponding to the input image; wherein, the image encoder uses relative position encoding in each layer to enhance the relative position information of the model output feature map.

3. The method according to claim 1, characterized in that The cross-attention adapter includes a local spatial prior module, a spatial feature injection module, and a multi-scale feature extraction module; The cross-attention adapter is used to extract multi-scale features of the input image, including: The local spatial prior module extracts multi-scale local spatial features of the input image through multiple deep convolutional neural networks; The spatial feature injection module injects the multi-scale local spatial features into the image encoder using a cross-attention mechanism; The multi-scale feature extraction module extracts multi-scale features from the image encoder using a cross-attention mechanism; The spatial feature injection module injects the multi-scale features into the image encoder.

4. The method according to claim 3, wherein The local spatial prior module extracts multi-scale local spatial features of the input image using multiple convolutional neural networks with different depths, including: The local spatial prior module extracts local spatial features of multiple scales from the input image through the multiple convolutional neural networks respectively; The local spatial prior module maps the local spatial features of multiple scales through the fully connected layers of the multiple convolutional neural networks to obtain local spatial features of multiple scales under the same feature dimension; The local spatial prior module splices the local spatial features of multiple scales under the same feature dimension to obtain the multi-scale local spatial features of the input image.

5. The method according to claim 3, characterized in that, The spatial feature injection module is a neural network based on a cross-attention mechanism; The spatial feature injection module injects the multi-scale local spatial features into the image encoder using a cross-attention mechanism, including: The spatial feature injection module uses a cross-attention mechanism to obtain an injection feature according to the multi-scale local spatial features and the output features of the i-th layer of the image encoder; where i is a positive integer; The spatial feature injection module injects the injection features into the (i + 1)-th layer of the image encoder.

6. The method according to claim 3, wherein The multi-scale feature extraction module is a neural network based on the cross-attention mechanism; The multi-scale feature extraction module extracts multi-scale features from the image encoder by using the cross-attention mechanism, including: The multi-scale feature extraction module receives the output features of the i-th layer and the spatial features of the i-th layer of the image encoder; The multi-scale feature extraction module uses the cross-attention mechanism to obtain the multi-scale features of the i-th layer according to the spatial features of the i-th layer and the output features of the i-th layer of the image encoder; The multi-scale feature extraction module uses a feed-forward neural network to perform a normalization operation on the second multi-scale features to obtain the multi-scale features of the (i + 1)-th layer.

7. The method according to claim 1, characterized in that The convolutional segmentation head is a neural network including a convolutional layer and an upsampling layer; The convolutional segmentation head is used to perform a segmentation operation on the feature map and output the semantic segmentation map of the input image, including: The convolutional segmentation head receives the low-resolution feature map output by the image encoder as input, and alternately performs convolutional operations and upsampling operations on the low-resolution feature map to obtain the semantic segmentation map of the input image.

8. The method according to claim 1, characterized in that The method further includes: Obtaining an image to be segmented; Inputting the image to be segmented into the image semantic segmentation model to obtain the segmentation result of the image to be segmented.

9. A fine-tuning device for an image semantic segmentation model, characterized in that, The device includes: A construction module for constructing an image semantic segmentation model to be fine-tuned, where the image semantic segmentation model to be fine-tuned includes a cross-attention adapter, an image encoder, and a convolutional segmentation head; wherein, the image encoder is used to extract features of an input image and output a feature map corresponding to the input image; the cross-attention adapter is used to extract the target multi-scale features of the input image; the convolutional segmentation head is used to perform a segmentation operation on the feature map and output the semantic segmentation map of the input image; An acquisition module for acquiring a preprocessed image semantic segmentation data set; A fine-tuning module for fine-tuning the cross-attention adapter in the image semantic segmentation model to be fine-tuned according to the image semantic segmentation data set and a loss function until the model convergence condition is reached, and obtaining a fine-tuned image semantic segmentation model.

10. The device according to claim 9, characterized in that, The image encoder is a neural network based on visual attention; The image encoder is specifically configured to alternately use a global attention mechanism and a window attention mechanism to extract features of the input image and output a feature map corresponding to the input image; wherein, the image encoder uses relative position encoding in each layer to enhance the relative position information of the model output feature map.