Less-annotation remote sensing image semantic segmentation method based on visual text guidance
By introducing visual text prior and high confidence feature mixing technology in remote sensing image semantic segmentation, the problem of poor segmentation performance caused by large differences within remote sensing image classes is solved, and higher segmentation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510227412.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing semantic segmentation technology under a small number of labeled images is difficult to effectively alleviate the problem of large differences in remote sensing images, resulting in poor segmentation performance.
A semantic segmentation method for semantic sensing images based on visual text guidance is proposed. The visual text model is used to extract visual text priors, and the robustness and accuracy of segmentation are improved through technologies such as high confidence feature mixing and Euclidean distance normalization loss accumulation gain.
By introducing a mixture of visual text priors and high confidence features, the problem of large differences in remote sensing images is alleviated, the performance and accuracy of semantic segmentation are improved, and the segmentation accuracy is higher than that of existing methods.
Smart Images

Figure CN120070895A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer vision, and particularly to a method and device for semantic segmentation of remotely sensed images with few annotations based on visual text guidance. Background Art
[0002] Few-shot semantic segmentation of remotely sensed images aims to achieve pixel-level discrimination of a category using a small number of annotated images. However, due to the large differences among remotely sensed targets of the same category, a small number of annotated data cannot cover all semantic information within the category, resulting in poor segmentation results. The semantic segmentation method under few annotated samples involved in the present application aims to extract generalizable visual text priors using a visual text model to alleviate the problem of intra-class differences. To further improve the robustness of guidance, high-confidence feature mixing is proposed to mix high-confidence query features with corresponding hierarchical support image encoding features, and Euclidean distance normalized discounted cumulative gain is used to improve the performance of parameter-free visual feature metrics. This technology has important applications in scenarios such as national defense technology, urban planning, and natural disaster prevention and control.
[0003] Existing semantic segmentation technologies under few annotated images can be divided into prototype-based methods and affinity-based methods. Although these methods have improved the metric performance of few-shot semantic segmentation to a certain extent, the degree of alleviating the problem of class differences is still limited. Summary of the Invention
[0004] In view of this, embodiments of the present application propose a method and device for semantic segmentation of remotely sensed images with few annotations based on visual text guidance, aiming to solve the problem that large intra-class differences in few-shot semantic segmentation affect the segmentation performance.
[0005] To achieve the above object, an embodiment of the present application provides a method for semantic segmentation of remotely sensed images with few annotations guided by visual text, including: inputting pre-acquired image data and text data into a semantic segmentation model, where the image data includes a support image, a support image label, and a query image, and the text data includes the name of the category to be segmented and the name of the background category. The semantic segmentation model includes a visual text model and a ResNet network; processing the image data using the visual feature encoder of the visual text model to obtain visual encoded features, where the visual encoded features include support image encoded features and query image encoded features, and the support image encoded features include foreground features; processing the support image encoded features and the query image encoded features respectively using a pre-constructed visual text prior decoupling model to obtain the support image visual text prior and the query image visual text prior of the support image and the query image with respect to the category to be segmented; processing the query image visual text prior using a pre-constructed high-confidence visual feature mixing model to obtain high-confidence features in the query image visual features, and determining a query prototype based on the high-confidence features and multi-level query image encoded features, and then performing weighted summation on the query prototype and the foreground features in the multi-level support image encoded features to obtain the mixed multi-level support image encoded features, where the multi-level support image encoded features and the multi-level query image encoded features are obtained by processing the query image and the support image based on the ResNet network; processing the query image visual features and the mixed multi-level support image encoded features using a preset multi-level prior calculation model to obtain a multi-level cosine affinity prior and a multi-level Euclidean distance normalized discounted cumulative gain prior; decoding the visual text prior, the multi-level cosine affinity prior, and the multi-level Euclidean distance normalized discounted cumulative gain prior using a preset multi-level prior decoding network to obtain the segmentation result of the input image.
[0006] Optionally, the background class category names in the text data are processed using the visual text model to obtain background class text encoding features, wherein the pre-constructed visual text prior decoupling model is used to process the support image encoding features and the query image encoding features respectively, and the support image visual text prior and the query image visual text prior of the support image and the query image with respect to the class to be segmented are correspondingly obtained, including: calculating the support image visual text similarity and the query image text similarity respectively according to the cosine similarities between the support image encoding features and the query image encoding features and the background class text encoding features; inputting the support image encoding features and the query image encoding features into the visual text prior decoupling model to correspondingly obtain the support image visual text prior and the query image visual text prior; wherein, the visual text prior decoupling model includes: a first 2D transposed convolution layer, a first normalization layer, a first activation layer, a second 2D transposed convolution layer, a second normalization layer, a second activation layer, a first 2D convolution layer, a third normalization layer, a third activation layer, a first max pooling layer, a second 2D convolution layer, a fourth normalization layer, a second max pooling layer, a third 2D convolution layer and a Sigmoid activation function layer connected in sequence.
[0007] Optionally, the method further includes: training the visual text prior decoupling model using the binary cross-entropy loss between the support image visual text prior and the support image label to obtain the final visual text prior decoupling model.
[0008] Optionally, the pre-constructed high-confidence visual feature mixing model is used to process the query image visual text prior, determine the high-confidence features in the query image visual features, and determine the query prototype based on the high-confidence features and the multi-level query image encoding features, and then the query prototype and the foreground features in the multi-level support image encoding features are weighted and summed to obtain the mixed multi-level support image encoding features, including: inputting the support image and the query image into the resnet network respectively to correspondingly obtain the multi-level support image encoding features and the multi-level query image encoding features; aligning the multi-level support image encoding features and the multi-level query image encoding features with the spatial resolution of the input image to obtain the aligned multi-level query image encoding features and the aligned multi-level support image encoding features; performing the following operations on each level of the aligned multi-level query image encoding features and the aligned multi-level support image encoding features to obtain the mixed multi-level support image encoding features: normalizing the query image visual text prior to obtain the normalized prior; calculating a threshold for the normalized prior to obtain a confidence level, setting the confidence level higher than a preset threshold to 1, and setting the confidence level lower than and equal to the preset threshold to 0; obtaining the query prototype according to the confidence level and the query image encoding features of the corresponding level; determining the mixed multi-level support image encoding features of the corresponding level based on the support image label, the query prototype and the support image encoding features of the corresponding level.
[0009] Optionally, determining the hierarchically mixed support image encoding features based on the query image tags, query prototypes, and support image encoding features includes: finding the target class regions of the support image encoding features at the corresponding levels according to the support image tags, and performing weighted mixing on the target class regions and the query prototypes to obtain the mixed support image encoding features at this level.
[0010] Optionally, before using the preset multi-level prior calculation model to process the query image visual features and the hierarchically mixed support image encoding features to obtain the visual relationship prior values, the method includes: after interpolating and flattening the support image tags, obtaining the corrected support image tags with the same spatial resolution as the visual features; respectively performing a Hadamard product operation on the hierarchically mixed support image encoding features and the corrected support image tags to obtain the mask features at the corresponding levels; respectively calculating the cosine similarities between the mask features at each level and the query image encoding features, and then performing mask mean calculation and dimensional transformation on the cosine similarities in sequence along the dimension of the corrected support image tags to obtain the cosine affinity priors at each level.
[0011] Optionally, before using the preset multi-level prior calculation model to process the query image visual features and the hierarchically mixed support image encoding features to obtain the visual relationship prior values, the method further includes: respectively calculating the per-pixel Euclidean distances between the hierarchically mixed support image encoding features and the query image encoding features to obtain the Euclidean distances at each level; calculating the number of foreground pixels of the corrected support image tags; determining the coordinate subsets of the hierarchically mixed support image encoding features that are closest to the Euclidean distances at each level of the query image encoding features, where the number of coordinate subsets is equal to the number of foreground pixels; performing a correlation score on the coordinate subsets, and calculating the discounted cumulative gain and the ideal discounted cumulative gain of the query image visual features as the target class according to the correlation score; obtaining the multi-level Euclidean distance normalized discounted cumulative gain prior according to the discounted cumulative gain and the ideal discounted cumulative gain of the query image visual features as the target class.
[0012] Optionally, the preset multi-level prior decoding network includes a 2D convolutional layer, a group normalization layer, and a ReLU function layer. Using the preset multi-level prior decoding network to decode the visual text prior, the multi-level cosine affinity priors, and the multi-level Euclidean distance normalized discounted cumulative gain priors to obtain the segmentation result of the input image includes: using the 2D convolutional layer, the group normalization layer, and the ReLU function layer to perform upsampling and cross-layer connection on the visual text prior, the multi-level cosine affinity priors, and the multi-level Euclidean distance normalized discounted cumulative gain priors at different levels to obtain the corresponding dual-channel query image prediction probability maps.
[0013] Optionally, the method further includes: calculating a scale-aware cross-entropy segmentation loss between the predicted probability map of the query image and the corrected support image label; obtaining a total loss function according to the sum of the binary cross-entropy loss and the scale-aware cross-entropy segmentation loss; training the resnet network and the vision-text model using the total loss function to obtain a semantic segmentation model determined by network parameters.
[0014] To achieve the above object, an embodiment of the present application further provides a few-annotated remote sensing image semantic segmentation device based on vision-text guidance, including: an input module, configured to input pre-acquired image data and text data into the semantic segmentation model, where the image data includes a support image, a support image label, and a query image, and the text data includes a class name to be segmented and a background class name, and the semantic segmentation model includes a vision-text model and a resnet network; an encoding module, configured to process the image data using a vision feature encoder of the vision-text model to obtain vision encoding features, where the vision encoding features include support image encoding features and query image encoding features, and the support image encoding features include foreground features; a first prior calculation module, configured to process the support image encoding features and the query image encoding features respectively using a pre-constructed vision-text prior decoupling model to correspondingly obtain a support image vision-text prior and a query image vision-text prior of the support image and the query image with respect to the class to be segmented; a feature mixing module, configured to process the query image vision-text prior using a pre-constructed high-confidence vision feature mixing model to obtain high-confidence features in the query image vision features, determine a query prototype based on the high-confidence features and multi-level query image encoding features, and then perform weighted summation on the query prototype and the foreground features in the multi-level support image encoding features to obtain mixed multi-level support image encoding features, where the multi-level support image encoding features and the multi-level query image encoding features are obtained by processing the query image and the support image using the resnet network; a second prior calculation module, configured to process the query image vision features and the mixed multi-level support image encoding features using a preset multi-level prior calculation model to obtain a multi-level cosine affinity prior and a multi-level Euclidean distance normalized discounted cumulative gain prior; an image segmentation module, configured to decode the vision-text prior, the multi-level cosine affinity prior, and the multi-level Euclidean distance normalized discounted cumulative gain prior using a preset multi-level prior decoding network to obtain a segmentation result of the input image.
[0015] The method and device for semantic segmentation of remotely sensed images with few annotations guided by visual text proposed in the embodiments of this application include: inputting pre-acquired image data and text data into a semantic segmentation model, where the image data includes a support image, a support image label, and a query image, and the text data includes the name of the category to be segmented and the name of the background category. The semantic segmentation model includes a visual text model and a ResNet network; processing the image data using the visual feature encoder of the visual text model to obtain visual encoded features, where the visual encoded features include a support image encoded feature and a query image encoded feature, and the support image encoded feature includes foreground features; using a pre-constructed visual text prior decoupling model to process the support image encoded feature and the query image encoded feature respectively, and correspondingly obtaining a support image visual text prior and a query image visual text prior of the support image and the query image with respect to the category to be segmented; using a pre-constructed high-confidence visual feature mixing model to process the query image visual text prior to obtain high-confidence features in the query image visual features, and determining a query prototype based on the high-confidence features and multi-level query image encoded features, and then performing weighted summation on the query prototype and the foreground features in the multi-level support image encoded features to obtain a mixed multi-level support image encoded feature, where the multi-level support image encoded feature and the multi-level query image encoded feature are obtained by processing the query image and the support image using the ResNet network; using a preset multi-level prior calculation model to process the query image visual features and the mixed multi-level support image encoded features to obtain a multi-level cosine affinity prior and a multi-level Euclidean distance normalized discounted cumulative gain prior; using a preset multi-level prior decoding network to decode the visual text prior, the multi-level cosine affinity prior, and the multi-level Euclidean distance normalized discounted cumulative gain prior to obtain the segmentation result of the input image. This application introduces visual text prior in the semantic segmentation of small samples of remotely sensed images, and uses the generality of the visual text model to alleviate the problem of large intra-class differences in remotely sensed images in small sample semantic segmentation, improving the segmentation performance. Moreover, the visual text prior designed in this application overcomes the defect that the visual text model does not have the ability of target localization. The high-confidence visual feature mixing model mixes the high-confidence part in the query features with the corresponding-level support image encoded features, improving the robustness of the measurement process of the corresponding-level support image encoded features and the query features. This application has higher segmentation accuracy compared with the existing remotely sensed small sample semantic segmentation methods and has high utilization value in the task of on-site security and protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is the flowchart of the method for semantic segmentation of remotely sensed images with few annotations guided by visual text provided in an embodiment of this application Figure 1 ;
[0017] Figure 2The process of the semantic segmentation method of remote sensing images with few annotations based on visual text guidance is provided in one embodiment of the present application. Figure 2 ;
[0018] Figure 3 It is a structural block diagram of a device for semantic segmentation of remote sensing images with few annotations based on visual text guidance provided in another embodiment of the present application. DETAILED DESCRIPTION
[0019] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. However, it will be appreciated by those skilled in the art that in the present application, many technical details are proposed in order to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can also be implemented. The division of the following embodiments is for the convenience of description, and the specific implementation of the present application should not be limited in any way, and the various embodiments can be combined and referenced with each other without contradiction.
[0020] An embodiment of the present application proposes a semantic segmentation method for remote sensing images with few annotations based on visual text guidance, which is applied to an electronic device, wherein the electronic device can be a terminal or a server. This embodiment and the following embodiments are all described by taking the server as an example. The implementation details of the semantic segmentation method for remote sensing images with few annotations based on visual text guidance proposed in this embodiment are described in detail below. The following content is only the implementation details provided for the convenience of understanding and is not necessary for the implementation of this solution.
[0021] It should be noted that this application designs a visual text prior decoupling module, using the visual feature encoder of the contrastive language-image pre-training (CLIP) model and text feature encoder To mine potential visual text priors with spatial localization capabilities.
[0022] The specific process of the semantic segmentation method of remote sensing images with few annotations based on visual text guidance proposed in this embodiment can be as follows: Figure 1 As shown, including:
[0023] S101, inputting pre-acquired image data and text data into a semantic segmentation model, wherein the image data includes supporting images, supporting image labels, and a query image, the text data includes the name of a category to be segmented and the name of a background category, and the semantic segmentation model includes a visual text model and a ResNet network;
[0024] Among them, the input of the semantic segmentation model includes the support image I s and its label M s (annotating the target area) and the query image I q (the image to be segmented), as well as the name of the category to be segmented (such as "building") text targert and the set of background class category names for prompting
[0025] S102. Process the image data using the visual feature encoder of the visual text model to obtain visual encoded features. Among them, the visual encoded features include support image encoded features and query image encoded features, and the support image encoded features include foreground features;
[0026] The output of the visual text model includes pixel-level visual encoded features and instance-level text encoded features Randomly shuffle the background class text encoded features sequentially to avoid introducing semantic-related information about the background class and causing overfitting.
[0027] S103. Use the pre-constructed visual text prior decoupling model to process the support image encoded features and the query image encoded features respectively, and correspondingly obtain the support image visual text prior and the query image visual text prior of the support image and the query image with respect to the category to be segmented;
[0028] The function of this step is to mine the visual-text correlation of the target class and generate a pixel-level prior map.
[0029] Reference Figure 1 , in the embodiment of the present application, use the visual text model to process the background class category names in the text data to obtain background class text encoded features. S103 may include the following execution process:
[0030] S1031. Calculate the support image visual text similarity and the query image text similarity respectively according to the cosine similarity between the support image encoded features and the query image encoded features and the background class text encoded features;
[0031] S1032. Input the support image encoded features and the query image encoded features into the visual text prior decoupling model, and correspondingly obtain the support image visual text prior and the query image visual text prior;
[0032] Among them, the visual text prior decoupling model includes: a first 2D transposed convolutional layer, a first normalization layer, a first activation layer, a second 2D transposed convolutional layer, a second normalization layer, a second activation layer, a first 2D convolutional layer, a third normalization layer, a third activation layer, a first max pooling layer, a second 2D convolutional layer, a fourth normalization layer, a second max pooling layer, a third 2D convolutional layer, and a Sigmoid activation function layer, which are connected in sequence.
[0033] In an embodiment of the present application, the method further includes: training the visual text prior decoupling model using the binary cross-entropy loss between the support image visual text prior and the support image label to obtain the final visual text prior decoupling model.
[0034] To ensure that the higher part of the Prior clip value reflects a high-confidence determination as the target class, otherwise it is the background, using the support image label as a constraint, calculate the binary cross-entropy loss between the support visual text prior and the label:
[0035]
[0036] where BCE is the binary cross-entropy loss.
[0037] Specifically, the processor generates a prompt text (such as "buildings in satellite images") for the class name C, encodes it as a text feature T C , randomly shuffles and concatenates the background class set {B 1 ,..., B n} to generate a confused background text feature T B , concatenates the target and background text features: T p = [T C ; T B , to obtain the prompt text feature
[0038] First, the processor calculates the visual text similarity, and the visual text similarity can be The specific calculation formula is as follows:
[0039]
[0040] where φ represents the transformation function for the feature COS(·) is the cosine similarity metric function. F vision includes multi-level support image encoded features and multi-level query image encoded features. The multi-level support image encoded features and multi-level query image encoded features are obtained by processing the query image and the support image based on the resnet network
[0041] After obtaining the visual text similarity, the processor needs to calculate the visual text prior of the target class to be segmented mined from Prob. This application designs a prior decoupling network for prior decoupling. The network is specifically composed of an upsampling part and a downsampling part. The visual text similarity Prob is mapped to obtain the visual text prior of the target class. The specific formula is as follows:
[0042]
[0043] Among them, convT, conv, and MaxPool represent 2D transposed convolution, 2D convolution, and max pooling respectively. N is instance normalization, and ReLU is the ReLU activation function. is the Sigmoid activation function, ensuring that the output value is in the range of [0, 1]. The visual text prior of the target class includes the visual text prior of the support image and the visual text prior of the query image.
[0044] Exemplarily, the upsampling path can be ConvTranspose2D → InstanceNorm → ReLU → ConvTranspose2D → InstanceNorm → ReLU; the downsampling path can be Conv2D → InstanceNorm → ReLU → MaxPool → Conv2D → InstanceNorm → MaxPool → Conv2D; the output layer can be Sigmoid activation.
[0045] S104 uses a pre-constructed high-confidence visual feature mixture model to process the visual text prior of the query image, obtains the high-confidence features in the query image visual features, and determines the query prototype based on the high-confidence features and the multi-level query image encoding features. Then, the query prototype and the foreground features in the multi-level support image encoding features are weighted and summed to obtain the mixed multi-level support image encoding features. Among them, the multi-level support image encoding features and the multi-level query image encoding features are obtained by processing the query image and the support image based on the resnet network.
[0046] The role of this step is to enhance the robustness of the corresponding level support image encoding features and alleviate the intra-class differences.
[0047] This step may include the following execution process:
[0048] S1041, input the support image and the query image into the resnet network respectively, and correspondingly obtain the multi-level support image encoding features and the multi-level query image encoding features.
[0049] S1042. Align the multi-level support image encoding features and the multi-level query image encoding features with the spatial resolution of the input image to obtain the aligned multi-level query image encoding features and the aligned multi-level support image encoding features;
[0050] S1043. Perform the following operations on each level of the aligned multi-level query image encoding features and the aligned multi-level support image encoding features to obtain the mixed multi-level support image encoding features:
[0051] S1044. Normalize the query image visual text prior to obtain the normalized prior;
[0052] S1045. Calculate the confidence level for the normalized prior, set the confidence level higher than the preset threshold to 1, and set the confidence level lower than and equal to the preset threshold to 0;
[0053] S1046. Obtain the query prototype based on the confidence level and the query image encoding features of the corresponding level;
[0054] S1047. Determine the mixed corresponding-level support image encoding features based on the support image label, the query prototype, and the corresponding-level support image encoding features.
[0055] Specifically, the high-confidence visual feature mixing model in this application uses the visual features of the high-confidence part of the query image visual text prior to perform weighted mixing with the query features to improve the robustness of the subsequent visual affinity prior. The input of the module is the visual feature The query visual text prior with the same spatial resolution as the visual feature after interpolation and the support image label M s ∈{0,1} H×W×1 .
[0056] First, perform the min-max normalization operation on :
[0057]
[0058] Calculate the threshold based on the normalized prior. The part higher than the threshold is considered to have a high confidence level and is set to 1, otherwise it is set to 0.
[0059]
[0060] Here, τ is the confidence level threshold, which is set to 0.7 in this application. Calculate the query prototype according to and F q
[0061]
[0062] Among them, MAP is the masked average pooling operation.
[0063] In the embodiments of the present application, according to the support image label, the target class region of the support image encoding feature at the corresponding level is found, and the target class region and the query prototype are weighted and mixed to obtain the mixed support image encoding feature at this level.
[0064] Specifically, the processor finds the target class region of the support image encoding feature at the corresponding level according to the support image label M s finds the target class region of the support image encoding feature at the corresponding level, and weights and mixes the features of this region and the query prototype to obtain the final multi-level support image encoding feature:
[0065] F s (i,j) = (1 - α)·F s (i,j) + α·P q if M s (i,j) = 1 (5)
[0066] α is the mixing weight, and in the present application, it is set to 0.5.
[0067] S105. Use the preset multi-level prior calculation model to process the query image visual feature and the mixed multi-level support image encoding feature to obtain the multi-level cosine affinity prior and the multi-level Euclidean distance normalized discounted cumulative gain prior;
[0068] In the embodiments of the present application, before step S105, the method further includes:
[0069] S1051. After interpolating and flattening the support image label, obtain the corrected support image label with the same spatial resolution as the visual feature;
[0070] S1052. Perform the Hadamard product operation on the mixed multi-level support image encoding feature and the corrected support image label respectively to obtain the mask feature at the corresponding level;
[0071] S1053. Calculate the cosine similarity between the mask feature at each level and the query image encoding feature respectively, and then perform mask mean calculation and dimension transformation on the cosine similarity along the dimension of the corrected support image label in turn to obtain the cosine affinity prior at each level.
[0072] Specifically, in addition to the visual text prior, the present application uses the cosine affinity prior and the Euclidean distance normalized discounted cumulative gain prior based on visual relationships, and the input is the flattened visual feature After interpolating and flattening the support image label, obtain the corrected support image label M with the same spatial resolution as the visual feature s ∈{0,1} HW×1 .
[0073] For the cosine affinity prior, first perform a Hadamard product operation on the final multi-level support image encoded features and the corrected support image labels to obtain masked features:
[0074]
[0075] where represents the Hadamard product, and then calculate the cosine similarity between them:
[0076]
[0077] S ∈ [0, 1] HW×HW , and then perform masked mean calculation on S along the support dimension and perform dimensional transformation to obtain the cosine affinity prior Prior cos ∈ [0, 1] H×W×1 .
[0078]
[0079] In the embodiments of the present application, before using the preset multi-level prior calculation model to process the query image visual features and the mixed multi-level support image encoded features to obtain the visual relationship prior value, the method further includes:
[0080] Calculate the per-pixel Euclidean distance between the mixed multi-level support image encoded features and the query image encoded features respectively to obtain the Euclidean distances at each level;
[0081] Calculate the number of foreground pixels of the corrected support image label;
[0082] Determine the coordinate subset of the mixed multi-level support image encoded features that is closest to the Euclidean distances at each level of the query image encoded features, where the number of coordinate subsets is equal to the number of foreground pixels;
[0083] Perform a correlation score on the coordinate subset, and calculate the discounted cumulative gain and the ideal discounted cumulative gain of the query image visual feature for the target class according to the correlation score;
[0084] Obtain the multi-level Euclidean distance normalized discounted cumulative gain prior according to the discounted cumulative gain and the ideal discounted cumulative gain of the query image visual feature for the target class.
[0085] Specifically, in addition to the visual text prior, the present application uses the cosine affinity prior based on visual relationship and the Euclidean distance normalized discounted cumulative gain prior, and the input is the flattened visual features After interpolating and flattening the support image label, obtain the corrected support image label M with the same spatial resolution as the visual features s ∈ {0, 1} HW×1 .
[0086] For the Euclidean distance normalized discounted cumulative gain prior, first calculate the pixel-by-pixel Euclidean distance between the support features and the query features:
[0087]
[0088] Then calculate to obtain M s The number of foreground pixels:
[0089] k = ∑M s (10)
[0090] For F q (i), obtain the coordinates of the top k support features closest to it:
[0091] indices = topk(-D(i,:))(11)
[0092] The function of topk(·) is to find the coordinates of the top k largest elements. Take the negative of D(i,:) to find the closest elements. Then, according to whether it falls within the foreground, obtain the correlation score rel = M s (indices). If it falls within the foreground, the correlation score is 1; otherwise, it is zero. Considering the influence of the sorting position, calculate the discounted cumulative gain:
[0093]
[0094] F q (i) for the ideal discounted cumulative gain of the target class is:
[0095]
[0096] It means that the top k features with the closest matches all fall within the foreground positions. Thus, calculate the normalized discounted cumulative gain:
[0097]
[0098] The closer the NDCG value is to 1, the more likely the feature F q (i) is the target class. According to this algorithm, obtain F q Euclidean distance normalized discounted cumulative gain prior Prior ndcg ∈[0,1] H×W×1 .
[0099] S106. Use the preset multi-level prior decoding network to decode the visual text prior, the multi-level cosine affinity prior, and the multi-level Euclidean distance normalized discounted cumulative gain prior to obtain the segmentation result of the input image.
[0100] In an embodiment of the present application, the preset multi-level prior decoding network includes a 2D convolutional layer, a group normalization layer, and a function layer, and step S106 may include the following execution process:
[0101] S1061, use the 2D convolutional layer, the group normalization layer, and the function layer to perform upsampling and cross-layer connection on visual text priors at different levels, multi-level cosine affinity priors, and multi-level Euclidean distance normalized discounted cumulative gain priors to obtain corresponding dual-channel query image prediction probability maps.
[0102] In an embodiment of the present application, the method further includes:
[0103] Calculate the scale-aware cross-entropy segmentation loss by calculating the query image prediction probability map and the corrected support image label;
[0104] According to the sum of the binary cross-entropy loss and the scale-aware cross-entropy segmentation loss, obtain the total loss function;
[0105] Use the total loss function to train the resnet network and the visual text model to obtain a semantic segmentation model with determined network parameters.
[0106] In this embodiment, the multi-level prior decoder is composed of 2D convolution, group normalization, and ReLU function. By performing upsampling and cross-layer connection on priors at different levels, a dual-channel query image prediction probability map is finally output. Compare the probability map with the image label M q Calculate the scale-aware cross-entropy segmentation loss:
[0107]
[0108] So the final total loss function is:
[0109]
[0110] By backpropagating the loss function, the network parameters are optimized.
[0111] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of the present application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, are all within the protection scope of this application.
[0112] The present application has the following advantages compared with other few-shot semantic segmentation technologies:
[0113] 1. The method of this application introduces visual-text prior in small-sample semantic segmentation of remote sensing images, alleviates the problem of large intra-class differences in remote sensing images in small-sample semantic segmentation by leveraging the generality of the visual-text model, and improves the segmentation performance. Moreover, the visual-text prior designed in this application overcomes the defect that the visual-text model lacks the ability of target localization.
[0114] 2. The algorithm has stronger robustness. The high-confidence visual feature mixing module designed in this application mixes the high-confidence part in the query feature with the support feature, improving the robustness of the support feature and query feature measurement process.
[0115] 3. Visual guidance is utilized more fully. The visual relationship prior designed in this application not only utilizes the cosine similarity between features, but also designs the Euclidean distance normalized discounted cumulative gain prior, improving the utilization degree of visual feature relationships.
[0116] 4. The segmentation result has higher accuracy. This application has higher segmentation accuracy compared with existing small-sample semantic segmentation methods for remote sensing, and has high utilization value in on-site security tasks.
[0117] Refer to Figure 2 , the implementation steps of the few-shot remote sensing image semantic segmentation method based on visual-text guidance provided by another embodiment of this application are as follows:
[0118] Step 1, feature extraction process.
[0119] First, input the support image, query image, and prompt text, and perform feature encoding on the image and prompt text respectively. For the image, the image encoder includes the visual encoder part of the visual-text model and the multi-level visual backbone part, and finally obtains a CLIP visual feature and a set of multi-level visual features for one image. The prompt text is encoded through the text encoder part of the visual-text model to obtain the CLIP text feature.
[0120] Step 2, visual-text prior decoupling process.
[0121] Input the CLIP visual features of the support image and query image and the CLIP text feature of the prompt text into the visual-text prior decoupling module. First, calculate the visual-text similarity of the support image and query image respectively, and then use the prior decoupling network to decouple the visual-text similarity of both to obtain the visual-text prior of the support image and query image regarding the class to be segmented. For the visual-text prior of the support image, calculate the binary cross-entropy with its true label to obtain the visual-text prior decoupling loss.
[0122] Step 3, high-confidence visual feature mixing process.
[0123] According to the query image visual text prior, find the high-confidence part in the query visual features, calculate the query prototype using masked mean pooling, and then perform weighted summation of the query prototype and the foreground part of the support visual features to obtain the mixed support visual features.
[0124] Step 4, the calculation process of visual relationship prior.
[0125] Perform visual relationship prior calculation on the query visual features and the mixed visual text features. It includes the cosine affinity prior based on cosine similarity and the Euclidean distance normalized discounted cumulative gain prior based on Euclidean distance.
[0126] Step 5, the multi-level prior decoding process.
[0127] Perform multi-level decoding, upsampling, and cross-layer connection on the visual text prior, multi-level cosine affinity prior, and Euclidean distance normalized discounted cumulative gain prior to obtain the final segmentation result, calculate the segmentation loss with the ground truth label, and sum it with the visual text prior decoupling loss to obtain the overall loss, thereby optimizing the model.
[0128] This application also provides simulation conditions and simulation content:
[0129] This application conducts simulations using Pytorch on a central processing unit of Intel(R) Xeon(R) Silver 4210R CPU @ 2.40GHz, with 128G of memory and a Linux operating system. The data used in the simulations is a publicly available dataset.
[0130] The simulation data is the remote sensing image dataset iSAID-5i. The dataset contains 15 categories, with an image size of 256×256 pixels. 10 of the categories are used for training, and the remaining 5 categories are used for validation. There are a total of three partitioning methods: fold0, fold1, and fold2. This application conducts simulations under two different conditions: 1-shot and 5-shot, that is, using 1 support image and using 5 support images, to verify the superiority of the performance of this application. The visual backbone is Resnet50 pre-trained on ImageNet, and the visual text model is RemoteCLIP-ViT-L-14.
[0131] To prove the effectiveness of the method, this application is compared with other advanced few-shot semantic segmentation methods. The evaluation metric is mIoU, and the comparison results are shown in Table 1:
[0132] Table 1
[0133]
[0134] As can be seen from Table 1, this application is superior to other methods in terms of the mIoU metric.
[0135] Reference Figure 3 , based on the above embodiments, this application also proposes a semantic segmentation device for remotely sensed images guided by visual text with few annotations. The semantic segmentation device 1000 for remotely sensed images includes:
[0136] An input module 1001 is configured to input pre-acquired image data and text data into a semantic segmentation model. Among them, the image data includes a support image, a support image label, and a query image, and the text data includes a category name to be segmented and a background category name. The semantic segmentation model includes a visual text model and a ResNet network;
[0137] An encoding module 1002 is configured to process the image data using a visual feature encoder of the visual text model to obtain visual encoding features. Among them, the visual encoding features include support image encoding features and query image encoding features, and the support image encoding features include foreground features;
[0138] A first prior calculation module 1003 is configured to respectively process the support image encoding features and the query image encoding features using a pre-constructed visual text prior decoupling model to correspondingly obtain a support image visual text prior and a query image visual text prior of the support image and the query image with respect to the category to be segmented;
[0139] A feature mixing module 1004 is configured to process the query image visual text prior using a pre-constructed high-confidence visual feature mixing model to obtain high-confidence features in the query image visual features, and determine a query prototype based on the high-confidence features and multi-level query image encoding features, and then perform weighted summation on the query prototype and the foreground features in the multi-level support image encoding features to obtain mixed multi-level support image encoding features. Among them, the multi-level support image encoding features and the multi-level query image encoding features are obtained by processing the query image and the support image based on the ResNet network;
[0140] A second prior calculation module 1005 is configured to process the query image visual features and the mixed multi-level support image encoding features using a preset multi-level prior calculation model to obtain a multi-level cosine affinity prior and a multi-level Euclidean distance normalized discounted cumulative gain prior;
[0141] An image segmentation module 1006 is configured to decode the visual text prior, the multi-level cosine affinity prior, and the multi-level Euclidean distance normalized discounted cumulative gain prior using a preset multi-level prior decoding network to obtain a segmentation result of the input image.
[0142] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.
[0143] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0144] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and details without departing from the spirit and scope of this application.
Claims
1. A semantic segmentation method for remote sensing images with few annotations based on visual text guidance, characterized in that: include: Inputting pre-acquired image data and text data into the semantic segmentation model, wherein the image data includes supporting images, supporting image labels, and query images, the text data includes the names of the categories to be segmented and the names of the background categories, and the semantic segmentation model includes a visual text model and a resnet network; Processing the image data using a visual feature encoder of a visual text model to obtain visual coding features, wherein the visual coding features include supporting image coding features and query image coding features, and the supporting image coding features include foreground features; The pre-built visual text prior decoupling model is used to process the encoding features of the support image and the encoding features of the query image respectively, and the visual text prior of the support image and the query image about the category to be segmented are obtained accordingly; The pre-built high-confidence visual feature hybrid model is used to process the visual text prior of the query image to obtain the high-confidence features in the visual features of the query image, and the query prototype is determined based on the high-confidence features and the multi-level query image coding features, and then the query prototype and the foreground features in the multi-level support image coding features are weighted summed to obtain the mixed multi-level support image coding features, wherein the multi-level support image coding features and the multi-level query image coding features are obtained based on the resnet network to process the query image and the support image; The preset multi-level prior calculation model is used to process the query image visual features and the mixed multi-level support image coding features to obtain the multi-level cosine affinity prior and the multi-level Euclidean distance normalized loss cumulative gain prior; The visual text prior, the multi-level cosine affinity prior and the multi-level Euclidean distance normalized discounted cumulative gain prior are decoded using a preset multi-level prior decoding network to obtain a segmentation result of the input image.
2. The method for semantic segmentation of remote sensing images with few annotations based on visual text guidance according to claim 1, characterized in that: The visual text model is used to process the background category names in the text data to obtain the background category text encoding features, characterized in that the pre-built visual text prior decoupling model is used to process the support image encoding features and the query image encoding features respectively, and the support image visual text prior and the query image visual text prior of the support image and the query image with respect to the category to be segmented are obtained correspondingly, including: Calculate the support image visual text similarity and the query image text similarity according to the cosine similarity between the support image encoding feature and the query image encoding feature and the background text encoding feature; Input the support image encoding features and the query image encoding features into the visual text prior decoupling model to obtain the support image visual text prior and the query image visual text prior respectively; Among them, the visual text prior decoupling model includes: a first 2D transposed convolution layer, a first normalization layer, a first activation layer, a second 2D transposed convolution layer, a second normalization layer, a second activation layer, a first 2D convolution layer, a third normalization layer, a third activation layer, a first maximum pooling layer, a second 2D convolution layer, a fourth normalization layer, a second maximum pooling layer, a third 2D convolution layer and a Sigmoid activation function layer connected in sequence.
3. The method for semantic segmentation of remote sensing images with few annotations based on visual text guidance according to claim 2, characterized in that: The method further comprises: The visual-text prior decoupling model is trained using the binary cross entropy loss between the supporting image visual-text prior and the supporting image label to obtain the final visual-text prior decoupling model.
4. The method for semantic segmentation of remote sensing images with few annotations based on visual text guidance according to claim 1, characterized in that: The method uses a pre-built high-confidence visual feature hybrid model to process the query image visual text prior, determines the high-confidence features in the query image visual features, and determines the query prototype based on the high-confidence features and the multi-level query image coding features, and then performs weighted summation on the query prototype and the foreground features in the multi-level support image coding features to obtain the mixed multi-level support image coding features, including: The support image and the query image are input into the resnet network respectively, and the multi-level support image encoding features and the multi-level query image encoding features are obtained correspondingly; Aligning the multi-level support image coding features and the multi-level query image coding features with the spatial resolution of the input image to obtain aligned multi-level query image coding features and aligned multi-level support image coding features; The following operations are performed on each level of the aligned multi-level query image coding features and the aligned multi-level support image coding features to obtain a mixed multi-level support image coding feature: Normalize the query image visual text prior to obtain a normalized prior; Perform threshold calculation on the normalized prior to obtain confidence, set the confidence higher than the preset threshold to 1, and set the confidence lower than or equal to the preset threshold to 0; According to the confidence and the query image encoding features of the corresponding level, the query prototype is obtained; The mixed corresponding-level support image coding features are determined based on the support image label, the query prototype and the corresponding-level support image coding features.
5. The method for semantic segmentation of remote sensing images with few annotations based on visual text guidance according to claim 4, characterized in that: The step of determining the mixed hierarchical supporting image coding features based on the query image label, the query prototype and the supporting image coding features includes: According to the supporting image label, the target class area of the supporting image coding feature of the corresponding level is found, and the target class area and the query prototype are weighted mixed to obtain the mixed supporting image coding feature of the level.
6. The method for semantic segmentation of remote sensing images with few annotations based on visual text guidance according to claim 4, characterized in that: Before using a preset multi-level prior calculation model to process the query image visual features and the mixed multi-level support image coding features to obtain a multi-level cosine affinity prior and a multi-level Euclidean distance normalized loss cumulative gain prior, the method includes: After interpolating and flattening the support image labels, a modified support image label with a spatial resolution consistent with the visual features is obtained; Performing Hadamard product operations on the mixed multi-level support image encoding features and the corrected support image labels respectively to obtain mask features of the corresponding levels; The cosine similarity of the mask features of each level and the encoding features of the query image is calculated respectively, and then the mask mean calculation and dimension transformation are performed on the cosine similarity along the dimension of the corrected support image label to obtain the cosine affinity prior of each level.
7. The method for semantic segmentation of remote sensing images with few annotations based on visual text guidance according to claim 6, characterized in that: Before using the preset multi-level prior calculation model to process the query image visual features and the mixed multi-level support image coding features to obtain the multi-level cosine affinity prior and the multi-level Euclidean distance normalized loss cumulative gain prior, the method further includes: Calculate the pixel-by-pixel Euclidean distance between the mixed multi-level support image coding features and the query image coding features to obtain the Euclidean distance of each level; Calculate the number of foreground pixels of the corrected support image labels; Determine a coordinate subset of the mixed multi-level support image coding feature that has the closest Euclidean distance to each level of the query image coding feature, wherein the number of coordinate subsets is equal to the number of foreground pixels; Scoring the relevance of the coordinate subsets, and calculating the loss cumulative gain and the ideal loss cumulative gain of the query image visual features for the target class based on the relevance score; According to the ideal discounted cumulative gain of the target class based on the discounted cumulative gain and the visual features of the query image, a multi-level Euclidean distance normalized discounted cumulative gain prior is obtained.
8. The method for semantic segmentation of remote sensing images with few annotations based on visual text guidance according to claim 3, characterized in that: The preset multi-level prior decoding network includes a 2D convolution layer, a group normalization layer and a ReLU function layer. The preset multi-level prior decoding network is used to decode the visual text prior, the multi-level cosine affinity prior and the multi-level Euclidean distance normalized discounted cumulative gain prior to obtain the segmentation result of the input image, including: A multi-level prior decoding network is used to upsample and cross-layer connect different levels of visual text priors, multi-level cosine affinity priors and multi-level Euclidean distance normalized discounted cumulative gain priors to obtain the corresponding dual-channel query image prediction probability map.
9. The method for semantic segmentation of remote sensing images with few annotations based on visual text guidance according to claim 8, characterized in that: The method further comprises: Compute the query image prediction probability map and the corrected support image labels to calculate the scale-aware cross entropy segmentation loss; The total loss function is obtained by summing the binary cross entropy loss and the scale-aware cross entropy segmentation loss. The total loss function is used to train the ResNet network and the visual text model to obtain a semantic segmentation model determined by the network parameters.
10. A device for semantic segmentation of remote sensing images with few annotations based on visual text guidance, characterized in that: include: An input module, used to input pre-acquired image data and text data into the semantic segmentation model, wherein the image data includes supporting images, supporting image labels, and query images, the text data includes the names of the categories to be segmented and the names of the background categories, and the semantic segmentation model includes a visual text model and a resnet network; An encoding module, used for processing image data using a visual feature encoder of a visual text model to obtain visual encoding features, wherein the visual encoding features include supporting image encoding features and query image encoding features, and the supporting image encoding features include foreground features; A first prior calculation module is used to process the support image coding features and the query image coding features respectively by using a pre-built visual text prior decoupling model, and obtain the support image visual text prior and the query image visual text prior of the support image and the query image with respect to the category to be segmented; A feature mixing module is used to process the visual text prior of the query image using a pre-built high-confidence visual feature mixing model to obtain high-confidence features in the visual features of the query image, and determine the query prototype based on the high-confidence features and the multi-level query image coding features, and then perform weighted summation on the query prototype and the foreground features in the multi-level support image coding features to obtain a mixed multi-level support image coding feature, wherein the multi-level support image coding features and the multi-level query image coding features are obtained by processing the query image and the support image based on the ResNet network; The second prior calculation module is used to process the query image visual features and the mixed multi-level support image coding features using a preset multi-level prior calculation model to obtain a multi-level cosine affinity prior and a multi-level Euclidean distance normalized loss cumulative gain prior; The image segmentation module is used to decode the visual text prior, the multi-level cosine affinity prior and the multi-level Euclidean distance normalized discounted cumulative gain prior using a preset multi-level prior decoding network to obtain a segmentation result of the input image.
Citation Information
Patent Citations
Novel image generation method jointly driven by text and semantic segmentation map
CN117557683A
Few-sample remote sensing image semantic segmentation method based on double-branch reinforcement network
CN119516186A
Panoptic segmentation with multi-dataset training and part-whole awareness
US20240378874A1
System and method for model-free, one-shot object pose estimation via coordinate regression
US20240404104A1
Cited By
Small sample industrial defect image segmentation method and system based on text background perception enhancement, and storage medium
CN120580259A
Small sample industrial defect image segmentation method and system based on text background perception enhancement and storage medium
CN120580259B