A Semi-Supervised Semantic Segmentation Method Based on Scale-Aware Attention

By introducing scale-aware attention module and semi-supervised learning method in semantic segmentation tasks, the problems of scale inconsistency and insufficient labeling in multi-scale target segmentation are solved, and better multi-scale target segmentation effect and model performance improvement are achieved.

CN115661463BActive Publication Date: 2025-06-20NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211432720.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2025-06-20
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problem of scale inconsistency in multi-scale target semantic segmentation tasks, and insufficient pixel-level annotation pictures, resulting in increased difficulty in model training and poor segmentation effect.

Method used

The semi-supervised semantic segmentation method based on scale perceived attention is adopted to dynamically allocate the weights of different scale features through the scale attention module, and use images without pixel-level labels to generate pseudo-labels, further improving the effect of multi-scale target semantic segmentation.

Benefits of technology

By dynamically adjusting the importance of output features in scale, the problem of scale inconsistency in semantic segmentation tasks is alleviated, the pseudo-label quality and multi-scale target segmentation capabilities of the model are improved, and the performance on multiple semantic segmentation test sets is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661463B_ABST
    Figure CN115661463B_ABST
Patent Text Reader

Abstract

The present invention provides a semi-supervised semantic segmentation method based on scale-aware attention, comprising the following steps: obtaining a multi-object semantic segmentation dataset, preprocessing the data with pixel-level labels and the data without pixel-level labels, and training a scale-aware attention network using the training set divided from the data with pixel-level labels; obtaining the scale importance of each image in the training set and the scale distribution of the target segmentation results, and training a confidence prediction model; inputting the data without pixel-level labels into the scale-aware attention network to output the scale importance and pseudo-labels; using the trained confidence prediction model to predict the confidence of the pseudo-labels; screening the pseudo-labels according to the predicted confidence, expanding the training set and retraining the scale-aware attention network. The present invention dynamically allocates weights to different-scale features by using a scale attention module, and further improves the effect of multi-scale object semantic segmentation by using images without pixel-level labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a semi-supervised semantic segmentation method based on scale-aware attention. Background Art

[0002] With the success of deep learning in various computer vision applications in recent years, deep learning has been widely applied to multi-scale object semantic segmentation tasks. A deep learning model designed for one task can be easily transferred to another task by training the model with new data. On the one hand, the target shapes are irregular and are interfered by complex backgrounds. Even for the same type of target, their sizes may vary greatly, which increases the difficulty of object segmentation based on models and manual operations. On the other hand, pixel-level annotation of targets requires a large amount of work, which greatly increases the difficulty of collecting a large number of finely annotated samples to meet the requirements of model training. These characteristics and difficulties lead to poor multi-scale object semantic segmentation effects.

[0003] FCN was the first to successfully apply the deep learning method to the semantic segmentation task. This fully convolutional network structure provides an end-to-end, pixel-to-pixel solution for semantic segmentation. An important contribution of it is that it replaces the last few fully connected layers of traditional CNNs with convolutional layers. And upsampling is introduced between the last few convolutional layers, so that an output with the same size as the input image can be obtained.

[0004] The encoder-decoder structure is another widely used deep segmentation architecture. The encoder part of U-Net is similar to the traditional CNN network, using 3×3 convolutional kernels and 2×2 max pooling for feature extraction. The decoder upsamples the deep features by a factor of 2 and then fuses them with the features of the same layer on the left side, and then uses 3×3 convolution for learning. Subsequently, the operations of upsampling, feature fusion, and using 3×3 convolution for learning are repeated until the resolution is the same as that of the input image. Finally, 1×1 convolution is used to obtain the final result. Different from FCN, the four upsampling operations in U-Net are all upsampling by a factor of 2, gradually restoring the resolution and avoiding the too sudden change of resolution caused by using a large upsampling factor in FCN. Skip connections are also introduced in U-Net to fuse feature maps of the same resolution, and this way of layer-by-layer multiple fusions enables high-level features and low-level features to be combined in a better way. The feature fusion method of U-Net is also different from that of FCN. It concatenates the features to be fused in the channel dimension instead of adding the corresponding pixel values. Due to the characteristics of medical images such as small data volume, fixed structure, and multi-modal, the advantages of the U-Net structure are fully demonstrated. U-Net has become a benchmark network in many medical image segmentation fields.

[0005] In recent years, the attention mechanism has demonstrated great advantages in the field of deep learning. Existing attention mechanisms include channel attention mechanisms, spatial attention mechanisms, and combinations of the two. The SE block is a typical channel attention mechanism, STN and Non-local are relatively commonly used spatial attention mechanisms, and CBAM combines the advantages of channel attention mechanisms and spatial attention mechanisms. However, current attention mechanisms cannot address the problem of inconsistent scales in multi-object segmentation tasks. Summary of the Invention

[0006] Aiming at the technical problems of inconsistent scales and insufficient pixel-level labeled images in the prior art, the present invention provides a semi-supervised semantic segmentation method based on scale-aware attention, which dynamically assigns weights to different-scale features using a scale attention module and further improves the effect of multi-scale object semantic segmentation using images without pixel-level labels.

[0007] The technical solution adopted by the present invention is as follows: A semi-supervised semantic segmentation method based on scale-aware attention, comprising the following steps:

[0008] Step 1: Obtain a multi-object semantic segmentation dataset, preprocess the data with pixel-level labels and the data without pixel-level labels, divide the preprocessed data with pixel-level labels into a training set and a test set, and augment the images in the training set;

[0009] Step 2: Use the augmented training set to train a scale-aware attention network, and under the guidance of the cross-entropy loss function, output a segmentation result map. The scale-aware attention network includes an encoder-decoder structure, a multi-scale feature connection module, and a scale attention module connected in sequence;

[0010] Step 3: Use the trained scale-aware attention network in Step 2 to obtain the scale importance of each image in the training set and the scale distribution of the target segmentation results, and train a confidence prediction model. The confidence prediction model outputs the confidence of the label under the guidance of the mean square error loss function;

[0011] Step 4: Input the preprocessed data without pixel-level labels into the trained scale-aware attention network in Step 2, and output the scale importance of each unlabeled image and the corresponding pseudo-label;

[0012] Step 5: Statistically analyze the scale distribution of the target segmentation results of the pseudo-labels of the unlabeled images, and input them together with the scale importance of the unlabeled images into the confidence prediction model to predict the confidence of the pseudo-labels of each unlabeled image;

[0013] Step 6: Screen the pseudo-labels according to the predicted confidence levels, and divide them into reliable pseudo-labels and unreliable pseudo-labels. The unlabeled images corresponding to the reliable pseudo-labels form a reliable pseudo-label image set, and the unlabeled images corresponding to the unreliable pseudo-labels form an unreliable pseudo-label image set;

[0014] Use the reliable pseudo-label image set to expand the training set, and retrain the well-trained scale-aware attention network in Step 2;

[0015] The unreliable pseudo-label image set is input into the retrained scale-aware attention network to generate corresponding pseudo-labels, forming a reliable pseudo-label data set;

[0016] Step 7: Use the reliable pseudo-label data set to expand the training set again, and retrain the retrained scale-aware attention network in Step 6 to obtain the final scale-aware attention network.

[0017] Further, in Step 1, the expansion process includes: rotating the pictures, and the rotation angles include 90°, 180°, and 270°, and saving the rotated pictures; flipping the pictures vertically and horizontally, and saving the flipped pictures.

[0018] Further, in Step 2, the encoder part of the encoder-decoder structure uses the backbone network EfficientNet-B0.

[0019] Further, in Step 2, the multi-scale feature connection module regards the features from different levels of the decoder as features of multiple scales and connects them to form multi-scale features. The multi-scale feature connection module uses 1×1 convolution and corresponding upsampling multiples to align the features from different scales in both the spatial and channel dimensions.

[0020] Further, in Step 2, the scale attention module includes a spatial path and a channel path, dynamically adjusts the input multi-scale features in both the spatial and channel dimensions, and assigns different weights to features of different scales.

[0021] Further, Step 2.1: The operations of the spatial path of the scale attention module include:

[0022] Step 2.11: Given the input feature F containing S scales in , for the features of each scale perform the in-channel averaging operation:

[0023]

[0024] where is The feature of the j-th channel, where C is the total number of channels, is the feature map of the i-th scale;

[0025] Step 2.12: After concatenation in the channel dimension, it is input into a 7×7 convolutional layer, and then the convolutional result is input into the sigmoid function to obtain the weight map F of the spatial path spa_att :

[0026]

[0027] where, [] represents the operation of concatenating features, and conv 7×7 is the 7×7 convolutional operation, and σ represents the sigmoid function;

[0028] Step 2.2: The operations of the channel path of the scale attention module include:

[0029] Step 2.21: Perform global average pooling (GAP) operation on the given input feature F containing S scales in to compress the features and obtain the global feature of the j-th channel corresponding to the i-th scale

[0030]

[0031] where, represents the parameter value at the position (h, w) of the j-th channel and the i-th scale of the input feature F in , and H and W represent the height and width of the image respectively;

[0032] Step 2.22: Perform an operation of taking the average within the channel once:

[0033]

[0034] The obtained is the result of taking the average of the global features of the i-th scale;

[0035] Step 2.23: Use the MLP layer to learn the scale-specific channel importance, and then input it into the sigmoid activation function to obtain the weight map F of the channel dimension cha_att :

[0036]

[0037] where, mlp represents the MLP layer;

[0038] Step 2.3: Calculate the per-scale product using the results of the two paths to obtain the importance of each scale i

[0039]

[0040] Step 2.4: Importance of each scale and the input features at each scale are multiplied element-wise to obtain the activation features at each scale

[0041]

[0042] Furthermore, in Step 2, the cross-entropy loss function is as follows:

[0043]

[0044] where y represents the label, represents the prediction result, and k represents the number of classes.

[0045] Furthermore, in Step 3, the confidence prediction model only contains two layers of convolution.

[0046] Furthermore, in Step 3, the mean squared error loss function is:

[0047]

[0048] where y represents the label value, represents the predicted value.

[0049] Furthermore, in Step 6, the top 50% of the pseudo-labels with higher confidence are selected as reliable pseudo-labels, and the remaining pseudo-labels are used as unreliable pseudo-labels.

[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0051] 1. The present invention utilizes the scale importance in the scale attention module and the scale distribution of the targets in the pseudo-labels to improve the quality of the generated pseudo-labels, and reliable pseudo-labels for unlabeled data can be obtained only through two retrainings; the unlabeled images with pseudo-labels and the original labeled data are combined, and the scale-aware attention network is retrained, which can effectively improve the model's ability to simultaneously segment multi-scale targets.

[0052] 2. The scale attention module in the present invention can dynamically adjust the importance of the output features in terms of scale, assign different weights to features of different scales, and can alleviate the scale inconsistency problem existing in the semantic segmentation task.

[0053] 3. The present invention has achieved good performance improvements on multiple semantic segmentation test sets and has good universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Flow chart of an embodiment of the present invention;

[0055] Figure 2 Structural diagram of the scale-aware attention network of an embodiment of the present invention;

[0056] Figure 3 Structural diagram of the scale attention module of an embodiment of the present invention;

[0057] Figure 4 Comparison chart of multi-scale object segmentation effects between an embodiment of the present invention and a baseline model on the IDRiD dataset. Detailed implementation manners

[0058] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] An embodiment of the present invention provides a semi-supervised semantic segmentation method based on scale-aware attention. As Figure 1 shown, it includes the following steps:

[0060] Step 1: Obtain a multi-object semantic segmentation dataset, such as the IDRiD dataset and the DDR dataset, and preprocess the data with pixel-level labels and the data without pixel-level labels. Specifically, each image in the dataset is processed to a certain size, such as 1440×960 pixels. This size can retain most of the image information and will not cause insufficient video memory to support model inference due to the overly large image size.

[0061] Divide the preprocessed data with pixel-level labels into a training set and a test set, and augment the images in the training set. The augmentation process includes: rotating the images in the training set by angles including 90°, 180°, and 270°, and saving the rotated images; flipping the images in the training set vertically and horizontally, and saving the flipped images; using the images in the training set and the training set images processed as above together as the training set D l , and complete the data augmentation.

[0062] Step 2: Use the training set D l to train the scale-aware attention network Net f , and obtain a segmentation result map under the guidance of the cross-entropy loss function. As Figure 2 shown, the scale-aware attention network Net f includes an encoder-decoder structure, a multi-scale feature connection module, and a scale attention module connected in sequence. The encoder-decoder structure is used to extract features from the image, the multi-scale feature connection module connects features from different scales, and the scale attention module efficiently fuses and learns the obtained multi-scale features.

[0063] The encoder part of the encoder-decoder structure uses the backbone network EfficientNet-B0 with better feature learning ability to replace the traditional U-Net.

[0064] Regarding the features from different levels of the decoder as features of multiple scales, in the multi-scale feature connection module, first, for other levels except the shallowest one, use 1×1×16 convolution to scale the number of channels of their features to 16, and then perform an upsampling operation to increase the features with lower resolution to the same resolution as the input image, achieving the effect of aligning the features from different scales in both the spatial and channel dimensions.

[0065] Multi-scale features, that is, the input feature F containing S scales in are input into the scale attention module, as Figure 3 shown. The scale attention module contains a spatial path and a channel path, operates on the features in the scale dimension, dynamically adjusts the multi-scale input features, obtains the importance of the output features in terms of scale, and assigns different weights to the features of different scales to solve the problem of scale inconsistency existing in multi-object semantic segmentation.

[0066] Among them, the operations of the spatial path of the scale attention module include:

[0067] Step 2.11: Given the input feature F containing S scales in , for the features of each scale perform an average operation within the channel:

[0068]

[0069] where is the feature of the j-th channel of , C is the total number of channels, and

[0070] is the feature map of the i-th scale; After concatenation in the channel dimension, input it into a 7×7 convolutional layer, and then input the convolutional result into the sigmoid function to map the weight to the range [0,1], obtaining the weight map F spa_att :

[0071]

[0072] where [] represents the operation of concatenating features, conv 7×7 is the 7×7 convolutional operation, and σ represents the sigmoid function.

[0073] The operations of the channel path of the scale attention module include:

[0074] Step 2.21: Perform global average pooling operation on the given input feature F containing S scales in to compress the feature and obtain the global feature corresponding to the j-th channel of the i-th scale

[0075]

[0076] where, represents the parameter value at the position (h, w) of the j-th channel of the i-th scale of the input feature F in , and H and W represent the height and width of the image respectively;

[0077] Step 2.22: Perform an in-channel averaging operation once:

[0078]

[0079] The obtained is the result of averaging the global features of the i-th scale;

[0080] Step 2.23: Use the MLP layer to learn scale-specific channel importance, and then input it into the sigmoid activation function to map the weights to the range [0, 1], obtaining the weight map F cha_att :

[0081]

[0082] where, mlp represents the MLP layer.

[0083] Step 2.3: Calculate the per-scale product using the results of the spatial path and the channel path to obtain the importance of each scale i

[0084]

[0085] Step 2.4: Element-wise multiply the importance of each scale and the input feature at each scale to obtain the activated feature at each scale

[0086]

[0087] The cross-entropy loss function is as follows:

[0088]

[0089] where, y represents the label, Represents the prediction result, and k represents the number of categories.

[0090] Step 3: Use the trained scale-aware attention network in Step 2 to obtain the scale importance of each image in the training set D l and statistically analyze the scale distribution of the target segmentation results of the labels in each image. Train a confidence prediction model using the scale importance and scale distribution. The confidence prediction model only contains two layers of convolution. The confidence prediction model outputs the confidence of the label under the guidance of the mean square error loss function.

[0091] The mean square error loss function is:

[0092]

[0093] where y represents the label value, represents the predicted value.

[0094] Step 4: Input the preprocessed pixel-level label-free data into the trained scale-aware attention network Net f to obtain the scale importance and corresponding pseudo-labels of each unlabeled image.

[0095] Step 5: Statistically analyze the scale distribution of the target segmentation results of the pseudo-labels of the unlabeled images, and input them together with the scale importance of the unlabeled images into the trained confidence prediction model to predict the confidence of the pseudo-labels of each unlabeled image.

[0096] Step 6: Screen the pseudo-labels according to the predicted confidence, select the top 50% of the pseudo-labels with higher confidence as reliable pseudo-labels, and the corresponding unlabeled images form a reliable pseudo-label image set D u1 ; the 50% of the pseudo-labels with lower confidence are used as unreliable pseudo-labels, and the corresponding unlabeled images form an unreliable pseudo-label image set

[0097] Merge the reliable pseudo-label image set D u1 with the training set D l to form a new training set (D u1 ∪D l ), input the images in it into the trained scale-aware attention network Net f for training to obtain a retrained scale-aware attention network Net u .

[0098] Input the images in the unreliable pseudo-label image set into the retrained scale-aware attention network Net u to generate corresponding pseudo-labels, forming a reliable pseudo-label data set D u2 .

[0099] Step 7: Merge the reliable pseudo-labeled dataset D u2 into the new training set (D u1 ∪D l ), to obtain the final training set (D u1 ∪D u2 ∪D l ), and retrain the retrained scale-aware attention network Net u to obtain the final scale-aware attention network Net.

[0100] Step 8: After the model training is completed, perform multi-scale object semantic segmentation tasks on the test set, and calculate AUPR and mAUPR to prove the effectiveness of the model.

[0101] Among them, AUPR is the area under the precision-recall curve (Area Under Precision Recall curve). To introduce AUPR, first give the definitions of precision and recall:

[0102]

[0103]

[0104] If the vertical axis is set to Precision and the horizontal axis is set to Recall, and all (recall, precision) matching pairs are plotted in the graph, the PR curve is obtained. The area of the region enclosed by the PR curve and the horizontal and vertical coordinate axes is AUPR. The average AUPR (mAUPR) is used to evaluate the overall performance of multi-object semantic segmentation. It is obtained by calculating the average value of the AUPR of all n types of objects, and it can be expressed as follows:

[0105]

[0106] As Figure 4 shown, it is a comparison chart of the multi-scale object segmentation effects of the embodiment of the present invention and the baseline model on the IDRiD dataset. From left to right in each row are: the original image, the label, the prediction of EfficientNet-B0, the prediction of the scale-aware attention network, and the prediction of the semi-supervised semantic segmentation method based on scale-aware attention. The comparison areas are marked with boxes, and EX, HE, SE, and MA are the objects of 4 categories in the dataset.

[0107] The existing advanced methods, including VRT, iFLYTEK-MIG, U-Net, U-Net++, DeepLabv3+, PSPNet, HED, FCRN, CASENet, L-seg, Scale-Aware Attention Network, and the semi-supervised semantic segmentation method based on scale-aware attention of the present invention, are compared on the IDRiD dataset, and the results are shown in Table 1.

[0108] Table 1 Comparison results of the present method and existing multi-object semantic segmentation methods on the IDRiD test set

[0109]

[0110] The existing multi-scale object semantic segmentation methods, including U-Net, U-Net++, DeepLabv3+, PSPNet, HED, FCRN, CASENet, L-seg, Scale-Aware Attention Network, and the semi-supervised semantic segmentation method based on scale-aware attention of the present invention, are compared on the DDR dataset, and the results are shown in Table 2.

[0111] Table 2 Comparison results of the present method and existing multi-object semantic segmentation methods on the DDR test set

[0112]

[0113] The present invention has been described in detail through the embodiments above. However, the content described above is only an exemplary embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. The protection scope of the present invention is defined by the claims. All those who utilize the technical solutions of the present invention, or those skilled in the art who, inspired by the technical solutions of the present invention, design similar technical solutions within the essence and protection scope of the present invention to achieve the above technical effects, or make equivalent changes and improvements to the scope of the application, etc., should still fall within the patent coverage protection scope of the present invention.

Claims

1. A semi-supervised semantic segmentation method based on scale-aware attention, characterized in that: It includes the following steps: Step 1: Obtain a multi-object semantic segmentation dataset, preprocess the data with pixel-level labels and the data without pixel-level labels, divide the preprocessed data with pixel-level labels into a training set and a test set, and augment the images in the training set; Step 2: Use the augmented training set to train a scale-aware attention network, and output a segmentation result map under the guidance of a cross-entropy loss function. The scale-aware attention network includes an encoder-decoder structure, a multi-scale feature connection module, and a scale attention module connected in sequence; The scale attention module includes a spatial path and a channel path, dynamically adjusts the input multi-scale features in both the spatial and channel dimensions, and assigns different weights to features of different scales; Step 3: Use the trained scale-aware attention network in Step 2 to obtain the scale importance of each image in the training set and the scale distribution of the target segmentation results, and train a credibility prediction model. The credibility prediction model outputs the credibility of the label under the guidance of a mean squared error loss function; Step 4: Input the preprocessed data without pixel-level labels into the trained scale-aware attention network in Step 2, and output the scale importance of each unlabeled image and the corresponding pseudo-labels; Step 5: Statistically analyze the scale distribution of the target segmentation results of the pseudo-labels of the unlabeled images, and input them together with the scale importance of the unlabeled images into the credibility prediction model to predict the credibility of the pseudo-labels of each unlabeled image; Step 6: Screen the pseudo-labels according to the predicted credibility, divide them into reliable pseudo-labels and unreliable pseudo-labels. The unlabeled images corresponding to the reliable pseudo-labels form a reliable pseudo-label image set, and the unlabeled images corresponding to the unreliable pseudo-labels form an unreliable pseudo-label image set; Use the reliable pseudo-label image set to augment the training set, and retrain the trained scale-aware attention network in Step 2; Input the unreliable pseudo-label image set into the retrained scale-aware attention network to generate corresponding pseudo-labels, forming a reliable pseudo-label dataset; Step 7: Use the reliable pseudo-label dataset to augment the training set again, and retrain the retrained scale-aware attention network in Step 6 to obtain the final scale-aware attention network.

2. The semi-supervised semantic segmentation method based on scale-aware attention according to claim 1, characterized in that: In Step 1, the augmentation process includes: rotating the images, with the rotation angles including 90°, 180°, and 270°, and saving the rotated images; flipping the images vertically and horizontally, and saving the flipped images.

3. The semi-supervised semantic segmentation method based on scale-aware attention according to claim 1, characterized in that: In Step 2, the encoder part of the encoder-decoder structure uses the backbone network EfficientNet-B0.

4. The semi-supervised semantic segmentation method based on scale-aware attention according to claim 1, characterized in that: In Step 2, the multi-scale feature connection module regards the features from different levels of the decoder as features of multiple scales and connects them to form multi-scale features. The multi-scale feature connection module uses 1×1 convolutions and corresponding upsampling multiples to align the features from different scales in both the spatial and channel dimensions.

5. The semi-supervised semantic segmentation method based on scale-aware attention according to claim 1, characterized in that: Step 2.1: The operations of the spatial path of the scale attention module include: Step 2.11: Given the input features containing S scales , for the features of each scale , perform the in-channel average operation: ; Among them, is the feature of the j th channel, C is the total number of channels, is the feature map of the i th scale; Step 2.12: After concatenation in the channel dimension, input it into a 7×7 convolutional layer, and then input the convolutional result into the sigmoid function to obtain the weight map of the spatial path : ; Among them, [] represents the concatenation operation on features, is a 7×7 convolution operation, represents the sigmoid function; Step 2.2: The operations of the channel path of the scale attention module include: Step 2.21: For the given input features containing S scales perform global average pooling operation to obtain the global features corresponding to the i th scale and the j th channel : ; Among them, represents the th i scale and the j th channel position of the input feature, at ([[]] h, w ), the parameter value, H and W respectively represent the height and width of the image;​ Step 2.22: Perform an in-channel average operation: ; Obtained is the result of averaging the global features of the i th scale; Step 2.23: Use the MLP layer to learn scale-specific channel importance and then input it into the sigmoid activation function to obtain a weight map in the channel dimension : ; Among them, mlp represents the MLP layer; Step 2.3: Calculate the per-scale product using the results of the two paths to obtain the importance at each scale i : ; Step 2.4: Importance of each scale and the input features at each scale are multiplied element-wise to obtain the activation features at each scale : 。 6. The semi-supervised semantic segmentation method based on scale-aware attention according to claim 1, wherein: In Step 2, the cross-entropy loss function is as follows: ; Among them, represents a label, represents a prediction result, k indicates the number of categories.

7. The semi-supervised semantic segmentation method based on scale-aware attention according to claim 1, wherein: In Step 3, the credibility prediction model only contains two layers of convolution.

8. The semi-supervised semantic segmentation method based on scale-aware attention according to claim 1, wherein: In Step 3, the mean squared error loss function is: ; Among them represents the tag value, represents the predicted value.

9. The semi-supervised semantic segmentation method based on scale-aware attention according to claim 1, wherein: In Step 6, select the top 50% of the pseudo-labels with higher credibility as reliable pseudo-labels, and the remaining pseudo-labels as unreliable pseudo-labels.

Citation Information

Patent Citations

  • Contour perception multi-organ segmentation network construction method based on class-by-class convolution operation

    CN112465827A

  • SAR image change detection method based on multi-scale differential feature attention mechanism

    CN114926746A