An image segmentation method based on boundary enhancement

Through the boundary-enhanced image segmentation network model, the problems of boundary information extraction and feature fusion error in the prior art are solved, and a higher image segmentation accuracy is achieved.

CN116205927BActive Publication Date: 2025-07-08XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310165505.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-07-08
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

The existing image segmentation method based on deep learning algorithms has problems of error and insufficient fusion in boundary information extraction and fusion of features at different scales, resulting in unsatisfactory segmentation results.

Method used

The boundary enhanced image segmentation network model based on the encoder-decoder framework is adopted, and boundary features are extracted and supervised through the boundary extraction module, and feature fusion is carried out in the scale, space, and channel dimensions in combination with the multi-scale attention aggregation module.

Benefits of technology

The accuracy of image segmentation is improved, and the spatial information and semantic information of the segmentation result are enhanced through accurate boundary information extraction and multi-dimensional feature fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116205927B_ABST
    Figure CN116205927B_ABST
Patent Text Reader

Abstract

The present invention discloses an image segmentation method based on boundary enhancement, comprising: establishing an image segmentation network model with boundary enhancement, and segmenting an input image by using the trained network model; wherein, the encoder of the image segmentation network model includes a first feature extraction module, a boundary extraction module and a second feature extraction module, which are mainly used for extracting the boundary features and boundary labels of the input image to obtain feature maps of different scales; the decoder of the image segmentation network model includes a bidirectional mutual enhancement module and a multi-scale attention aggregation module, which are mainly used for performing attention aggregation processing on the enhanced feature maps based on the scale dimension, the spatial dimension and the channel dimension to obtain a multi-dimensional fusion feature map. The boundary information extracted by this method is more accurate, enabling the obtained multi-dimensional fusion feature map to highlight the spatial information and semantic information that can more effectively improve the accuracy of the segmentation result, thereby improving the accuracy of image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to an image segmentation method based on boundary enhancement. Background Art

[0002] The goal of image segmentation is to segment the input image according to semantic information and predict the semantic category of each pixel from a given label set. With the gradual intelligence of modern life, more and more applications need to infer relevant semantic information from images for subsequent processing, such as augmented reality, autonomous driving, video surveillance, etc. Therefore, accurate segmentation of images becomes very important.

[0003] Traditional image segmentation usually uses traditional machine learning algorithms such as clustering and random forests to obtain image features. In recent years, with the rapid development of professional computing chips, the computing cost has been rapidly reduced, making it possible to widely use deep learning algorithms, thereby significantly improving the accuracy of image segmentation without increasing costs. Therefore, image segmentation methods based on deep learning algorithms have received extensive attention from scholars.

[0004] For example, Jianlong Hou et al. proposed a boundary-sensitive network based on dynamic hybrid gradient convolution and coordination sensitivity in the paper "BSNet: Dynamic Hybrid Gradient Convolution Based Boundary-Sensitive Network for Remote Sensing Image Segmentation" (IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-22, 2022). Chengli Peng et al. proposed a cross-fusion network for extracting multi-scale semantic information in the paper "CrossFusion Net: A Fast Semantic Segmentation Network for Small-Scale Semantic Information Capturing in Aerial Scenes" (IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-13, 2022). Aijin Li et al. proposed a semantic boundary awareness network in the paper "Multitask Semantic Boundary Awareness Network for Remote Sensing Image Segmentation" (IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-14, 2022). Guohui Deng et al. proposed a class-constrained coarse-to-fine attentional deep network in the paper "CCANet: Class-Constraint Coarse-to-Fine Attentional Deep Network for Subdecimeter Aerial Image Semantic Segmentation" (IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-20, 2022).Rui Li et al. proposed a multi-attention network in the paper "Multi-attention Network for Semantic Segmentation of Fine-Resolution Remote Sensing Images" (IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-13, 2022).

[0005] However, the above image segmentation method based on deep learning algorithms still has the following defects:

[0006] First, when extracting boundary information from the feature map in the existing scheme, there are errors in the boundary information in the feature map obtained by operations such as convolution and downsampling, resulting in incorrect spatial information for restoring boundary details.

[0007] Second, in the process of aggregating features of different scales, only simple concatenation or summation operations are used to directly fuse the feature maps of different scales after upsampling, without considering the influence degree of features of different scales on the segmentation result, as well as the different proportions of low-level spatial information and high-level semantic information in features of different scales. Therefore, the spatial information and semantic information in the feature maps of different scales cannot be fully fused, resulting in an unsatisfactory image segmentation result. Summary of the Invention

[0008] In order to solve the above problems existing in the prior art, the present invention provides an image segmentation method based on boundary enhancement. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0009] An image segmentation method based on boundary enhancement, comprising:

[0010] Establishing a boundary-enhanced image segmentation network model based on an encoder-decoder framework;

[0011] Training the image segmentation network model, and using the trained network model to segment the input image to obtain a segmentation result map;

[0012] Wherein, the encoder of the image segmentation network model includes a first feature extraction module, a boundary extraction module, and a second feature extraction module;

[0013] The first feature extraction module is used to extract features from the input image to obtain a first feature map;

[0014] The boundary extraction module is used to extract the boundary features of the first feature map, and at the same time extract the boundary labels of the image labels corresponding to the input image, and use the boundary labels to supervise the output boundary features to obtain a boundary feature map;

[0015] The second feature extraction module is used to perform multi-scale feature extraction on the first feature map and the boundary feature map to obtain feature maps of different scales;

[0016] The decoder of the image segmentation network model includes a bidirectional mutual enhancement module and a multi-scale attention aggregation module; the bidirectional mutual enhancement module is used to process the feature maps of different scales to obtain enhanced feature maps of different scales;

[0017] The multi-scale attention aggregation module is used to perform attention aggregation processing on the enhanced feature maps based on the scale dimension, spatial dimension, and channel dimension to obtain a multi-dimensional fusion feature map.

[0018] Advantages of the present invention:

[0019] On the one hand, when the image segmentation method based on boundary enhancement provided by the present invention extracts the boundaries of the feature maps, it simultaneously extracts the boundaries of the label images, obtains boundary labels, and supervises the boundaries extracted from the feature maps through the boundary labels, so that the extracted boundary information is more accurate; on the other hand, when performing feature fusion, through three-dimensional channel attention processing of scale, space, and channel, the obtained multi-dimensional fusion feature map highlights the spatial information and semantic information that can more effectively improve the accuracy of the segmentation result, thereby improving the accuracy of image segmentation.

[0020] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Brief Description of the Drawings

[0021] Figure 1 is a schematic flow chart of an image segmentation method based on boundary enhancement provided by an embodiment of the present invention;

[0022] Figure 2 is a schematic structural diagram of an image segmentation network model provided by an embodiment of the present invention;

[0023] Figure 3 is a schematic structural diagram of a bidirectional mutual enhancement sub-network provided by an embodiment of the present invention. Detailed Embodiments

[0024] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0025] Embodiment 1

[0026] Please refer toFigure 1 , Figure 1 is a schematic flowchart of an image segmentation method based on boundary enhancement provided by an embodiment of the present invention, which includes:

[0027] Establish a boundary-enhanced image segmentation network model based on an encoder-decoder framework;

[0028] Train the image segmentation network model, and use the trained network model to segment the input image to obtain a segmentation result map;

[0029] Among them, the encoder of the image segmentation network model includes a first feature extraction module, a boundary extraction module, and a second feature extraction module;

[0030] The first feature extraction module is used to extract features from the input image to obtain a first feature map;

[0031] The boundary extraction module is used to extract the boundary features of the first feature map, and at the same time extract the boundary labels of the image labels corresponding to the input image, and use the boundary labels to supervise the output boundary features to obtain a boundary feature map;

[0032] The second feature extraction module is used to perform multi-scale feature extraction on the first feature map and the boundary feature map to obtain feature maps of different scales;

[0033] The decoder of the image segmentation network model includes a bidirectional mutual enhancement module and a multi-scale attention aggregation module; the bidirectional mutual enhancement module is used to process feature maps of different scales to obtain enhanced feature maps of different scales;

[0034] The multi-scale attention aggregation module is used to perform attention aggregation processing on the enhanced feature maps based on the scale dimension, spatial dimension, and channel dimension to obtain a multi-dimensional fusion feature map.

[0035] In this embodiment, the first feature extraction module is a convolutional layer, which includes 1 3×3 convolutional operation Conv, 1 batch normalization BN, and 1 rectified linear unit ReLU. Among them, the calculation formula of the rectified linear function ReLu is as follows:

[0036]

[0037] where x is an element in the input feature map.

[0038] Furthermore, the boundary extraction module in this embodiment can use the Laplace operator, Sobel operator, Canny operator, or LoG operator to implement the extraction of boundary labels.

[0039] Preferably, the Laplace operator is used in this embodiment to construct the boundary extraction module. Please refer to Figure 2 ,Figure 2 It is a schematic structural diagram of the image segmentation network model provided by an embodiment of the present invention; among them, the boundary extraction module includes 2 Laplacian operators and 1 convolutional layer; among them,

[0040] The first Laplacian operator is used to extract the boundary features of the first feature map, and the second Laplacian operator is used to extract the boundary labels of the image labels corresponding to the input image;

[0041] The convolutional layer is used to process the boundary features and supervise the output of the first feature extraction module based on the boundary labels during the model training stage to obtain the boundary feature map.

[0042] It should be noted that the extraction of edges from the image labels using the Laplacian operator and the supervision of the boundary features output by the convolutional layer are only performed during the network training stage. When using the trained network for inference, the boundary extraction module only uses the Laplacian operator to extract the boundaries of the first feature map output by the first feature extraction module, and then processes the boundary features through the convolutional layer. Finally, the boundary feature map output by the convolutional layer is added to the first feature map and input into the second feature extraction module for processing.

[0043] In this embodiment, the second feature extraction module adopts a convolutional neural network architecture, which can be any one of ResNet-50, ResNet-152, ResNeXt-50, ResNeXt-101, or ResNeXt-152.

[0044] Preferably, in this embodiment, a convolutional neural network architecture based on ResNet-101 is used as the second feature extraction module, and its structural diagram is as Figure 2 shown, which includes a downsampling layer, the first stage of ResNet-101, the second stage of ResNet-101, the third stage of ResNet-101, and the fourth stage of ResNet-101; among them,

[0045] The first stage of ResNet-101 contains 3 residual blocks, the second stage of ResNet-101 contains 4 residual blocks, the third stage of ResNet-101 contains 23 residual blocks, and the fourth stage of ResNet-101 contains 3 residual blocks.

[0046] More specifically, the residual block includes: 1 1×1 convolutional operation, 1 3×3 convolutional operation, and 1 1×1 convolutional operation.

[0047] In this embodiment, the downsampling layer is also the maximum pooling, and its stride is 2. In the first stage of ResNet-101, downsampling is achieved through maximum pooling. In the second stage, the third stage, and the fourth stage of ResNet-101, downsampling is achieved by setting the stride of the first convolutional operation in the stage to 2.

[0048] After adding the boundary feature map output by the boundary extraction module and the first feature map output by the first feature extraction module, it is sent to the second feature extraction module for processing, and then high-level features and low-level features of different sizes can be obtained.

[0049] In this embodiment, when extracting the boundary of the feature map, the boundary of the label image is also extracted at the same time to obtain the boundary label, and the boundary extracted from the feature map is supervised by the boundary label, so that the extracted boundary information is more accurate.

[0050] Furthermore, the bidirectional mutual enhancement module includes a number of bidirectional mutual enhancement sub-networks with the same structure. Each bidirectional mutual enhancement sub-network includes a first enhancement module and a second enhancement module, which are respectively used to enhance two feature maps with different sizes.

[0051] Specifically, as Figure 2 shown, based on the ResNet-101 convolutional neural network architecture adopted by the second feature extraction module in this embodiment, two bidirectional mutual enhancement sub-networks are set in the bidirectional mutual enhancement module. The input of one bidirectional mutual enhancement sub-network is the output feature maps of the first stage and the third stage of ResNet-101; the input of the other bidirectional mutual enhancement sub-network is the output feature maps of the second stage and the fourth stage of ResNet-101.

[0052] Furthermore, please refer to Figure 3 , Figure 3 is the structural schematic diagram of the bidirectional mutual enhancement sub-network provided by the embodiment of the present invention. It includes a first enhancement module and a second enhancement module. Among them,

[0053] The first enhancement module first processes the input feature map with the first size through two convolutional layers to obtain the second feature map with the first size; then it processes the input feature map with the first size through a convolutional layer, a pooling layer, and a Sigmoid activation layer in sequence to obtain the third feature map with the second size;

[0054] Correspondingly, the second enhancement module first processes the input feature map with the second size using two convolutional layers to obtain a fourth feature map with the second size; then processes the input feature map with the second size through a convolutional layer, an upsampling layer, and a Sigmoid activation layer in sequence to obtain a fifth feature map with the first size;

[0055] The first enhancement module is also used to multiply the second feature map and the fifth feature map, and obtain a first enhanced feature map with the first size through a convolutional layer;

[0056] The second enhancement module is also used to multiply the third feature map and the fourth feature map, and obtain a second enhanced feature map with the second size through a convolutional layer.

[0057] For example, for the first enhancement module, the size of the input feature map (i.e., the first size) is H×W×C: for the input feature map, the features therein are processed through two convolutional layers to obtain a processed feature map of H×W×C with the same size as the input size, that is, the second feature map. In addition, through a convolutional layer with a stride of 2, a pooling layer with a stride of 2, and a Sigmoid activation layer, the input feature map is processed to obtain an activated feature map of H / 4×W / 4×C, that is, the third feature map.

[0058] For the second enhancement module, the size of the input feature map (i.e., the second size) is H / 4×W / 4×C: for the input feature map, the features therein are processed through two convolutional layers to obtain a processed feature map of H / 4×W / 4×C with the same size as the input size, that is, the fourth feature map. In addition, through a convolutional layer, an upsampling operation with an upsampling multiple of 4, and a Sigmoid activation layer, the input feature map is processed to obtain an activated feature map of H×W×C, that is, the fifth feature map.

[0059] Among them, the calculation formula of the Sigmoid activation function is as follows:

[0060]

[0061] Among them, x is the input feature map.

[0062] Multiply the processed feature map (i.e., the second feature map) of size H×W×C with the activated feature map (i.e., the fifth feature map) of size H×W×C, and through a convolutional layer, obtain an output feature map of size H×W×C, i.e., the first enhanced feature map. Multiply the processed feature map (i.e., the third feature map) of size H / 4×W / 4×C with the activated feature map (i.e., the fourth feature map) of size H / 4×W / 4×C, and through a convolutional layer, obtain an output feature map of size H / 4×W / 4×C, i.e., the second enhanced feature map.

[0063] In this embodiment, the feature maps of the first stage and the third stage of ResNet-101 are input into a bidirectional mutual enhancement sub-network, and the second stage and the fourth stage of ResNet-101 are input into another bidirectional mutual enhancement sub-network, so as to obtain four scales of enhanced feature maps with the same sizes as those of the first, second, third, and fourth stages of ResNet-101.

[0064] Further, please continue to refer to Figure 2 , where the multi-scale attention aggregation module includes a multi-scale fusion sub-network, a scale dimension attention sub-network, a spatial dimension attention sub-network, and a channel dimension attention sub-network; where

[0065] The multi-scale fusion sub-network is used to perform size transformation and feature concatenation on the enhanced feature maps of different scales;

[0066] The scale dimension attention sub-network is used to perform global average pooling on the concatenated feature maps in the scale dimension to obtain a feature map after scale dimension attention processing;

[0067] The spatial dimension attention sub-network is used to perform global average pooling and maximum pooling respectively on the feature map after scale dimension attention processing in the spatial dimension to obtain a feature map after spatial dimension attention processing;

[0068] The channel dimension attention sub-network is used to perform global average pooling on the feature map after spatial dimension attention processing in the channel dimension to obtain a multi-dimensional fusion feature map.

[0069] Specifically, as Figure 2 shown, the multi-scale fusion sub-network performs size transformation on the four scales of enhanced feature maps output by the bidirectional mutual enhancement module to make the sizes of the four scales of feature maps all the same as the size of the output feature map of the first stage of ResNet-101. Then, the feature maps with transformed sizes are concatenated to obtain a feature map with dimensions (scale S, channel C, spatial H×W).

[0070] For the cascaded feature maps, the scale dimension attention sub-network performs global average pooling in the scale dimension to obtain a feature vector with a size of (S×1×1×1). Then, the feature vector is processed through a convolutional layer (with a kernel size of 1×1) to obtain the attention vector in the scale dimension, and the attention vector in the scale dimension is multiplied by the cascaded feature maps of the multi-scale fusion sub-network to obtain the feature maps processed by the scale dimension attention.

[0071] The spatial dimension attention sub-network first processes the feature maps processed by the scale dimension attention through a convolutional layer, and then performs global average pooling and maximum pooling respectively on the processed feature maps in the spatial dimension to obtain two spatial attention maps with a size of (1×1×H×W). The two obtained spatial attention maps are processed through a convolutional layer to output a spatial dimension attention map with a size of (1×1×H×W), and the spatial dimension attention map is multiplied by the feature maps processed by the scale dimension attention to obtain the feature maps processed by the spatial dimension attention.

[0072] The channel dimension attention sub-network first processes the feature maps processed by the spatial dimension attention through a convolutional layer, and then performs global average pooling on the processed feature maps in the channel dimension to obtain a feature vector with a size of (1×C×1×1). Then, the feature vector is processed through a fully connected layer and a convolutional layer (with a kernel size of 1×1) to obtain the attention vector in the channel dimension, and the attention vector in the channel dimension is multiplied by the feature maps processed by the spatial dimension attention to obtain the output feature maps of the multi-scale attention aggregation module.

[0073] In this embodiment, three-dimensional channel attention processing is performed through the scale, spatial, and channel dimensions, so that the obtained multi-dimensional fusion feature maps can highlight the spatial information and semantic information that can more effectively improve the accuracy of the segmentation results.

[0074] It can be understood that the decoder of the image segmentation network model provided in this embodiment further includes an upsampling layer and a convolutional layer after the multi-scale attention aggregation module, which are used to upsample the output feature maps of the multi-scale attention aggregation module with an upsampling ratio of 4, and then obtain the segmentation result map through a convolutional layer.

[0075] On the one hand, when the image segmentation method based on boundary enhancement provided by the present invention extracts the boundary of the feature map, it simultaneously extracts the boundary of the label image to obtain the boundary label, and supervises the boundary extracted from the feature map through the boundary label, so that the extracted boundary information is more accurate. On the other hand, during feature fusion, through three-dimensional attention processing of scale, space, and channel, the obtained multi-dimensional fusion feature map highlights the spatial information and semantic information that can more effectively improve the accuracy of the segmentation result, thereby improving the accuracy of image segmentation.

[0076] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. An image segmentation method based on boundary enhancement, characterized in that Including: Establish a boundary-enhanced image segmentation network model based on the encoder-decoder framework; Train the image segmentation network model, and use the trained network model to segment the input image to obtain a segmentation result map; Among them, the encoder of the image segmentation network model includes a first feature extraction module, a boundary extraction module, and a second feature extraction module; The first feature extraction module is used to extract features from the input image to obtain a first feature map; The boundary extraction module is used to extract the boundary features of the first feature map, and at the same time extract the boundary label of the image label corresponding to the input image, and use the boundary label to supervise the output boundary features to obtain a boundary feature map; The second feature extraction module is used to perform multi-scale feature extraction on the first feature map and the boundary feature map to obtain feature maps of different scales; The decoder of the image segmentation network model includes a bidirectional mutual enhancement module and a multi-scale attention aggregation module; The bidirectional mutual enhancement module is used to process the feature maps of different scales to obtain enhanced feature maps of different scales; The multi-scale attention aggregation module is used to perform attention aggregation processing on the enhanced feature maps based on the scale dimension, spatial dimension, and channel dimension to obtain a multi-dimensional fusion feature map.

2. The image segmentation method based on boundary enhancement according to claim 1, wherein The first feature extraction module is a convolutional layer, which includes 1 3×3 convolutional operation Conv, 1 batch normalization BN, and 1 rectified linear unit ReLU.

3. The image segmentation method based on boundary enhancement according to claim 1, wherein The boundary extraction module uses the Laplace operator, Sobel operator, Canny operator, or LoG operator to extract the boundary label.

4. The image segmentation method based on boundary enhancement according to claim 1, wherein The boundary extraction module includes 2 Laplace operators and 1 convolutional layer; among them, The first Laplace operator is used to extract the boundary features of the first feature map, and the second Laplace operator is used to extract the boundary label of the image label corresponding to the input image; The convolutional layer is used to process the boundary features and supervise the output of the first feature extraction module based on the boundary label during the model training stage to obtain a boundary feature map.

5. The image segmentation method based on boundary enhancement according to claim 1, wherein The second feature extraction module adopts a convolutional neural network architecture, which can be any one of ResNet-50, ResNet-152, ResNeXt-50, ResNeXt-101, or ResNeXt-152.

6. The image segmentation method based on boundary enhancement according to claim 1, wherein, The second feature extraction module adopts a convolutional neural network architecture based on ResNet-101, which includes a downsampling layer, the first stage of ResNet-101, the second stage of ResNet-101, the third stage of ResNet-101, and the fourth stage of ResNet-101; among them, The first stage of ResNet-101 contains 3 residual blocks, the second stage of ResNet-101 contains 4 residual blocks, the third stage of ResNet-101 contains 23 residual blocks, and the fourth stage of ResNet-101 contains 3 residual blocks.

7. The method for image segmentation based on boundary enhancement according to claim 1, wherein The two-way mutual enhancement module includes a number of two-way mutual enhancement sub-networks with the same structure. Each two-way mutual enhancement sub-network includes a first enhancement module and a second enhancement module, which are respectively used to perform enhancement processing on two feature maps of different sizes. Among them, The first enhancement module first processes the input feature map with the first size through two convolutional layers to obtain a second feature map with the first size; then, it sequentially processes the input feature map with the first size through a convolutional layer, a pooling layer, and a Sigmoid activation layer to obtain a third feature map with the second size; Correspondingly, the second enhancement module first processes the input feature map with the second size through two convolutional layers to obtain a fourth feature map with the second size; then, it sequentially processes the input feature map with the second size through a convolutional layer, an upsampling layer, and a Sigmoid activation layer to obtain a fifth feature map with the first size; The first enhancement module is also used to multiply the second feature map and the fifth feature map, and obtain a first enhanced feature map with the first size through a convolutional layer; The second enhancement module is also used to multiply the third feature map and the fourth feature map, and obtain a second enhanced feature map with the second size through a convolutional layer.

8. The method for image segmentation based on boundary enhancement according to claim 6, wherein The two-way mutual enhancement module includes two two-way mutual enhancement sub-networks. The input of one two-way mutual enhancement sub-network is the output feature maps of the first stage and the third stage of ResNet-101; the input of the other two-way mutual enhancement sub-network is the output feature maps of the second stage and the fourth stage of ResNet-101.

9. The image segmentation method based on boundary enhancement according to claim 1, wherein The multi-scale attention aggregation module includes a multi-scale fusion sub-network, a scale dimension attention sub-network, a spatial dimension attention sub-network, and a channel dimension attention sub-network. Among them, The multi-scale fusion sub-network is used to perform size transformation and feature concatenation on the enhanced feature maps of different scales; The scale dimension attention sub-network is used to perform global average pooling on the concatenated feature maps in the scale dimension to obtain a feature map processed by scale dimension attention; The spatial dimension attention sub-network is used to perform global average pooling and maximum pooling respectively on the feature map processed by scale dimension attention in the spatial dimension to obtain a feature map processed by spatial dimension attention; The channel dimension attention sub-network is used to perform global average pooling on the feature map processed by spatial dimension attention in the channel dimension to obtain a multi-dimensional fusion feature map.

Citation Information

Patent Citations

  • Semantic image segmentation method and system based on edge enhancement

    CN111462126A

  • Image segmentation method and system based on edge auxiliary calculation and mask attention

    CN114565770A