A Real-Time Semantic Segmentation Method Based on Attention and Multi-Scale Feature Extraction
Through the real-time semantic segmentation method based on attention and multi-scale feature extraction, the spatial detail and semantic information branch module, the lightweight residual attention module and the deep aggregation pyramid pooling module are used to solve the problem of insufficient accuracy in the real-time semantic segmentation algorithm, and efficient segmentation of small goals is achieved.
Patent Information
- Application Number
- CN202310802309.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-06-30
AI Technical Summary
Existing real-time semantic segmentation algorithms are difficult to achieve an effective balance between accuracy and real-time, especially in small-scale target segmentation, which is insufficient accuracy and high error rate.
Real-time semantic segmentation method based on attention and multi-scale feature extraction is adopted to extract feature information through the spatial detail branch module and the semantic information branch module, and feature expression capabilities are enhanced by lightweight residual attention module and deep aggregation pyramid pooling module, and feature fusion is performed in combination with attention fusion module.
While maintaining real-time, the segmentation accuracy is significantly improved, especially the segmentation accuracy for small targets, achieving an effective balance between accuracy and real-time.
Smart Images

Figure CN117115435B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image semantic segmentation, and particularly to a real-time semantic segmentation method based on attention and multi-scale feature extraction. Background Art
[0002] Semantic segmentation is a basic computer vision task and plays a crucial role in many practical applications, such as medical image segmentation, autonomous driving, and robotics. With the gradual popularization and application of fully convolutional networks in the field of image semantic segmentation, a series of novel network models have been successively proposed. However, the structures of many semantic segmentation models with excellent performance are extremely complex and cumbersome, and are not suitable for being deployed on mobile device platforms with limited computing resources and low latency. Therefore, how to perform fast and accurate semantic segmentation in more real-time scenarios faces new challenges.
[0003] With the continuous increase in the deployment requirements of mobile devices, real-time semantic segmentation has received increasing attention, and many real-time semantic segmentation models based on lightweight convolutional neural networks have been proposed, mainly divided into encoder-decoder structures and multi-branch structures. In the encoder-decoder structure, the encoder part is mainly used to extract image features, and the decoder part is used to sequentially restore image details. However, the encoder-decoder structure needs to process a large amount of information and has high requirements for computing resources and memory management; in the multi-branch structure, image feature information is extracted through different branches and then fused. It utilizes the feature expression capabilities of different branches, can effectively achieve information interaction, and thus improves the accuracy of semantic segmentation. In addition, with the rapid development of the attention mechanism in recent years, attention has been widely applied to real-time semantic segmentation network models, and the extracted features are corrected to better retain valuable feature information.
[0004] In summary, how to effectively balance accuracy and real-time performance, make full use of the feature expression capabilities of different branch structures, and solve the problems of insufficient accuracy and large error rate in segmenting small targets in the current semantic segmentation algorithms in real-time scenarios has become an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] Aiming at the deficiencies of the above-mentioned prior art, the present invention provides a real-time semantic segmentation method based on attention and multi-scale feature extraction, which can effectively balance accuracy and real-time performance, make full use of the feature expression capabilities of different branch structures, and solve the problems of insufficient accuracy and large error rate in segmenting small targets in the current semantic segmentation algorithms in real-time scenarios.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A real-time semantic segmentation method based on attention and multi-scale feature extraction includes the following steps:
[0008] S1. Obtain the original image;
[0009] S2. Perform initial feature extraction on the original image, and extract the initial feature map of the original image through a shared convolutional layer;
[0010] S3. Respectively perform information extraction processing on the obtained initial feature map through a spatial detail branch module SDB and a semantic information branch module SIB; wherein, the spatial detail information of the initial feature map is extracted through the spatial detail branch module SDB, and the semantic context information of the initial feature map is extracted through the semantic information branch module SIB;
[0011] S4. Fuse the extracted spatial detail information and semantic context information of the initial feature map to obtain a fused feature map;
[0012] S5. Send the fused feature map into an image feature classifier to obtain a final feature segmentation image.
[0013] Preferably, in S2, the shared convolutional layer includes three 3×3 convolutional layers connected in sequence, wherein the convolution stride of the first convolutional layer is 2, and the convolution strides of the second and third convolutional layers are 1.
[0014] Preferably, in S3, the spatial detail branch module SDB includes three 3×3 convolutional layers connected in sequence.
[0015] Preferably, in S3, the semantic new information branch module SIB includes a lightweight residual attention module LRA and a deep aggregation pyramid pooling module DAPPM connected in sequence; wherein, the lightweight residual attention module LRA is used to extract context information; the deep aggregation pyramid pooling module DAPPM is used to preserve image edge features;
[0016] During the processing of the semantic new information branch module SIB, the lightweight residual attention module LRA is called multiple times for corresponding processing, and then the deep aggregation pyramid pooling module DAPPM is called for corresponding processing.
[0017] Preferably, the lightweight residual attention module LRA includes a 3×3 convolutional layer, a 3×1 depth convolutional layer, a 1×3 depth convolutional layer, a 3×1 dilated convolutional layer, a 1×3 dilated convolutional layer, a 1×1 recovery convolutional layer, and a compression attention sub-module SA connected in sequence. Finally, the output end of the compression attention sub-module SA and the input feature map of the lightweight sub-module are image-superimposed through a superimposition layer ADD and then output to a channel shuffle module.
[0018] Preferably, the processing process of the compression attention sub-module SA includes:
[0019] First, for the input feature map X in , use global average pooling to compress the two-dimensional features of each channel;
[0020] Then, assign weights to the channel features through the designed attention convolutional layer ACONV; among them, the attention convolutional layer ACONV learns weights through an additional residual path to recalibrate the channels of the output feature X out , and its calculation formula is:
[0021] X out = X a × X res + X a ;
[0022] X a = Up(P Aconv (P Aconv (P APool (X in ))));
[0023] X res = P Conv (X in );
[0024] In the formula, P APool represents global average pooling, P Aconv represents channel attention convolution, Up() represents upsampling, P Conv represents residual convolution, X a represents the output of the attention convolution channel, and X res represents the output of the residual convolution.
[0025] Preferably, the depth aggregation pyramid pooling module DAPPM includes multiple pooling branches for performing pooling processing on the input feature map at different scales, and a connection layer whose input end is respectively connected to the output ends of each pooling branch;
[0026] Each pooling branch includes a pooling layer and a 1×1 upsampling convolutional layer connected in sequence. The pooling processing scales of the pooling layers in each pooling branch are different, and in addition to the first pooling branch, each of the other pooling branches also includes a 3×3 residual convolutional layer. The input end of the 3×3 residual convolutional layer is respectively connected to the output end of the upsampling convolutional layer in its own pooling branch and the output end of the adjacent previous pooling branch, and the output end of the 3×3 residual convolutional layer serves as its output end in the pooling branch;
[0027] The connection layer is used to connect the output features of each pooling branch; the output of the connection layer is also superimposed on the input processing graph to obtain the output of the depth aggregation pyramid pooling module DAPPM; the input processing graph is the feature graph obtained by performing 1×1 convolution on the input feature graph of the depth aggregation pyramid pooling module DAPPM.
[0028] Preferably, the expressions for the outputs of multiple pooling branches of the depth aggregation pyramid pooling module DAPPM are:
[0029]
[0030] In the formula, Y i represents the output of the i-th pooling branch, n represents the number of pooling branches, X represents the input feature graph of the depth aggregation pyramid pooling module DAPPM, C 1×1 represents 1×1 convolution, C 3×3 represents 3×3 convolution, Up represents upsampling, P j,k represents a pooling layer with a kernel size of j and a stride of k, P APool is global average pooling.
[0031] Preferably, in S4, the attention fusion module AFM is used to fuse the spatial detail information and context information from top to bottom.
[0032] Preferably, in S4, the processing process of the attention fusion module AFM includes:
[0033] First, the following calculations are performed on the output feature map S1 of the spatial detail branch module SDB and the output feature map S2 of the semantic information branch module SIB to obtain the process feature maps S H and S W in the image height H direction and the image width W direction respectively:
[0034]
[0035]
[0036] In the formula, C 1×1 represents 1×1 convolution, and represent pooling in the image height H direction and pooling in the image width W direction respectively, S1 is the output feature map of the spatial detail branch module SDB, and S2 is the output feature map of the semantic information branch module SIB;
[0037] Then, the following processing is performed on the process feature maps S H and S W respectively to obtain the corresponding attention weights T H and T W :
[0038] T H = λ(C 1×1 (S H ));
[0039] T W = λ(C 1×1 (S W ));
[0040] Wherein, λ represents sigmoid processing;
[0041] Finally, the output Z of the attention fusion module AFM is obtained through the following calculation:
[0042] Z = (S1 + S2) × T H × T W .
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] 1. Aiming at the insufficient accuracy in the semantic segmentation algorithm in real-time scenarios, the present invention proposes a real-time semantic segmentation method based on attention and multi-scale feature extraction. It includes an initial stage, a Spatial Detail Branch (SDB), a Semantic Information Branch (SIB), and an Attention Fusion Module (AFM). This method makes full use of the feature expression capabilities of different branches and introduces an attention mechanism to highlight important features. In order to obtain semantic information at different scales, a lightweight residual attention module is designed, which combines depthwise separable convolution and dilated convolution, and a channel attention module is introduced to refine features. At the same time, convolutional layers are designed in the spatial detail information branch to extract local information, channel attention and depth aggregation pyramid pooling modules are introduced in stages in the semantic information branch to enhance semantic information, and an attention fusion module is designed to fuse the feature information between different branches from top to bottom, realizing effective interaction of feature information at different scales and improving the segmentation effect.
[0045] Compared with the existing real-time semantic segmentation methods that rely on encoder-decoder structures, the present invention makes full use of the feature expression capabilities of the double-branch structure and designs an attention fusion module to fuse the feature information between different branches from top to bottom, thereby realizing effective interaction of feature information at different scales. It optimizes the segmentation accuracy of small targets in images, making the model have a more general and robust application range.
[0046] In summary, the present method can effectively balance accuracy and real-time performance, make full use of the feature expression capabilities of different branch structures, and solve the problems of insufficient accuracy and large error rate in segmenting small targets in current semantic segmentation algorithms for real-time scenarios.
[0047] 2. Considering the effectiveness of multi-scale feature extraction, the present method designs a lightweight residual attention module LRA, combines depthwise separable convolution and dilated convolution, and introduces a channel attention module to refine features. At the same time, convolution layers are designed in the spatial detail information branch to extract local information, and channel attention and depthwise pyramid pooling modules are introduced stage by stage in the semantic information branch to enhance semantic information. An effective balance between segmentation accuracy and inference speed is effectively achieved.
[0048] 3. Currently, common residual feature extraction modules include the one-dimensional non-bottleneck module (Non-bottleneck-1D) and the depthwise asymmetric bottleneck module (Depth-wise Asymmetric Bottleneck, DAB). Non-bottleneck-1D uses one-dimensional decomposed convolution to accelerate and reduce the number of parameters, and DAB designs a dual-branch structure using depthwise separable convolution and dilated convolution. However, these methods ignore the correlation between context information and do not consider the acquisition of multi-scale information, which may lead to classification errors for small targets. The present invention designs a lightweight residual attention module LRA, which captures more significant information features by introducing a squeeze-and-attention (SA) sub-module to improve its representation ability. The LRA module designed by the present invention effectively improves the representation ability of traditional residual modules by selectively weighting the feature mapping channels through the introduction of the SA sub-module. At the same time, the SA does not fully compress the spatial detail information and retains the detail features, making it more suitable for lightweight networks with fewer layers and high segmentation accuracy requirements.
[0049] 4. When performing information fusion, according to the excitation strategy of the attention mechanism, the AFM module can not only capture cross-channel information but also capture spatial detail information at different positions, occupying very few parameters while maintaining high-efficiency fusion. Through this module, the features of the two branches can be fully fused, and the feature information can be adaptively highlighted in both channels and space. Finally, a fine segmentation image is predicted on the classifier. In this way, the AFM can fully fuse the features of the two branches and adaptively highlight the feature information in both the channel and spatial dimensions simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, where:
[0051] Figure 1 is the flowchart of the present invention;
[0052] Figure 2 is the schematic diagram of the overall network structure used in the present invention;
[0053] Figure 3 is the schematic diagram of the structure of the feature extraction module in the embodiment;
[0054] Figure 4 is the schematic diagram of the structure of the Deep Aggregation Pyramid Pooling Module (DAPPM) in the embodiment;
[0055] Figure 5 is the schematic diagram of the structure of the Attention Fusion Module (AFM) in the embodiment;
[0056] Figure 6 is the schematic diagram of the result comparison of adding different modules in the verification experiment of the embodiment;
[0057] Figure 7 is the schematic diagram of the result comparison of different algorithms on the Cityscapes dataset in the verification experiment of the embodiment. Detailed implementation manners
[0058] The following is a further detailed description through specific implementation manners:
[0059] Embodiment:
[0060] As Figure 1 shown, in this embodiment, a real-time semantic segmentation method based on attention and multi-scale feature extraction is disclosed. The structure of the network used in this method is as Figure 2 shown, and mainly includes an initial stage, a Spatial Detail Branch (SDB), a Semantic Information Branch (SIB), and an Attention Fusion Module (AFM).
[0061] This method includes the following steps:
[0062] S1. Obtain the original image;
[0063] S2. Perform initial feature extraction on the original image, and extract the initial feature map of the original image through a shared convolutional layer;
[0064] Specifically, in step S2, the shared convolutional layer includes three sequentially connected 3×3 convolutional layers, where the convolution stride of the first convolutional layer is 2, and the convolution strides of the second and third convolutional layers are 1. Since only one downsampling operation is performed in the initial stage, the spatial detail information of the original image can be effectively preserved.
[0065] S3. Process the obtained initial feature maps through a spatial detail branch module SDB and a semantic information branch module SIB respectively for information extraction; wherein, extract the spatial detail information of the initial feature map through the spatial detail branch module SDB, and extract the semantic context information of the initial feature map through the semantic information branch module SIB.
[0066] Extract spatial detail information and semantic information through SDB and SIB respectively, share the initial stage to achieve information interaction between different branches, and facilitate subsequent feature fusion.
[0067] In specific implementation, the spatial detail branch SDB mainly consists of 3 3×3 convolutional layers. To obtain more features, the number of channels in the middle convolutional layer is expanded, and this operation can remove redundancy and enhance the extraction of spatial detail information features.
[0068] The semantic new information branch SIB mainly consists of a channel attention module (Channel Attention Module, CAM), a lightweight residual attention module (Lightweight residual attention module, LRA), and a deep aggregation pyramid pooling module (Deep Aggregation Pyramid Pool Module, DAPPM). At the same time, in this branch, two downsampling and upsampling operations are respectively introduced to process the feature maps to change the resolution of the feature maps at different stages and ensure the removal of redundant information in the image during the extraction of semantic feature information.
[0069] The semantic new information branch module SIB includes a lightweight residual attention module LRA and a deep aggregation pyramid pooling module DAPPM connected in sequence; wherein, the lightweight residual attention module LRA is used to extract context information; the deep aggregation pyramid pooling module DAPPM is used to preserve image edge features. In addition, since the channels contain rich feature information and interference noise, therefore, SIB first emphasizes the features that need to be highlighted by introducing CAM, and at the same time suppresses the interference noise. The processing process of the semantic new information branch SIB includes: first, emphasize the features that need to be highlighted by the channel attention module CAM, and at the same time suppress the interference noise; then, extract context information through the lightweight residual attention module LRA; then, use the deep aggregation pyramid pooling module DAPPM after the feature extraction stage to improve the model's ability to obtain global information and preserve image edge features. To ensure the suppression effect on noise, a channel attention module CAM can also be used after the deep aggregation pyramid pooling module DAPPM.
[0070] In order to extract more refined context information, during the processing of the semantic new information branch module SIB, the lightweight residual attention module LRA is called multiple times for corresponding processing first, and then the deep aggregation pyramid pooling module DAPPM is called for corresponding processing. As a specific implementation manner, the LRA module designed in the present invention adopts depthwise separable convolution and dilated convolution, and calls the LRA module with a dilation rate of 2 three times and the dilation rate sequence of {4, 4, 8, 8, 16, 16} six times at the 1 / 4 and 1 / 8 feature map resolutions respectively, gradually increasing the receptive field and enhancing the representation of detailed features. This structural design not only helps to extract local and global features and enhance the context connection between pixels, but also can obtain context information similar to that of a deep network in the shallow layer with fewer parameters.
[0071] Currently, the commonly used residual feature extraction modules include the one-dimensional non-bottleneck module (Non-bottleneck-1D) and the depth-wise asymmetric bottleneck module (Depth-wise Asymmetric Bottleneck, DAB), as shown in Figure 3 Figures 3(a) and 3(b) in [reference]. Non-bottleneck-1D uses one-dimensional decomposed convolution to accelerate and reduce the number of parameters, and DAB designs a double-branch structure using depthwise separable convolution and dilated convolution. However, these methods ignore the correlation between context information and do not consider the acquisition of multi-scale information, which may lead to classification errors for small targets. The present invention designs a lightweight residual attention module LRA, as shown in Figure 3 Figure 3(c), which improves its representation ability by introducing squeeze-and-attention (SA) to capture more significant information features.
[0072] The lightweight residual attention module LRA includes a 3×3 convolutional layer, a 3×1 depth convolutional layer, a 1×3 depth convolutional layer, a 3×1 dilated convolutional layer, a 1×3 dilated convolutional layer, a 1×1 recovery convolutional layer, and a squeeze-and-attention sub-module SA connected in sequence. Finally, after the output end of the squeeze-and-attention sub-module SA and the input feature map of the lightweight sub-module are image-superimposed through a superimposition layer ADD, the output is sent to the channel shuffle module.
[0073] The processing process of the lightweight residual attention module LRA includes:
[0074] First, for the input feature map, the number of channels is reduced to half of the original using the 3×3 convolutional layer;
[0075] Then, the downsampled feature map is successively subjected to 3×1 depthwise convolution layer and 1×3 depthwise convolution layer to perform 3×1 and 1×3 depthwise convolutions, and then successively through 3×1 dilated convolution layer and 1×3 dilated convolution layer to perform 3×1 and 1×3 dilated convolutions. In this way, depthwise separable convolution and dilated convolution are combined to reduce the number of parameters to a greater extent. At the same time, dilated convolutions with different dilation rates are used in different network layers to expand the receptive field and obtain multi-scale context information;
[0076] Finally, a 1×1 restoration convolution layer is used to restore the number of channels, and a squeeze-and-attention sub-module SA is introduced to strengthen the feature representation. After the image end output by the squeeze-and-attention sub-module SA and the input feature map of the lightweight sub-module are image-superimposed through the addition layer ADD, it is output to the channel shuffle module (Channel Shuffle), and channel shuffle is used to enhance information interaction.
[0077] Each channel in the feature map obtained after image convolution usually represents different features. These features have different effects on segmentation, but each channel maintains the same weight, which cannot reflect the dependence relationship between different channels and is not conducive to the extraction of target feature information. The present invention introduces a squeeze-and-attention sub-module SA, which extracts more salient features by introducing local and global weighting mechanisms, and establishes channel context dependencies to enhance the expression between features. The squeeze-and-attention sub-module SA is as shown in Figure 3 (c). The processing process of the squeeze-and-attention sub-module SA includes:
[0078] First, for the input feature map X in , global average pooling is used to compress the two-dimensional features of each channel; then, an attention convolution layer (ACONV) is designed to assign weights to the channel features, and an additional residual path is designed to learn the weights to recalibrate the channels of the output feature X out , specifically:
[0079] X out =X a ×X res +X a (1)
[0080] wherein, the calculation methods of X a and X res are respectively:
[0081] X a =Up(P Aconv (P Aconv (P APool (X in )))) (2)
[0082] X res =P Conv (Xin ) (3)
[0083] In the formula, P APool represents global average pooling, P Aconv represents channel attention convolution, Up() represents upsampling, and P Conv represents residual convolution, and X a represents the output of the attention convolution channel, and X res represents the output of the residual convolution.
[0084] In the LRA module designed by the present invention, by introducing the compression attention sub-module SA, the selective weighting of the feature map channels effectively improves the representation ability of the traditional residual module. At the same time, SA does not completely compress the spatial detail information, retains the detail features, and is more suitable for lightweight networks with fewer layers and high segmentation accuracy requirements.
[0085] In order to aggregate the context information of different regions, after the feature extraction stage, the DAPPM is used to improve the model's ability to obtain global information and preserve the image edge features.
[0086] To better capture the semantic information of different scales in the image, the present invention adopts the depth aggregation pyramid pooling module (DAPPM module) to further extract the context information from the low-resolution feature map.
[0087] The structure of the depth aggregation pyramid pooling module DAPPM is as Figure 4 shown, including multiple pooling branches for performing different-scale pooling processing on the input feature map respectively, and a connection layer with its input end connected to the output ends of each pooling branch respectively;
[0088] Each pooling branch includes a pooling layer and a 1×1 upsampling convolutional layer connected in sequence. The pooling processing scales of the pooling layers in each pooling branch are different, and each pooling branch except the first pooling branch further includes a 3×3 residual convolutional layer. The input end of the 3×3 residual convolutional layer is respectively connected to the output end of the upsampling convolutional layer in its own pooling branch and the output end of the adjacent previous pooling branch, and the output end of the 3×3 residual convolutional layer serves as the output end of its pooling branch;
[0089] The connection layer is used to connect the output features of each pooling branch; the output of the connection layer is also superimposed on the input processing graph to obtain the output of the depth aggregation pyramid pooling module DAPPM; the input processing graph is the feature graph obtained after the input feature graph of the depth aggregation pyramid pooling module DAPPM is convolved by 1×1.
[0090] The feature map is processed through pooling operations of different scales to obtain feature information at different scales. Considering that a single 3×3 or 1×1 convolution is not sufficient to mix all multi-scale context information, the present invention first upsamples the feature map and then designs more 3×3 convolutions to fuse context information of different scales in a hierarchical residual manner. If the input feature map is X, the expressions of the outputs of each pooling branch of the depth aggregation pyramid pooling module DAPPM are as follows:
[0091]
[0092] In the formula, n represents the scale size, C 1×1 represents a 1×1 convolution, C 3×3 represents a 3×3 convolution, Up represents upsampling, P j,k represents a pooling layer with a kernel size of j and a stride of k, P APool is global average pooling.
[0093] Finally, all feature maps are connected and compressed through a 1×1 convolution, and a 1×1 convolution is added from the input end to further optimize the output features.
[0094] S4. Fuse the feature information extracted from different branches in the feature extraction stage.
[0095] In the feature fusion part, traditional methods such as concatenation inevitably ignore the correlation between features. To better fuse the features between different branches, the present invention uses AFM to guide the effective fusion of spatial detail information and context information from top to bottom. According to the excitation strategy of the attention mechanism, this module can not only capture cross-channel information but also capture spatial detail information at different positions, occupying very few parameters while maintaining high-efficiency fusion. Through this module, the features of the two branches can be fully fused, and the feature information can be adaptively highlighted in channels and space. Finally, a fine-segmented image is predicted on the classifier.
[0096] How to effectively fuse the feature information between different branches is a key issue in the dual-branch network. Commonly used feature fusion methods mainly include channel merging (Concat) or pixel-level addition (Add). However, these methods ignore the position difference and information diversity between the features of each layer, and also do not consider the correlation between pixels, which is prone to classification errors and thus reduces the segmentation performance of the algorithm. The present invention uses AFM to efficiently fuse local features and global features. The structure of the attention fusion module AFM is as Figure 5 shown.
[0097] AFM aggregates feature information of different scales, retains both global and local feature information, and can effectively improve the semantic segmentation performance without introducing too much computational complexity.
[0098] Specifically, first, the following calculations are performed on the output feature map S1 of the spatial detail branch module SDB and the output feature map S2 of the semantic information branch module SIB to obtain the process feature maps S H and S W :
[0099]
[0100]
[0101] where C 1×1 represents a 1×1 convolution, and and represent pooling in the image height H direction and pooling in the image width W direction respectively, s1 is the output feature map of the spatial detail branch module SDB, and S2 is the output feature map of the semantic information branch module SIB;
[0102] Then, the following processing is performed on the process feature maps S H and S W respectively to obtain the corresponding attention weights T H and T W :
[0103] T H = λ(C 1×1 (S H ));
[0104] T W = λ(C 1×1 (S W ));
[0105] where λ represents sigmoid processing;
[0106] The above processing uses a reduction factor r (i.e., split operation) to reduce the channel dimension, performs channel reduction on S H , S W to obtain the corresponding attention weights T H and T W .
[0107] Finally, the output Z of the attention fusion module AFM is obtained through the following calculation:
[0108] Z = (S1 + S2) × T H × T W .
[0109] In this way, AFM can fully fuse the features of the two branches and adaptively highlight the feature information in both the channel and spatial dimensions.
[0110] S5. Feed the fused feature map into the image feature classifier to obtain the final feature segmentation image.
[0111] Compared with the existing real-time semantic segmentation methods that rely on the encoder-decoder structure, the present invention fully utilizes the feature expression ability of the dual-branch structure and designs an attention fusion module to fuse the feature information between different branches from top to bottom, thereby realizing the effective interaction of feature information at different scales. The segmentation accuracy of small targets in the image is optimized, making the model have a more general and robust application range. Considering the effectiveness of multi-scale feature extraction, the present method designs a lightweight residual attention module LRA, combines depthwise separable convolution and dilated convolution, and introduces a channel attention module to refine the features; at the same time, a convolutional layer is designed in the spatial detail information branch to extract local information, and channel attention and depth aggregation pyramid pooling modules are introduced periodically in the semantic information branch to enhance the semantic information. An effective balance between segmentation accuracy and inference speed is achieved effectively.
[0112] In addition, the LRA module designed by the present invention effectively improves the representation ability of the traditional residual module by introducing a squeeze attention sub-module SA for selective weighting of the feature mapping channels. At the same time, SA does not completely compress the spatial detail information and retains the detail features, which is more suitable for lightweight networks with fewer layers and high segmentation accuracy requirements. Moreover, when performing information fusion, according to the excitation strategy of the attention mechanism, the AFM module can not only capture cross-channel information but also capture the spatial detail information at different positions, occupying very few parameters while maintaining high-efficiency fusion. Through this module, the features of the two branches can be fully fused, and the feature information can be adaptively highlighted in both channels and space. Finally, a fine segmentation image is predicted on the classifier. In this way, the AFM can fully fuse the features of the two branches and adaptively highlight the feature information in both the channel and spatial dimensions at the same time.
[0113] In summary, the present method can achieve an effective balance between accuracy and real-time performance, fully utilize the feature expression ability of different branch structures, and solve the problems of insufficient accuracy and large error rate in segmenting small targets in the current semantic segmentation algorithms for real-time scenarios.
[0114] To verify the effect of the technical solution disclosed by the present invention, in this embodiment, two common datasets, Cityscapes and CamVid, are used for experiments and the results are analyzed to verify the effectiveness of the algorithm.
[0115] Cityscapes is a large dataset of urban street scenes, which is widely used in the field of semantic segmentation. The experiment uses 5000 finely annotated images for training, validation and testing, with the numbers being 2975, 500 and 1525 respectively, and contains 19 categories. CamVid is a street view dataset of driving cars, which contains 11 categories and 701 finely annotated images, and these images are divided into 367 training samples, 101 validation samples and 233 test samples.
[0116] The Mean Intersection over Union (MIoU) and Frames Per Second (FPS) are used as metrics for accuracy and inference speed. These two evaluation metrics are the main standard metrics in the current field of real-time semantic segmentation.
[0117] The experimental environment is based on Pytorch 1.11.0 and python 3.7.9, and the experiment is carried out on a single GTX 4090 GPU. The experimental settings are as follows: The Stochastic Gradient Descent (SGD) optimizer is used to optimize the network with a batch size of 8, the maximum number of training epochs is 1000, the momentum and weight decay are set to 0.9 and 1e-4 respectively, and the learning rate adopts the "poly" strategy. The adaptive learning rate is adjusted after each iteration as follows:
[0118]
[0119] where lr is the learning rate after each iteration, lr init is the initial learning rate, iter is the current iteration index, max_iter is the maximum number of iterations in each epoch, power is the momentum, and the initial learning rate is set to 4.5e-2.
[0120] Regarding data augmentation, random horizontal flipping, mean attenuation and random cropping are used for the input images during training, and the random values are numbers in {0.75, 1.0, 1.25, 1.5, 1.75, 2.0}. The Cityscapes dataset is randomly cropped to a resolution of 512×1024 for training.
[0121] To verify the performance of each module of the present invention, the present invention uses MIoU and FPS as evaluation criteria to conduct experimental comparative analysis on each module separately. The experimental results are shown in Table 1. It includes a semantic information branch SIB, a spatial detail information branch SDB, and a feature fusion part. Among them, SIB mainly includes CAM and DAPPM, SDB mainly considers 3 additional 3×3 convolutions added, and the feature fusion part mainly includes Cat, Add, and AFM. In addition, since the feature extraction module LRA of the present invention uses dilated convolution, in order to further verify the influence under different dilation rate settings, two additional comparative experiments are added.
[0122] Table 1 Ablation experiment results on the Cityscapes dataset
[0123]
[0124] First, only use the LRA module designed by the present invention to extract image features and perform feature fusion, and then output the results, obtaining an accuracy of 72.9% and a speed of 163 frames / s. Second, on this basis, introduce the additional convolutional layer in SDB to extract richer spatial detail information, obtaining an accuracy of 73.1% and a speed of 155 frames / s. Then, consider introducing CAM and DAPPM respectively to highlight the expression of important feature channel information. When only CAM is introduced, an accuracy of 73.4% and a speed of 154 frames / s are obtained. When only DAPPM is introduced, an accuracy of 73.6% and a speed of 144 frames / s are obtained. Finally, on the basis of introducing CAM and DAPPM, remove the influence of the convolutional layer in SDB, obtaining an accuracy of 74.1% and a speed of 149 frames / s.
[0125] To more intuitively show the performance of each module, Figure 6 The segmentation maps with different modules added are shown. It can be seen that when an additional convolution is added to SDB, the expression of spatial detail information in the image is more accurate; after adding the channel attention module CAM, when similar targets appear in the image, the mutual interference between targets is reduced, avoiding classification errors; after adding the pyramid pooling module DAPPM, when target overlaps appear in the image, by obtaining multi-scale feature information, some occluded small targets can be correctly segmented to a certain extent.
[0126] In the feature fusion stage, to explore the impact of the attention fusion module AFM of the present invention, two groups of comparative experiments were designed, using the common feature fusion methods Cat and Add. Using the Cat fusion method, an accuracy of 73.6% and a speed of 146 frames / s were obtained. Using the Add fusion method, an accuracy of 72.4% and a speed of 152 frames / s were obtained. The AFM fusion method adopted by the present invention can achieve an accuracy of 74.4% and a speed of 138 frames / s. It can be seen that AFM takes into account the differences and diversities between different branch features, has a higher segmentation accuracy, and does not need to sacrifice too much inference speed.
[0127] In the feature extraction stage of AMFENet, at the 1 / 8 feature map resolution, the LRA module with the dilation rate sequence {r = 4, 4, 8, 8, 16, 16} was adopted. To verify the effectiveness of this design, two additional groups of comparative experiments were added, and the dilation rates were set to a fixed dilation rate {r = 4, 4, 4, 4, 4, 4} and relatively prime dilation rates {r = 3, 3, 7, 7, 13, 13} respectively. Using the LRA module with the fixed dilation rate {r = 4, 4, 4, 4, 4, 4}, an accuracy of 70.9% and a speed of 132 frames / s were obtained. Using the LRA module with relatively prime dilation rates {r = 3, 3, 7, 7, 13, 13}, an accuracy of 73.0% and a speed of 141 frames / s were obtained. It can be seen that the fixed dilation rate is limited by the size of its receptive field, which is not conducive to capturing multi-scale feature information, and using multiple convolutions with the same dilation rate will produce a grid effect, reducing the segmentation accuracy. The relatively prime dilation rates still have a relatively small receptive field compared to the present invention, which is not conducive to extracting context information in a larger range.
[0128] The above experimental studies have proved that the designs of each module in AMFENet have improved the overall segmentation accuracy of the model and enhanced the robustness of the model, fully verifying the effectiveness and rationality of the AMFENet network structure design.
[0129] Table 2 Comparison results of different algorithms on the Cityscapes dataset
[0130]
[0131] To verify the effectiveness of the present invention, AMFENet was compared and analyzed with other advanced real-time semantic segmentation algorithms on the Cityscapes dataset, and the results are shown in Table 2. The experimental results show that, since AMFENet introduces an attention module to strengthen feature representation and realizes feature interaction at different scales, AMFENet achieves the optimal performance in terms of segmentation accuracy and has a sub-optimal inference speed, effectively achieving a balance between accuracy and speed. Although BiseNet-v2 has the fastest speed, its accuracy is 1.8% lower than that of AMFENet, and its feature representation ability is inferior to AMFENet. ENet and ESPNet perform optimally in terms of model parameters, but they lose a large amount of image edge detail information in the decoding stage, and both the segmentation accuracy and speed are significantly lower than those of AMFENet. In addition, several other comparison methods with fewer parameters than AMFENet perform worse than AMFENet in terms of segmentation accuracy and speed. BiseNet and DFANet, which have similar numbers of parameters to AMFENet, both require pre-training, and their segmentation accuracy and speed are significantly lower than those of AMFENet, and there are differences in overall performance compared with AMFENet. Through comprehensive comparative analysis of the three performance indicators, AMFENet of the present invention takes into account both segmentation accuracy and inference speed while keeping the model lightweight, and effectively achieves a balance between accuracy and real-time performance.
[0132] Figure 7 The subjective comparison effects of DABNet, FBSNet, and AMFENet in terms of segmentation accuracy are shown. From the segmentation results in the first row, it can be seen that AMFENet can clearly and accurately segment the utility poles on both sides of the road; while there are segmentation errors or failures in the segmentation results of DABNet and FBSNet, indicating that the present invention has a more accurate segmentation effect on small targets at close range. From the segmentation results in the second and third rows, it can be seen that AMFENet can accurately segment the traffic lights in the distance and can also well segment a group of people in the image without interference with each other; DABNet and FBSNet cannot segment the traffic lights in the distance, and DABNet has weak anti-interference ability when segmenting the crowd, and artifacts and other phenomena will appear in the segmentation results, fully indicating that the present invention has a better segmentation effect on some small targets at a distance and has strong anti-interference ability when segmenting multiple targets. From the segmentation results in the fourth row, it can be seen that AMFENet can accurately segment the occluded and overlapping parts in the image, avoiding mutual interference between different categories; DABNet and FBSNet have classification errors when segmenting the overlapping parts and have a very poor segmentation effect on small categories such as utility poles. The experimental results show that the present invention demonstrates more excellent semantic segmentation ability and classification recognition ability.
[0133] Table 3 shows the comparison between AMFENet and other advanced real-time semantic segmentation algorithms on the CamVid dataset. The experimental results show that in the low-resolution CamVid dataset, the design of AMFENet combined with a dual-branch network structure optimizes the segmentation effect of small targets in images, and the introduced attention module effectively highlights the expression of important features. AMFENet achieves the best in both segmentation accuracy and inference speed, and also has a certain advantage in the number of parameters. Although ENet has the lowest number of parameters, the speed of AMFENet is better than that of ENet, and the accuracy is 16.4% higher than that of ENet. The number of parameters of ICNet is about 5 times more than that of AMFENet, and the segmentation accuracy is similar to that of AMFENet, but the speed is much lower than that of AMFENet. Although the number of parameters of DABNet is lower than that of AMFENet, it is significantly inferior to AMFENet in both segmentation accuracy and speed. The results of DFANet are also inferior to those of AMFENet in all aspects. It can be seen that AMFENet can also achieve excellent segmentation performance on the CamVid dataset, which fully demonstrates the strong robustness of AMFENet.
[0134] Table 3 Comparison results of different algorithms on the CamVid dataset
[0135]
[0136]
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and do not limit the technical solutions. Those of ordinary skill in the art should understand that any modifications or equivalent replacements made to the technical solutions of the present invention without departing from the purpose and scope of the present technical solution shall be covered by the scope of the claims of the present invention.
Claims
1. A real-time semantic segmentation method based on attention and multi-scale feature extraction, characterized in that, It includes the following steps: S1. Obtain the original image; S2. Conduct initial feature extraction on the original image, and extract the initial feature map of the original image through a shared convolutional layer; S3. Respectively perform information extraction processing on the obtained initial feature map through a spatial detail branch module SDB and a semantic information branch module SIB; wherein, extract the spatial detail information of the initial feature map through the spatial detail branch module SDB, and extract the semantic context information of the initial feature map through the semantic information branch module SIB; S4. Fuse the extracted spatial detail information and semantic context information of the initial feature map to obtain a fused feature map; S5. Send the fused feature map into an image feature classifier to obtain the final feature segmentation image; Among them, in S3, the semantic new information branch module SIB includes a lightweight residual attention module LRA and a depth aggregation pyramid pooling module DAPPM connected in sequence; wherein, the lightweight residual attention module LRA is used to extract context information; the depth aggregation pyramid pooling module DAPPM is used to preserve image edge features; During the processing of the semantic information branch module SIB, first call the lightweight residual attention module LRA multiple times for corresponding processing, and then call the depth aggregation pyramid pooling module DAPPM for corresponding processing; In S4, an attention fusion module AFM is used to fuse the spatial detail information and context information from top to bottom; the processing process of the attention fusion module AFM includes: First, the following calculations are performed on the output feature map S1 of the spatial detail branch module SDB and the output feature map S2 of the semantic information branch module SIB to obtain the process feature maps S in the image height H direction and the image width W direction respectively H and S W : where C 1×1 represents a 1×1 convolution, and represent pooling in the image height H direction and pooling in the image width W direction respectively. S1 is the output feature map of the spatial detail branch module SDB, and S2 is the output feature map of the semantic information branch module SIB; Then, process the process feature maps S H and S W separately as follows to obtain the corresponding attention weights T H and T W : T H = λ(C 1×1 (S H )); T W = λ(C 1×1 (S W )); In the formula, λ represents sigmoid processing; Finally, the output Z of the attention fusion module AFM is obtained through the following calculation: Z = (S1 + S2) × T H × T W 。 2. The real-time semantic segmentation method based on attention and multi-scale feature extraction according to claim 1, characterized in that: In S2, the shared convolutional layer includes three 3×3 convolutional layers connected in sequence, wherein the convolution stride of the first convolutional layer is 2, and the convolution strides of the second and third convolutional layers are 1.
3. The real-time semantic segmentation method based on attention and multi-scale feature extraction according to claim 1, characterized in that: In S3, the spatial detail branch module SDB includes three 3×3 convolutional layers connected in sequence.
4. The real-time semantic segmentation method based on attention and multi-scale feature extraction according to claim 1, characterized in that: The lightweight residual attention module LRA includes a 3×3 convolutional layer, a 3×1 depth convolutional layer, a 1×3 depth convolutional layer, a 3×1 dilated convolutional layer, a 1×3 dilated convolutional layer, a 1×1 recovery convolutional layer, and a compression attention sub-module SA connected in sequence. Finally, the output end of the compression attention sub-module SA and the input feature map of the lightweight sub-module are image-superimposed through a superimposed layer ADD and then output to a channel shuffle module.
5. The real-time semantic segmentation method based on attention and multi-scale feature extraction according to claim 4, characterized in that: The processing process of the compression attention sub-module SA includes: First, for the input feature map X in , use global average pooling to compress the two-dimensional features of each channel; Then, weights are assigned to the channel features through the designed attention convolutional layer ACONV; among them, the attention convolutional layer ACONV learns weights through an additional residual path to recalibrate the output feature X out channels, and its calculation formula is: X out = X a × X res + X a ; X a = Up(P Aconv (P Aconv (P APool (X in )))); X res = P Conv (X in ) Wherein, P APool represents global average pooling, P Aconv represents channel attention convolution, Up() represents upsampling, P Conv represents residual convolution, X a represents the output of the attention convolution channel, X res represents the output of the residual convolution.
6. The real-time semantic segmentation method based on attention and multi-scale feature extraction according to claim 5, characterized in that: The depth aggregation pyramid pooling module DAPPM includes multiple pooling branches for respectively performing pooling processing on the input feature map at different scales, and a connection layer whose input end is respectively connected to the output ends of each pooling branch; Each pooling branch includes a pooling layer and a 1×1 upsampling convolutional layer connected in sequence. The pooling scales of the pooling layers in each pooling branch are different. In addition to the first pooling branch, each of the other pooling branches also includes a 3×3 residual convolutional layer. The input end of the 3×3 residual convolutional layer is respectively connected to the output end of the upsampling convolutional layer in its own pooling branch and the output end of the previous adjacent pooling branch. The output end of the 3×3 residual convolutional layer serves as its output end in the pooling branch. The connection layer is used to connect the output features of each pooling branch. The output of the connection layer is also superimposed on the input processing graph to obtain the output of the Deep Aggregation Pyramid Pooling Module (DAPPM). The input processing graph is the feature graph obtained after 1×1 convolution of the input feature graph of the Deep Aggregation Pyramid Pooling Module (DAPPM).
7. The real-time semantic segmentation method based on attention and multi-scale feature extraction according to claim 6, characterized in that: The expressions for the outputs of the pooling branches of the Deep Aggregation Pyramid Pooling Module (DAPPM) are as follows: Where Y i represents the output of the i-th pooling branch, n represents the number of pooling branches, X represents the input feature map of the depth aggregation pyramid pooling module DAPPM, and C 1×1 represents a 1×1 convolution, C 3×3 represents a 3×3 convolution, Up represents upsampling, and P j,k represents a pooling layer with a kernel size of j and a stride of k, and P APool is global average pooling.
Citation Information
Patent Citations
Semantic segmentation method of attention mechanism based on deep learning
CN112287940A
Lightweight multi-scale feature fusion real-time image semantic segmentation method and system
CN114445430A