Multi-scale attention and median enhanced semantic segmentation method

Through the semantic segmentation method of multi-scale attention and median enhancement, the problem of low efficiency of global information capture and cross-channel information interaction in the prior art is solved, clearer edge segmentation and more efficient information fusion are achieved, and the safety and reliability of autonomous driving are improved.

CN120580438AActive Publication Date: 2025-09-02NANCHANG INST OF TECH

Patent Information

Application Number
CN202511078951.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-02
Publication Date
2025-09-02
Estimated Expiration
2045-08-02

AI Technical Summary

Technical Problem

The existing semantic segmentation methods are difficult to effectively capture the global information of the image in autonomous driving, especially the interaction efficiency of edge details and cross-channel information, which leads to the inability of autonomous driving vehicles to accurately perceive objects in complex scenarios, affecting safety and reliability.

Method used

The semantic segmentation method of multi-scale attention and median enhancement is adopted to extract features through the ResNeSt50 backbone network, combining split attention blocks, multi-scale feature enhancement modules and median enhancement spatial channel attention modules to realize dynamic weight allocation and efficient information fusion of features, improving edge segmentation effect and cross-channel information interaction.

Benefits of technology

It significantly improves the edge segmentation effect and cross-channel information fusion efficiency, improves the accuracy of objects perceived by autonomous vehicles in complex scenarios, and enhances the safety and reliability of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580438A_ABST
    Figure CN120580438A_ABST
Patent Text Reader

Abstract

The invention records a multi-scale attention and median enhancement semantic segmentation method, and the method comprises the following steps: S1, extracting the initial features of an original image, and forming the output features of an Inception block; s2, inputting the output features into a multi-scale feature enhancement module, and outputting multi-scale features; s3, inputting the multi-scale features into a median enhancement space channel attention module, and outputting an attention feature map; s4, generating a final fusion feature of the high-level feature and the low-level feature through the cross feature fusion block; s5, performing up-sampling on the attention feature map, adding the attention feature map with the final fusion feature, and performing up-sampling to obtain a fusion feature map; s6, performing up-sampling and splicing on the fused feature map, and outputting a spliced feature map; and S7, performing up-sampling and full-connection layer expansion on the spliced feature map to obtain a segmentation prediction map of the original image. According to the method, feature distribution information can be captured more comprehensively, and a clearer and more accurate original image edge segmentation result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and autonomous driving technology, and particularly to a semantic segmentation method with multi-scale attention and median enhancement. Background Art

[0002] Semantic segmentation is the process of dividing an image into regions of different colors based on semantic meaning and labeling the segmented regions according to semantic categories, thereby obtaining an image with pixel-wise semantic annotations. In autonomous driving, semantic segmentation is crucial for identifying and localizing various objects. By segmenting semantic objects, autonomous vehicles can extract detailed information about objects and their spatial relationships, such as pedestrians, vehicles, road signs, and obstacles, enabling the vehicle to make informed decisions and take appropriate actions.

[0003] When existing semantic segmentation methods based on urban traffic are applied to the field of autonomous driving, due to defects or limitations in the segmentation model, they can only use local information to understand the input image and have difficulty in obtaining deep global information. Specifically: First, object edge segmentation is poor. In semantic segmentation, the hierarchical structure of convolutional neural networks can effectively capture the global context of an image. However, because edge details typically only account for approximately 1%-5% of the total image pixels, as the receptive field expands, these edge details are easily diluted by the smoothing operations of low-level features. According to statistics from the Cityscapes dataset, edge pixels (the 2-pixel-wide area surrounding the object) only account for 1.8% of all annotated pixels in the image, but their misclassification rate is approximately 3.7 times that of non-edge regions. This indicates that while the attention mechanism can improve the overall mIoU, the improvement in edge regions is significantly lower than that in non-edge regions, indicating that existing methods still have limitations in modeling image edge features.

[0004] Second, cross-channel information exchange is inefficient. In semantic segmentation, traditional convolutional layers struggle to efficiently establish global dependencies between channels, making it difficult for the model to fully integrate complementary semantic information distributed across different channels. For example, while depthwise separable convolution reduces computational effort by approximately 50%-70% by decoupling spatial and channel processing, it sacrifices channel interaction capabilities.

[0005] In summary, existing semantic segmentation methods can lead to semantic information bias, especially in complex urban traffic scenarios such as multi-target occlusion and coexistence of cross-scale objects. The semantic information bias will be gradually amplified with the depth of the network, resulting in the inability of autonomous driving vehicles to accurately and clearly perceive surrounding objects, thereby making incorrect driving decisions and control commands, reducing the safety and reliability of autonomous driving. Summary of the Invention

[0006] The present invention aims to solve the problems existing in the prior art at least to a certain extent. Therefore, the present invention provides a semantic segmentation method with multi-scale attention and median enhancement. The specific technical solution is as follows.

[0007] A semantic segmentation method with multi-scale attention and median enhancement, comprising the following steps: S1, using ResNeSt50 as the backbone network to extract the initial features of the original image, splitting the initial features into multiple base groups through the split attention block, and each base group is further split into multiple subgroups; Based on the global context information of the original image, a weighted combination is performed on the subgroups corresponding to each basis array to obtain the feature representation of each basis array; Merge the feature representations of all base arrays to form the output features of the Inception block of the ResNeSt50 convolutional neural network; S2: Input the output features of the Inception block into the multi-scale feature enhancement module of the ResNeSt50 convolutional neural network to extract features and obtain 5 feature maps; splice the 5 feature maps according to the channel dimensions to obtain the multi-scale features of the original image; S3, inputting the multi-scale features into the median enhanced spatial channel attention module and outputting an attention feature map; S4, aggregates high-level features and low-level features through the cross-feature fusion block to generate the final fusion features of high-level features and low-level features; S5, after upsampling the attention feature map, adding it to the final fusion feature and upsampling it to obtain a fusion feature map; S6, upsampling the fused feature map, concatenating it with the middle-layer features of ResNeSt50 after 1×1 convolution, and outputting the concatenated feature map; In S7, the spliced ​​feature map undergoes a 3×3 convolution and is expanded through upsampling and a fully connected layer to obtain the segmentation prediction map of the original image.

[0008] Furthermore, step S1 includes the following specific steps: S101, inputting an original image into a ResNeSt50 convolutional neural network, scaling the pixel values ​​of the original image to [0, 1] by normalization, and standardizing according to the ImageNet mean, and padding the edges of the original image of non-standard size with zeros to a minimum size divisible by the step size; S102, extracting initial features of the original image through the initial convolution layer of the ResNeSt50 convolutional neural network; S103, performing feature grouping and split attention operation on the initial features through the split attention block to form output features of the Inception block of the ResNeSt50 convolutional neural network.

[0009] Furthermore, step S103 includes the following specific operation steps: dividing the initial features into a plurality of base groups, and each base group is further subdivided into a plurality of subgroups; Based on the global context information of the original image, the channel weights of the original image are output, and the weighted sum of the subgroups is performed to obtain the feature representation of each basis array; The feature representations of all base arrays are weighted and fused to form the output features of the Inception block of the ResNeSt50 convolutional neural network.

[0010] Furthermore, the channel weights of the original image are output through the following formula: , Where, For the The channel weights of the basis array; is the number of base arrays; is the Sigmoid function; is a fully connected layer; For average pooling: For the The characteristics of the base array; Furthermore, the weighted sum of multiple subgroup outputs of the same base array is performed using the following formula: , Where, For the The characteristics of the base array; is the number of base arrays; For the The first r Output features of subgroups; The number of subgroups for each base group; r The index of the subgroup, ranging from 1 to ; Furthermore, the feature representations of all basis arrays are weightedly fused using the following formula: , Where, is the weighted fusion result of all basis array feature representations; For the The channel weights of the basis array; For the The characteristics of the base array; kis the index of the base array, ranging from 1 to K .

[0011] Furthermore, step S3 includes the following specific steps: S301, performing global average pooling, global maximum pooling, and global median pooling on the multi-scale features using a channel attention mechanism to obtain three different pooling results of the multi-scale features; S302, performing nonlinear transformation and dimensionality reduction-increase processing on the three different pooling results of the multi-scale features through a shared multi-layer perceptron; S303, use Sigmoid The activation function compresses the output value of the multilayer perceptron into the range of [0, 1], resulting in three processed feature maps; S304: Add the three processed feature maps element-wise to obtain the initial channel attention feature. The formula is as follows: , Where, is the initial channel attention feature; is the Sigmoid function; is a multilayer perceptron; It is a multi-scale feature; S305: Multiply the initial channel attention feature and the multi-scale feature element-wise to output the final channel attention feature. The formula is: , where is the final channel attention feature; is element-wise multiplication; It is a multi-scale feature; S306, passing the final channel attention feature through multiple deep convolutional layers of different sizes to extract the basic features of the final channel attention feature; S307, the basic features of the final channel attention feature are added element-wise, and then fused with the feature map of the final channel attention feature after 5×5 convolution processing to obtain a spatial feature map. The formula is as follows: , Where, is the spatial feature map; N Depthwise convolutions of different sizes; Indicates the number of depthwise convolutions; is the final channel attention feature; S308, the spatial feature map is subjected to 1×1 convolution processing and element-wise multiplication with the final channel attention feature to output an attention feature map. The formula is: , where is the attention feature map; is the spatial feature map; is element-wise multiplication; is the final channel attention feature map.

[0012] Furthermore, step S4 includes the following specific steps: S401, a high-level feature of the ResNeSt50 convolutional neural network is upsampled to expand the spatial size to the same size as the low-level feature; S402: Use a 3×3 convolution kernel and a dilated convolution with a dilation rate of 2 to expand the receptive field of high-level features. S403, batch normalization is performed on the high-level features after the dilated convolution process to obtain optimized high-level features. The formula is: , where is the optimized high-level feature; It is batch normalization processing; It is a high-level feature of the ResNeSt50 convolutional neural network; It is a dilated convolution with a 3×3 convolution kernel and a dilation rate of 2; In S404, a low-level feature of the ResNeSt50 convolutional neural network has its number of channels adjusted by 1×1 convolution to match the number of channels of the high-level feature; S405, batch normalization is performed to further optimize the feature distribution after 1×1 convolution adjustment, and the optimized low-level features are obtained. The formula is: , where is the optimized low-level feature; It is batch normalization processing; It is a low-level feature of the ResNeSt50 convolutional neural network; S406, adding the optimized high-level features to the optimized low-level features, based on nonlinear ReLU Activation function, generating the final fusion feature, the formula is: , where is the final fusion feature; is the optimized low-level feature; is the optimized high-level feature.

[0013] Furthermore, in step S6, the spliced ​​feature map is outputted by the following formula: , Where, is the concatenated feature map; is the fusion feature map; It is a branch of the ResNeSt50 convolutional neural network.

[0014] Based on the above technical solution, the semantic segmentation method described in the present invention has the following beneficial effects: 1. In the semantic segmentation method described in the present invention, the split attention mechanism of the ResNeSt50 convolutional neural network dynamically assigns weights to feature subgroups through global context information, achieving effective information integration within feature groups and providing good basic features for subsequent information processing.

[0015] 2. The semantic segmentation method described in the present invention adopts a multi-scale feature enhancement module with a parallel multi-branch structure, such as 1x1 convolution, serial dilated convolutions with different void ratios, and strip pooling, which can efficiently capture contextual information of different scales, such as local details, medium-range semantics, long-range dependencies, and strip structures such as roads or poles, and enhance the model's adaptability to objects with large scale changes in traffic scenes, such as nearby vehicles and distant pedestrians, through splicing and fusion.

[0016] 3. The semantic segmentation method described in this invention can more comprehensively capture feature distribution information by introducing a median-enhanced spatial channel attention module, combined with global average pooling, maximum pooling, and median pooling. Median pooling is insensitive to noise and outliers, making it particularly suitable for addressing common issues such as blurred edges and uneven lighting in traffic scenes. It allows the network to focus more on true boundaries, effectively alleviating the ambiguity and fragmentation problems of existing methods in object contour segmentation, such as vehicle edges, pedestrian outlines, lane lines, traffic signs, and cones, thereby achieving clearer and more accurate edge segmentation results.

[0017] 4. The semantic segmentation method described in this invention selectively fuses features from different layers of the backbone network, such as high-level semantic features and low-level detail features, through a cross-feature fusion block. Through dilated convolutions to expand the receptive field, channel adjustment, feature addition, and nonlinear activation, high-level semantic information guides low-level features and low-level detail information complements high-level features, significantly improving fusion efficiency and feature quality.

[0018] 5. The semantic segmentation method described in the present invention has significantly improved the accuracy of some important scenarios compared to the commonly used DeepLabv3+ and PSPNet segmentation methods. The semantic segmentation method of the present invention and the DeepLabv3+ and PSPNet segmentation methods were all trained on a dual-card RTX 4060 Laptop GPU under the PyTorch framework. The training settings were: the batch size was set to 8, the maximum number of iterations was 120, the initial learning rate was 0.001, the "adam" optimizer was used for parameter update, the weight decay coefficient was 0.0001, and the "poly" learning rate strategy was used. A larger learning rate was used initially, and the learning rate was gradually reduced with the number of iterations. The following table shows the performance comparison of the semantic segmentation method of the present invention and the DeepLabv3+ and PSPNet segmentation methods in different scenarios: , It can be seen from the above table that the semantic segmentation method described in the present invention has outstanding advantages in various scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the overall framework of the method described in the present invention; Figure 2 Schematic diagram of the Block structure of ResNeSt50; Figure 3 Schematic diagram of the multi-scale feature enhancement module structure; Figure 4 Schematic diagram of the median-enhanced spatial channel attention module structure; Figure 5 Schematic diagram of the cross-feature fusion block structure; Figure 6 Schematic diagram of the segmentation effect of the original image. DETAILED DESCRIPTION

[0020] Before describing the specific embodiments, the following explanations or glossaries are given: 1. Certain terms are used in the specification and claims to refer to specific components. Those skilled in the art will understand that they may use different terms to refer to the same component. This specification and claims do not distinguish components based on differences in terms, but rather on differences in their functions. Unless otherwise defined, technical or scientific terms used in this disclosure should have the ordinary meanings understood by persons of ordinary skill in the art to which this disclosure pertains.

[0021] 2. A three-channel image refers to an image with three color channels: red, green, and blue. It is also called an RGB image, where RGB stands for Red, Green, and Blue.

[0022] 3. The original image in this embodiment refers to an image that has not been processed in any way or an original image taken by a camera installed on a vehicle.

[0023] 4. ResNeSt50 is a convolutional neural network structure. Res is the abbreviation of residual, and 50 means that there are 50 convolutional layers in the entire network.

[0024] The following is combined with Figure 1 To the attached Figure 6 , this embodiment is described in detail.

[0025] This embodiment describes a semantic segmentation method with multi-scale attention and median enhancement. Figure 1 The overall framework diagram of the semantic segmentation method is shown, which includes the following steps.

[0026] S1 uses ResNeSt50 as the backbone network to extract the initial features of the original image, and divides the initial features into multiple base groups through the split attention block. Each base group is further split into multiple subgroups. Figure 2 Figure 4 shows a schematic diagram of the Block structure of ResNeSt50.

[0027] Based on the global context information of the original image, weight evaluation is performed and the subgroups corresponding to each basis array are weighted combined to obtain the feature representation of each basis array.

[0028] The feature representations of all base arrays are merged to form the output features of the Inception block of the ResNeSt50 convolutional neural network.

[0029] The specific steps are as follows: S101, input the original image to the ResNeSt50 convolutional neural network. The original image can also be described as a three-channel image. The pixel values ​​of the original image are scaled to [0, 1] through normalization and standardized according to the ImageNet mean. The edges of the original image of non-standard size are padded with zeros to the minimum size that can be divided by the step size.

[0030] S102, through the initial convolution layer of the ResNeSt50 convolutional neural network, that is, the convolution layer of the 7×7 convolution kernel, reduces the spatial resolution of the original image, retains the basic texture features of the original image, and extracts the initial features of the original image.

[0031] For ease of description, this embodiment simply refers to the “initial features of the original image” as “initial features”.

[0032] S103, performing feature grouping and split attention operations on the initial features through the split attention block, including the following steps: Divide the initial features into multiple base groups, and each base group is further divided into multiple subgroups; Based on the global context information of the original image, the channel weights of the original image are output, and the weighted sum of the subgroups is performed to obtain the feature representation of each basis array; The feature representations of all base arrays are weighted and fused to form the output features of the Inception block of the ResNeSt50 convolutional neural network.

[0033] The specific formula for this step is as follows: ① Output the channel weight of the original image through the following formula: , Where, For the The channel weights of the basis array; is the number of base arrays; is the Sigmoid function; is a fully connected layer; For average pooling: For the The characteristics of the base array.

[0034] ② Perform weighted sum on multiple subgroup outputs of the same base array using the following formula: , Where, For the The characteristics of the base array; is the number of base arrays; For the The first r Output features of subgroups; The number of subgroups for each base group; r is the index of the subgroup, ranging from 1 to .

[0035] ③ Perform weighted fusion on the feature representations of all basis arrays using the following formula: , Where, is the weighted fusion result of all basis array feature representations; For the The channel weights of the basis array; For the The characteristics of the base array; k is the index of the base array, ranging from 1 to K .

[0036] S2, inputs the output features of the Inception block into the multi-scale feature enhancement module of the ResNeSt50 convolutional neural network; Figure 3 The figure shows the structure of the multi-scale feature enhancement module. The multi-scale feature enhancement module uses a parallel 1×1 convolution kernel, three serial dilated convolution kernels with dilation rates of 6, 12, and 18, and an SP strip pooling kernel to extract the output features of the Inception block in the ResNeSt50 convolutional neural network, obtaining five feature maps. The five feature maps are spliced ​​according to their channel dimensions to obtain the multi-scale features of the original image.

[0037] For the convenience of subsequent description, this embodiment simply refers to “multi-scale features of the original image” as “multi-scale features”.

[0038] S3 inputs the multi-scale features into the median-enhanced spatial channel attention module and outputs the attention feature map after the multi-scale features are processed by the median-enhanced spatial channel attention, so that the multi-scale features are optimized in both channel and spatial dimensions, improving the effect and robustness of feature extraction. Figure 4 The schematic diagram of the median enhanced spatial channel attention module structure is shown, which specifically includes the following steps: S301: Use the channel attention mechanism to perform global average pooling, global maximum pooling, and global median pooling on the multi-scale features to obtain three different pooling results of the multi-scale features; S302, the three different pooling results of the multi-scale features are subjected to nonlinear transformation and dimensionality reduction-increase processing through a shared multi-layer perceptron. Through dimensionality reduction-increase processing, the dependencies between channels can be automatically learned, thereby highlighting important features. Specifically: the multi-layer perceptron contains two 1×1 convolutional layers and one ReLU Activation function. The first 1×1 convolutional layer reduces the feature dimension of the multi-scale feature from C to C / n, where C is the feature dimension, n is the dimensionality reduction ratio; the second 1×1 convolutional layer restores the feature dimension of the multi-scale features to C.

[0039] S303, use Sigmoid The activation function compresses the output value of the multilayer perceptron into the range of [0, 1], and obtains three processed feature maps.

[0040] S304, perform element-wise addition of the three processed feature maps to obtain the initial channel attention features of the multi-scale features of the original image, as follows: , Where, is the initial channel attention feature; is the Sigmoid function; is a multilayer perceptron; It is a multi-scale feature.

[0041] For ease of description, this embodiment simply refers to the “initial channel attention features of the multi-scale features of the original image” as “initial channel attention features”.

[0042] S305, perform element-wise multiplication of the initial channel attention feature and the multi-scale feature to output the final channel attention feature of the multi-scale feature of the original image. The formula is: , Where, is the final channel attention feature; is element-wise multiplication; It is a multi-scale feature.

[0043] For ease of description, this embodiment simply refers to the “final channel attention feature of the multi-scale features of the original image” as the “final channel attention feature”.

[0044] S306, passing the final channel attention feature through a 5×5 deep convolutional layer to extract the basic features of the final channel attention feature, where the output size of the deep convolutional layer is the same as the input size; The basic features of the final channel attention feature are passed through multiple deep convolutional layers of different sizes, including 1×11, 1×7, etc., to further extract the basic features of the final channel attention feature.

[0045] S307, the basic features of the final channel attention features output by all deep convolutional layers in step S307 are added element-wise, and then fused with the feature map after the final channel attention features are processed by 5×5 convolution to obtain the final spatial feature map.

[0046] The formula is as follows: , Where, is the spatial feature map; N Depthwise convolutions of different sizes; Indicates the number of depthwise convolutions; is the final channel attention feature.

[0047] S308, the spatial feature map is subjected to 1×1 convolution processing and element-wise multiplication with the final channel attention feature to obtain an attention feature map after the multi-scale feature is subjected to median-enhanced spatial channel attention processing.

[0048] The formula is as follows: , Where, is the attention feature map; is the spatial feature map; is element-wise multiplication; is the final channel attention feature map.

[0049] S4, in order to achieve complementary advantages between multi-level features and effectively improve the long-range context learning capability of convolutional neural networks, selectively aggregates high-level semantic features, such as object categories, and low-level detail features, such as edges and textures, through the cross-feature fusion block. It also refines the feature maps of high-level and low-level features to generate the final fused features of high-level and low-level features.

[0050] For ease of description, this embodiment abbreviates “high-level semantic features” as “high-level features” and abbreviates “low-level detail features” as “low-level features”.

[0051] Figure 5 The cross-feature fusion block structure diagram is shown, which specifically includes the following steps: S401, a high-level feature of the ResNeSt50 convolutional neural network is upsampled to expand the spatial size to the same size as the low-level feature.

[0052] In S402, a 3×3 convolution kernel and a dilated convolution with a dilation rate of 2 are used to expand the receptive field of high-level features and capture broader contextual information.

[0053] S403: performing batch normalization on the high-level features after the dilated convolution processing to obtain optimized high-level features.

[0054] The formula is as follows: , Where, is the optimized high-level feature; It is batch normalization processing; It is a high-level feature of the ResNeSt50 convolutional neural network; It is a dilated convolution with a 3×3 convolution kernel and a dilation rate of 2.

[0055] In S404, a low-level feature of the ResNeSt50 convolutional neural network adjusts the number of channels through 1×1 convolution to match the number of channels of the high-level features.

[0056] S405: Perform batch normalization to further optimize the feature distribution after 1×1 convolution adjustment to obtain optimized low-level features.

[0057] The formula is as follows: , Where, is the optimized low-level feature; It is batch normalization processing; It is a low-level feature of the ResNeSt50 convolutional neural network.

[0058] S406, add the optimized high-level features to the optimized low-level features and introduce nonlinear ReLU The activation function generates the features after the fusion of high-level features and low-level features, which is defined as the final fusion feature.

[0059] The formula is as follows: , Where, is the final fusion feature; is the optimized low-level feature; is the optimized high-level feature; is a non-linear activation function.

[0060] S5, attention feature map after processing the median enhanced spatial channel attention module After upsampling, it is added to the cross feature fusion block and upsampled to obtain the fused feature map .

[0061] The formula is: , Where, is the fusion feature map; is the final fusion feature; is the attention feature map; S6, fusion feature map Then upsampled and the middle layer features of ResNeSt50 F 3 After 1×1 convolution processing, splicing is performed to obtain the spliced ​​feature map .

[0062] The formula is: , Where, is the concatenated feature map; is the fusion feature map; It is a branch of the ResNeSt50 convolutional neural network.

[0063] S7, the spliced ​​feature map is subjected to a 3×3 convolution and then expanded through upsampling and a fully connected layer to obtain the final segmentation prediction map of the original image. Figure 6 The figure shows the segmentation effect of the original image after the above steps.

[0064] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention as claimed.

Claims

1. A semantic segmentation method with multi-scale attention and median enhancement, characterized in that The following steps are involved: S1, using ResNeSt50 as the backbone network to extract the initial features of the original image, splitting the initial features into multiple base groups through the split attention block, and each base group is further split into multiple subgroups; Based on the global context information of the original image, a weighted combination is performed on the subgroups corresponding to each basis array to obtain the feature representation of each basis array; Merge the feature representations of all base arrays to form the output features of the Inception block of the ResNeSt50 convolutional neural network; S2: Input the output features of the Inception block into the multi-scale feature enhancement module of the ResNeSt50 convolutional neural network to extract features and obtain 5 feature maps; splice the 5 feature maps according to the channel dimensions to obtain the multi-scale features of the original image; S3, inputting the multi-scale features into the median enhanced spatial channel attention module and outputting an attention feature map; S4, aggregates high-level features and low-level features through the cross-feature fusion block to generate the final fusion features of high-level features and low-level features; S5, after upsampling the attention feature map, adding it to the final fusion feature and upsampling it to obtain a fusion feature map; S6, upsampling the fused feature map, concatenating it with the middle-layer features of ResNeSt50 after 1×1 convolution, and outputting the concatenated feature map; In S7, the spliced ​​feature map undergoes a 3×3 convolution and is expanded through upsampling and a fully connected layer to obtain the segmentation prediction map of the original image.

2. The semantic segmentation method with multi-scale attention and median enhancement according to claim 1, characterized in that: Step S1 includes the following specific steps: S101, inputting an original image into a ResNeSt50 convolutional neural network, scaling the pixel values ​​of the original image to [0, 1] by normalization, and standardizing according to the ImageNet mean, and padding the edges of the original image of non-standard size with zeros to a minimum size divisible by the step size; S102, extracting initial features of the original image through the initial convolution layer of the ResNeSt50 convolutional neural network; S103, performing feature grouping and split attention operation on the initial features through the split attention block to form output features of the Inception block of the ResNeSt50 convolutional neural network.

3. The semantic segmentation method with multi-scale attention and median enhancement according to claim 2, characterized in that: Step S103 includes the following specific steps: Divide the initial features into multiple base groups, and each base group is further divided into multiple subgroups; Based on the global context information of the original image, the channel weights of the original image are output, and the weighted sum of the subgroups is performed to obtain the feature representation of each basis array; The feature representations of all base arrays are weighted and fused to form the output features of the Inception block of the ResNeSt50 convolutional neural network.

4. The semantic segmentation method with multi-scale attention and median enhancement according to claim 3, characterized in that The channel weights of the original image are output using the following formula: , Where, For the The channel weights of the basis array; is the number of base arrays; is the Sigmoid function; is a fully connected layer; For average pooling: For the The characteristics of the base array; The weighted sum of multiple subgroup outputs of the same base array is calculated using the following formula: , Where, For the The characteristics of the base array; is the number of base arrays; For the The first r Output features of subgroups; The number of subgroups for each base group; r The index of the subgroup, ranging from 1 to ; The feature representations of all basis arrays are weightedly fused using the following formula: , Where, is the weighted fusion result of all basis array feature representations; For the The channel weights of the basis array; For the The characteristics of the base array; k is the index of the base array, ranging from 1 to K .

5. The semantic segmentation method with multi-scale attention and median enhancement according to claim 1, characterized in that Step S3 includes the following specific steps: S301, performing global average pooling, global maximum pooling, and global median pooling on the multi-scale features using a channel attention mechanism to obtain three different pooling results of the multi-scale features; S302, performing nonlinear transformation and dimensionality reduction-increase processing on the three different pooling results of the multi-scale features through a shared multi-layer perceptron; S303, use Sigmoid The activation function compresses the output value of the multilayer perceptron into the range of [0, 1], resulting in three processed feature maps; S304, perform element-wise addition of the three processed feature maps to obtain the initial channel attention feature, the formula is as follows: , Where, is the initial channel attention feature; is the Sigmoid function; is a multilayer perceptron; It is a multi-scale feature; S305: Perform element-wise multiplication of the initial channel attention feature and the multi-scale feature to output the final channel attention feature. The formula is: , where is the final channel attention feature; is element-wise multiplication; It is a multi-scale feature; S306, passing the final channel attention feature through multiple deep convolutional layers of different sizes to extract the basic features of the final channel attention feature; S307, the basic features of the final channel attention feature are added element-wise, and then fused with the feature map of the final channel attention feature after 5×5 convolution processing to obtain a spatial feature map. The formula is as follows: , Where, is the spatial feature map; N Depthwise convolutions of different sizes; Indicates the number of depthwise convolutions; is the final channel attention feature; S308, the spatial feature map is subjected to 1×1 convolution processing and element-wise multiplication with the final channel attention feature to output an attention feature map. The formula is: , where is the attention feature map; is the spatial feature map; is element-wise multiplication; is the final channel attention feature map.

6. The semantic segmentation method with multi-scale attention and median enhancement according to claim 1, characterized in that Step S4 includes the following specific steps: S401, a high-level feature of the ResNeSt50 convolutional neural network is upsampled to expand the spatial size to the same size as the low-level feature; S402: Use a 3×3 convolution kernel and a dilated convolution with a dilation rate of 2 to expand the receptive field of high-level features. S403, batch normalization is performed on the high-level features after the dilated convolution process to obtain optimized high-level features. The formula is: , where is the optimized high-level feature; It is batch normalization processing; It is a high-level feature of the ResNeSt50 convolutional neural network; It is a dilated convolution with a 3×3 convolution kernel and a dilation rate of 2; In S404, a low-level feature of the ResNeSt50 convolutional neural network has its number of channels adjusted by 1×1 convolution to match the number of channels of the high-level feature; S405, batch normalization is performed to further optimize the feature distribution after 1×1 convolution adjustment, and the optimized low-level features are obtained. The formula is: , where is the optimized low-level feature; It is batch normalization processing; It is a low-level feature of the ResNeSt50 convolutional neural network; S406, adding the optimized high-level features to the optimized low-level features, based on nonlinear ReLU Activation function, generating the final fusion feature, the formula is: , where is the final fusion feature; is the optimized low-level feature; is the optimized high-level feature.

7. The semantic segmentation method with multi-scale attention and median enhancement according to claim 1, characterized in that In step S6, the concatenated feature map is output using the following formula: , Where, is the concatenated feature map; is the fusion feature map; It is a branch of the ResNeSt50 convolutional neural network.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on attention multi-scale feature fusion

    CN111127493A

  • 2D human body posture estimation method based on multi-scale feature enhancement

    CN112131959A

  • Remote sensing image semantic segmentation method based on spatial detail perception and attention guidance

    CN117274608A

  • Expression recognition method and system based on multi-scale features and spatial attention

    US12354405B1

Cited By

  • Fire test powder mixing uniformity identification method and system

    CN121392375A

  • Reverse recovery period protection device based on high-voltage direct-current power transmission system

    CN121546507A

  • A reverse recovery period protection device based on a high voltage direct current transmission system

    CN121546507B