Image segmentation method and device, electronic equipment, storage medium and program product

By combining lightweight convolutional neural network, dual attention mechanism network and ASPP network, the problem of low image segmentation efficiency in road scenes in the prior art is solved, and efficient image segmentation effect is achieved.

CN120182290APending Publication Date: 2025-06-20CHINA MOBILE M2M +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510176654.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing image semantic segmentation model has poor efficiency when processing images in road scenes, and the coordination efficiency between various modules is not high.

Method used

The combination of preset lightweight convolutional neural network, dual attention mechanism network and hollow convolutional space pyramid ASPP network is used to optimize the image segmentation model through initial feature extraction, location feature and channel feature fusion, and context feature extraction at different scales.

Benefits of technology

The segmentation efficiency and accuracy of the image segmentation model in road scenes is improved, the system resource consumption is reduced, and the efficient understanding of road scene images is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182290A_ABST
    Figure CN120182290A_ABST
Patent Text Reader

Abstract

The invention provides an image segmentation method and device, electronic equipment, a storage medium and a program product, relates to the technical field of image processing, and can accurately and efficiently complete extraction of initial features of a to-be-segmented image through a small number of parameters and a centralized calculation process by presetting a lightweight convolutional neural network. According to the invention, through the parallel double attention mechanism network and ASPP network, feature extraction of important areas and key categories of initial features is realized, optimization of object segmentation boundaries in subsequent images to be segmented is facilitated, and segmentation precision is improved. According to the method, the lightweight convolutional neural network, the double attention mechanism network and the ASPP network are preset, so that understanding of the image content of the road scene in the to-be-segmented image is realized, and improvement of the efficiency of the image of the road scene is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to an image segmentation method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] When a computer information processing system is applied to a road scene, it is necessary to continuously collect images of the road and its surrounding environment, people, buildings, etc. When traditional algorithms analyze and process these images with different categories, sizes, no structure, and no rules, problems such as low real-time performance and low accuracy will occur. Most existing lightweight models generally only discuss application scenarios with single uses such as portrait segmentation, and do not involve application scenarios with high complexity and multiple influencing factors such as road scenes.

[0003] In existing image semantic segmentation models for road scenes, structures such as MobileNetV2, Efficient Channel Attention (ECA), Convolutional Block Attention Module (CBAM), and Squeeze-and-Excitation (SE) attention mechanisms are usually introduced. However, in existing image semantic segmentation models, the relative positions and parameter settings between each module are not reasonable, resulting in poor cooperation efficiency among the modules in the existing image semantic segmentation models. Therefore, the existing image semantic segmentation models have poor efficiency when processing images of road scenes. Summary of the Invention

[0004] The present invention provides an image segmentation method, apparatus, electronic device, storage medium, and program product to solve the defect that the existing image semantic segmentation model has poor efficiency when processing images of road scenes in the prior art, and to improve the efficiency of the image semantic segmentation model when processing images of road scenes.

[0005] In a first aspect, the present invention provides an image segmentation method, including: inputting an image to be segmented into an image segmentation model, and obtaining a segmentation result of the image to be segmented output by the image segmentation model; wherein, the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network, and an Atrous Spatial Pyramid Pooling (ASPP) network. The preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain initial features of the image to be segmented, and inputs the initial features into the dual attention mechanism network and the ASPP network respectively. The dual attention mechanism network extracts and fuses position features and channel features of the initial features to obtain fused attention features, and the ASPP network extracts context features of the initial features to obtain context features. The fused attention features, context features, and initial features are fused to obtain a segmentation result.

[0006] In one embodiment, the preset lightweight convolutional neural network includes a convolutional layer, a first inverted residual layer, and a second inverted residual layer using dilated convolution. The preset lightweight convolutional neural network is used to obtain initial features: based on the convolutional layer, feature extraction is performed on the image to be segmented to obtain a convolutional feature map; based on the first inverted residual layer, feature extraction is performed on the convolutional feature map to obtain an inverted residual feature map; based on the second inverted residual layer, dilated convolution feature extraction is performed on the inverted residual feature map to obtain initial features.

[0007] In one embodiment, the dual attention mechanism network is used to obtain fused attention features: the initial features are sequentially subjected to convolution, first-dimensional transformation, and transposition to obtain a first position feature matrix; the initial features are sequentially subjected to convolution and second-dimensional transformation to obtain a second position feature matrix; the initial features are sequentially subjected to convolution and third-dimensional transformation to obtain a third position feature matrix; feature fusion is performed on the first position feature matrix, the second position feature matrix, the third position feature matrix, and the initial features to obtain a position attention matrix; the initial features are sequentially subjected to fourth-dimensional transformation and transposition to obtain a first channel feature matrix; the initial features are subjected to fourth-dimensional transformation to obtain a second channel feature matrix; the initial features are subjected to fifth-dimensional transformation to obtain a third channel feature matrix; feature fusion is performed on the first channel feature matrix, the second channel feature matrix, the third channel feature matrix, and the initial features to obtain a channel attention matrix; matrix addition is performed on the position attention matrix and the channel attention matrix to obtain fused attention features.

[0008] In one embodiment, the ASPP network includes multiple dilated convolutional layers with different dilation factors. The ASPP network is used to obtain context features: based on the multiple dilated convolutional layers with different dilation factors, feature extraction of different scales is performed on the initial features to obtain multiple initial context features of different scales; feature fusion and feature extraction are performed on the multiple initial context features of different scales to obtain context features.

[0009] In one embodiment, the image segmentation model is used to obtain a segmentation result: feature fusion by matrix addition is performed on the fused attention features and the context features to obtain a first initial feature map; the first initial feature map is upsampled to obtain a second initial feature map; the second initial feature map and the initial features are concatenated to obtain a third initial feature map; convolution and upsampling are performed on the third initial feature map to restore the third initial feature map to the size of the image to be segmented, obtaining a segmentation result.

[0010] In one embodiment, feature fusion is performed on a first position feature matrix, a second position feature matrix, a third position feature matrix, and an initial feature to obtain a position attention matrix, including: multiplying and activating the first position feature matrix and the second position feature matrix to obtain a first position fusion matrix; multiplying the first position fusion matrix and the third position feature matrix to obtain a second position fusion matrix; adding the second position fusion matrix and the initial feature to obtain the position attention matrix.

[0011] In one embodiment, feature fusion is performed on a first channel feature matrix, a second channel feature matrix, a third channel feature matrix, and an initial feature to obtain a channel attention matrix, including: multiplying and activating the first channel feature matrix and the second channel feature matrix to obtain a first channel fusion matrix; multiplying the first channel fusion matrix and the third channel feature matrix to obtain a second channel fusion matrix; adding the second channel fusion matrix and the initial feature to obtain the channel attention matrix.

[0012] In a second aspect, the present invention provides an image segmentation device, including: a segmentation module, configured to input an image to be segmented into an image segmentation model and obtain a segmentation result of the image to be segmented output by the image segmentation model; wherein, the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network, and an atrous spatial pyramid pooling (ASPP) network. The preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain an initial feature of the image to be segmented, and inputs the initial feature into the dual attention mechanism network and the ASPP network respectively. The dual attention mechanism network extracts and fuses position features and channel features of the initial feature to obtain a fused attention feature, and the ASPP network extracts context features of the initial feature to obtain context features. The fused attention feature, the context features, and the initial feature are fused to obtain the segmentation result.

[0013] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for image segmentation as described in any one of the above is implemented.

[0014] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for image segmentation as described in any one of the above is implemented.

[0015] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for image segmentation as described in any one of the above is implemented.

[0016] The image segmentation method, device, electronic device, storage medium, and program product provided by the present invention can accurately and efficiently extract the initial features of the image to be segmented through a preset lightweight convolutional neural network with a small number of parameters and a concentrated calculation process. The present invention realizes the extraction of features of important regions and key categories of the initial features through a parallel dual-attention mechanism network and an ASPP network, which is beneficial to optimizing the object segmentation boundary in the subsequent image to be segmented and improving the segmentation accuracy. The present invention realizes the understanding of the image content of the road scene in the image to be segmented through a preset lightweight convolutional neural network, a dual-attention mechanism network, and an ASPP network, which is beneficial to improving the efficiency of the road scene image. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is one of the flowchart diagrams of the image segmentation method provided by the present invention.

[0019] Figure 2 is the second flowchart diagram of the image segmentation method provided by the present invention.

[0020] Figure 3 is the second flowchart diagram of the image segmentation method provided by the present invention.

[0021] Figure 4 is the flowchart diagram of the ASPP network for obtaining context features provided by the present invention.

[0022] Figure 5 is the structural diagram of the image segmentation device provided by the present invention.

[0023] Figure 6 is the structural diagram of the electronic device provided by the present invention. Detailed Embodiments

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0025] The following will be combined with Figures 1-6Describe an image segmentation method, apparatus, and electronic device according to the present invention.

[0026] Figure 1 is one of the schematic flowcharts of the image segmentation method provided by the present invention. As Figure 1 shown, the image segmentation method includes step S100.

[0027] S100: Input the image to be segmented into an image segmentation model, and obtain the segmentation result of the image to be segmented output by the image segmentation model; wherein, the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network, and an atrous spatial pyramid pooling (ASPP) network. The preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain the initial features of the image to be segmented, and inputs the initial features into the dual attention mechanism network and the ASPP network respectively. The dual attention mechanism network extracts and fuses the position features and channel features of the initial features to obtain the fused attention features. The ASPP network extracts the context features of the initial features to obtain the context features, and fuses the fused attention features, the context features, and the initial features to obtain the segmentation result.

[0028] The preset lightweight convolutional neural network of the present invention includes an improved MobileNetV2 network. The MobileNet series of networks mainly uses depthwise separable convolutions. Among them, MobileNetV2 has an inverted residual structure with a linear bottleneck, which further improves the network performance. Therefore, MobileNetV2 is more lightweight in structure than the extremely convolutional network (Xception) used in DeepLabV3+, and the calculation process is more concentrated. Applying it to road scenarios can reduce the system resource consumption by about 10 times while maintaining the segmentation accuracy.

[0029] For example, the complete MobileNetV2 network has 11 layers of structure. In order to further improve the real-time performance, the present invention simplifies and improves the MobileNetV2 network to obtain a preset lightweight convolutional neural network, so that the preset lightweight convolutional neural network is more suitable for image segmentation (including image segmentation of road scenarios).

[0030] The image segmentation model combines the improved MobileNetV2 network (preset lightweight convolutional neural network) with a parallel feature processing structure composed of an atrous spatial pyramid pooling (ASPP) network and a dual attention mechanism network (DAM) to complete the feature extraction of the image.

[0031] The parallel connection of the DAM network and the ASPP network can maximize the characteristics of each module. While retaining the ability of the ASPP module to capture image information at multiple scales, it can also selectively aggregate features at each position and redistribute resources between convolutional channels to achieve better segmentation accuracy. Table 1 shows the result comparison of different connection methods of the two modules.

[0032] Table 1 Influence of the connection method of the DAM network and the ASPP network on the segmentation result

[0033] As can be seen from Table 1, when the ASPP network and the DAM network are connected in parallel, the mean intersection over union (mIoU) of the image segmentation model is optimal.

[0034] A preset lightweight convolutional neural network is used to extract the initial features of the image to be segmented. The initial features are respectively sent to the dual attention mechanism network and the ASPP network. The dual attention mechanism network extracts and fuses the spatial attention features and channel attention features of the initial features to obtain the fused attention features with both spatial and channel features. The ASPP network extracts context features of different sizes from the initial features to obtain context features. The fused attention features, context features, and initial features are fused to obtain the segmentation result.

[0035] Furthermore, the image segmentation model is trained based on the preset model, using the sample image to be segmented and the label of the sample segmentation result of the sample image to be segmented. The preset model is constructed according to the improved MobileNetV2 network, the ASPP network, and the dual attention mechanism network.

[0036] The image segmentation method provided by the embodiments of the present invention can accurately and efficiently extract the initial features of the image to be segmented through the preset lightweight convolutional neural network with a small number of parameters and a concentrated calculation process. The present invention realizes the extraction of features of important regions and key categories of the initial features through the parallel dual attention mechanism network and the ASPP network, which is beneficial to optimizing the object segmentation boundary in the subsequent image to be segmented and improving the segmentation accuracy. The present invention realizes the understanding of the image content of the road scene in the image to be segmented through the preset lightweight convolutional neural network, the dual attention mechanism network, and the ASPP network, which is beneficial to improving the efficiency of the image of the road scene.

[0037] Based on the above embodiments, the preset lightweight convolutional neural network includes a convolutional layer, multiple first inverted residual layers, and a second inverted residual layer using dilated convolution. The preset lightweight convolutional neural network is used to obtain initial features: based on the convolutional layer, feature extraction is performed on the image to be segmented to obtain a convolutional feature map; based on the multiple first inverted residual layers, feature extraction is performed on the convolutional feature map multiple times to obtain an inverted residual feature map; based on the second inverted residual layer, dilated convolution feature extraction is performed on the inverted residual feature map to obtain initial features.

[0038] The preset lightweight convolutional neural network (improved MobileNetV2 network) of the present invention includes a convolutional layer, multiple first inverted residual layers, and a second inverted residual layer using dilated convolution. For example, the improved MobileNetV2 network includes 1 convolutional layer, 5 first inverted residual layers, and 2 second inverted residual layers using dilated convolution. The specific structure of the preset lightweight convolutional neural network (improved MobileNetV2 network) is shown in Table 2.

[0039] Table 2 Specific structure of the preset lightweight convolutional neural network

[0040] Performing initial feature extraction on the image to be segmented according to the preset lightweight convolutional neural network includes the following steps.

[0041] (1) Input the image to be segmented into the image segmentation model, and through the operation of a convolutional layer with a size of 3x3, a stride of 2, and 32 convolutional kernels, a convolutional feature map with an output channel number of 32 is obtained.

[0042] (2) Input the convolutional feature map into a first inverted residual layer with a stride of 1 to obtain a first initial inverted residual feature map with an output channel number of 16.

[0043] (3) Input the first initial inverted residual feature map into two first inverted residual layers with a stride of 2 to obtain a second initial inverted residual feature map with an output channel number of 24.

[0044] (4) Input the second initial inverted residual feature map into three first inverted residual layers with a stride of 2 to obtain a third initial inverted residual feature map with an output channel number of 32.

[0045] (5) Input the third initial inverted residual feature map into four first inverted residual layers with a stride of 2 to obtain a fourth initial inverted residual feature map with an output channel number of 64.

[0046] (6) Input the fourth initial inverted residual feature map into three first inverted residual layers with a stride of 1 to obtain an inverted residual feature map with an output channel number of 96.

[0047] (7) Input the inverted residual feature map into three second inverted residual layers with a stride of 2 and using dilated convolution to obtain a dilated convolution feature map with an output channel number of 160.

[0048] (8) Input the dilated convolution feature map into a second inverted residual layer with a stride of 1 and using dilated convolution to obtain an initial feature with an output channel number of 320.

[0049] The present invention obtains the initial feature through a convolutional layer, a first inverted residual layer, and a second inverted residual layer using dilated convolution, and can accurately and efficiently complete the extraction of the initial feature with a small number of parameters and a concentrated calculation process, which is beneficial to improving the efficiency of subsequent image segmentation of the image to be segmented.

[0050] Based on the above embodiments, the dual attention mechanism network is used to obtain the fused attention feature: perform convolution, first-dimensional transformation, and transposition on the initial feature in sequence to obtain a first position feature matrix; perform convolution and second-dimensional transformation on the initial feature in sequence to obtain a second position feature matrix; perform convolution and third-dimensional transformation on the initial feature in sequence to obtain a third position feature matrix; perform feature fusion on the first position feature matrix, the second position feature matrix, the third position feature matrix, and the initial feature to obtain a position attention matrix; perform fourth-dimensional transformation and transposition on the initial feature in sequence to obtain a first channel feature matrix; perform fourth-dimensional transformation on the initial feature to obtain a second channel feature matrix; perform fifth-dimensional transformation on the initial feature to obtain a third channel feature matrix; perform feature fusion on the first channel feature matrix, the second channel feature matrix, the third channel feature matrix, and the initial feature to obtain a channel attention matrix; perform matrix addition on the position attention matrix and the channel attention matrix to obtain the fused attention feature.

[0051] The dual attention mechanism network includes a position attention module (Position Attention Module, PAM) and a channel attention module (Channel Attention Module, CAM). After the initial feature enters the dual attention mechanism network, it will be sent to PAM and CAM respectively to extract channel feature information and position feature information.

[0052] Based on the above embodiments, perform feature fusion on the first position feature matrix, the second position feature matrix, the third position feature matrix, and the initial feature to obtain a position attention matrix. Specifically, perform matrix multiplication and activation on the first position feature matrix and the second position feature matrix to obtain a first position fusion matrix; perform matrix multiplication on the first position fusion matrix and the third position feature matrix to obtain a second position fusion matrix; perform matrix addition on the second position fusion matrix and the initial feature to obtain a position attention matrix.

[0053] Such as Figure 2 AndFigure 3 As shown, the position attention module successively performs convolution, first - dimension transformation, and transposition on the initial feature to obtain the first position feature matrix. The position attention module successively performs convolution and second - dimension transformation on the initial feature to obtain the second position feature matrix. The position attention module successively performs convolution and third - dimension transformation on the initial feature to obtain the third position feature matrix. The first - dimension transformation, second - dimension transformation, and third - dimension transformation are dimension transformations of different dimensions.

[0054] The position attention module performs matrix multiplication and activation on the first position feature matrix and the second position feature matrix to obtain the first position fusion matrix. The position attention module performs matrix multiplication on the first position fusion matrix and the third position feature matrix to obtain the second position fusion matrix. The position attention module performs matrix addition on the second position fusion matrix and the initial feature to obtain the position attention matrix.

[0055] The present invention realizes the accurate acquisition of the position attention matrix through convolution feature extraction, different - dimension transformation, and feature fusion of the initial feature.

[0056] Based on the above - mentioned embodiment, feature fusion is performed on the first channel feature matrix, the second channel feature matrix, and the initial feature to obtain the channel attention matrix. Specifically, matrix multiplication and activation are performed on the first channel feature matrix and the second channel feature matrix to obtain the first channel fusion matrix; matrix multiplication is performed on the first channel fusion matrix and the second channel feature matrix to obtain the second channel fusion matrix; matrix addition is performed on the second channel fusion matrix and the initial feature to obtain the channel attention matrix.

[0057] As Figure 2 and Figure 3 shown, the channel attention module successively performs dimension transformation and transposition on the initial matrix to obtain the first channel feature matrix. The channel attention module performs dimension transformation on the initial feature to obtain the second channel feature matrix.

[0058] The channel attention module performs matrix multiplication and activation on the first channel feature matrix and the second channel feature matrix to obtain the first channel fusion matrix. The channel attention module performs matrix multiplication on the first channel fusion matrix and the second channel feature matrix to obtain the second channel fusion matrix. The channel attention module performs matrix addition on the second channel fusion matrix and the initial feature to obtain the channel attention matrix.

[0059] The dual - attention network mechanism performs matrix addition on the position attention matrix and the channel attention matrix to obtain the fused attention feature.

[0060] In the embodiments of the present invention, the initial features are convolved and dimensionally transformed in different dimensions through the position attention module to obtain multiple position feature maps, achieving accurate extraction of the position features of the initial features. In the embodiments of the present invention, the channel attention module is used to accurately extract the channel features of the initial features.

[0061] Based on the above embodiments, the ASPP network includes multiple dilated convolutional layers with different dilation factors. The ASPP network is used to obtain context features: based on the multiple dilated convolutional layers with different dilation factors, feature extraction at different scales is performed on the initial features to obtain multiple initial context features at different scales; feature fusion and feature extraction are performed on the multiple initial context features at different scales to obtain context features.

[0062] The ASPP network includes multiple dilated convolutional layers with different dilation factors. For example, the ASPP network includes a dilated convolutional layer with a dilation factor (rate) of 1, a dilated convolutional layer with a rate of 6, a dilated convolutional layer with a rate of 12, and a dilated convolutional layer with a rate of 18.

[0063] As Figure 4 shown, the ASPP network performs global average pooling processing, dilated convolutional processing with a 1×1 convolutional kernel and a dilation factor of 1, and upsampling on the initial features to obtain the first initial context feature. The ASPP network performs dilated convolutional processing with a 1×1 convolutional kernel and a dilation factor of 1 on the initial features to obtain the second initial context feature. The ASPP network performs dilated convolutional processing with a 3×3 convolutional kernel and a dilation factor of 6 on the initial features to obtain the third initial context feature. The ASPP network performs dilated convolutional processing with a 3×3 convolutional kernel and a dilation factor of 12 on the initial features to obtain the fourth initial context feature. The ASPP network performs dilated convolutional processing with a 3×3 convolutional kernel and a dilation factor of 18 on the initial features to obtain the fifth initial context feature. The first initial context feature, the second initial context feature, the third initial context feature, the fourth initial context feature, and the fifth initial context feature are initial context features at different scales. Feature fusion and convolutional processing with a 1×1 convolutional kernel (feature extraction) are performed on the first initial context feature, the second initial context feature, the third initial context feature, the fourth initial context feature, and the fifth initial context feature to obtain context features.

[0064] The dilation factor (rate) of the ASPP network also directly affects the distance between elements in the dilated convolution, thereby affecting the segmentation accuracy. Moreover, the number of parameters of the image segmentation model is also affected by the number of channels in the ASPP network. Therefore, in order to reduce consumption and improve real-time performance, the present invention appropriately adjusts the combination of the dilation factor (rate) and the number of channels in the ASPP network: the dilation factor (rate) is set to 1, 6, 12, 18, and the number of channels is 128. Table 3 shows the performance comparison of different dilation factors of the ASPP network.

[0065] Table 3 Performance Comparison of Different Dilation Factors of the ASPP Network

[0066] As can be seen from Table 3, when the combination of the dilation factors of the ASPP network is 1 + 6 + 12 + 18, the mean intersection over union (mIoU) of the ASPP network is optimal.

[0067] The present invention extracts features from the initial features by combining dilated convolution layers with different dilation factors, obtaining initial context features of different scales, and realizing comprehensive context feature extraction of the initial features.

[0068] Based on the above embodiments, the image segmentation model is used to obtain a segmentation result: perform feature fusion by matrix addition of the fused attention feature and the context feature to obtain a first initial feature map; perform upsampling on the first initial feature map to obtain a second initial feature map; splice the second initial feature map and the initial feature to obtain a third initial feature map; perform convolution and upsampling on the third initial feature map to restore the third initial feature map to the size of the image to be segmented, obtaining the segmentation result.

[0069] As Figure 2 shown, perform feature fusion by matrix addition of the fused attention feature and the context feature to obtain a first initial feature map that simultaneously includes position attention features, channel attention features, and context features. Perform bilinear interpolation upsampling on the first initial feature map to obtain a second initial feature map with the same size as the initial feature. Splice the second initial feature map and the initial feature to obtain a third initial feature map. Perform convolution and bilinear interpolation upsampling on the third initial feature map to restore the third initial feature map to the size of the image to be segmented, obtaining the segmentation result.

[0070] The present invention obtains the segmentation result by fusing the fused attention feature and the context feature again, realizing the joint analysis of the position attention feature, the channel attention feature, and the context feature, which is beneficial to improving the accuracy of the segmentation result.

[0071] The image segmentation device provided by the present invention will be described below. The image segmentation device described below can be correspondingly referred to the image segmentation method described above.

[0072] As Figure 5 shown, an image segmentation device includes: a segmentation module 501, configured to input an image to be segmented into an image segmentation model, and obtain a segmentation result of the image to be segmented output by the image segmentation model; wherein, the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network, and an atrous spatial pyramid pooling (ASPP) network. The preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain initial features of the image to be segmented, and inputs the initial features into the dual attention mechanism network and the ASPP network respectively. The dual attention mechanism network extracts and fuses position features and channel features of the initial features to obtain fused attention features. The ASPP network extracts context features of the initial features to obtain context features, and fuses the fused attention features, context features, and initial features to obtain a segmentation result.

[0073] The image segmentation device provided by the embodiments of the present invention can accurately and efficiently complete the extraction of the initial features of the image to be segmented with a small number of parameters and a concentrated calculation process through the preset lightweight convolutional neural network. The present invention realizes the extraction of features of important regions and key categories of the initial features through the parallel dual attention mechanism network and the ASPP network, which is beneficial to optimizing the object segmentation boundary in the subsequent image to be segmented and improving the segmentation accuracy. The present invention realizes the understanding of the image content of the road scene in the image to be segmented through the preset lightweight convolutional neural network, the dual attention mechanism network, and the ASPP network, which is beneficial to improving the efficiency of the image of the road scene.

[0074] In one embodiment, the preset lightweight convolutional neural network includes a convolutional layer, a first inverted residual layer, and a second inverted residual layer using atrous convolution. The segmentation module 501 is configured to: based on the convolutional layer, perform feature extraction on the image to be segmented to obtain a convolutional feature map; based on the first inverted residual layer, perform feature extraction on the convolutional feature map to obtain an inverted residual feature map; based on the second inverted residual layer, perform atrous convolution feature extraction on the inverted residual feature map to obtain initial features.

[0075] In one embodiment, the segmentation module 501 is configured to: successively perform convolution, first - dimensional transformation, and transpose on the initial features to obtain a first position feature matrix; successively perform convolution and second - dimensional transformation on the initial features to obtain a second position feature matrix; successively perform convolution and third - dimensional transformation on the initial features to obtain a third position feature matrix; perform feature fusion on the first position feature matrix, the second position feature matrix, the third position feature matrix, and the initial features to obtain a position attention matrix; successively perform fourth - dimensional transformation and transpose on the initial features to obtain a first channel feature matrix; perform fourth - dimensional transformation on the initial features to obtain a second channel feature matrix; perform fifth - dimensional transformation on the initial features to obtain a third channel feature matrix; perform feature fusion on the first channel feature matrix, the second channel feature matrix, the third channel feature matrix, and the initial features to obtain a channel attention matrix; perform matrix addition on the position attention matrix and the channel attention matrix to obtain a fused attention feature.

[0076] In one embodiment, the ASPP network includes multiple dilated convolutional layers with different dilation factors. The segmentation module 501 is configured to: based on the multiple dilated convolutional layers with different dilation factors, perform feature extraction of different scales on the initial features to obtain multiple initial context features of different scales; perform feature fusion and feature extraction on the multiple initial context features of different scales to obtain context features.

[0077] In one embodiment, the segmentation module 501 is configured to: perform feature fusion by matrix addition on the fused attention feature and the context feature to obtain a first initial feature map; perform upsampling on the first initial feature map to obtain a second initial feature map; splice the second initial feature map and the initial features to obtain a third initial feature map; perform convolution and upsampling on the third initial feature map to restore the third initial feature map to the size of the image to be segmented, and obtain a segmentation result.

[0078] In one embodiment, the segmentation module 501 is configured to: perform matrix multiplication and activation on the first position feature matrix and the second position feature matrix to obtain a first position fusion matrix; perform matrix multiplication on the first position fusion matrix and the third position feature matrix to obtain a second position fusion matrix; perform matrix addition on the second position fusion matrix and the initial features to obtain a position attention matrix.

[0079] In one embodiment, the segmentation module 501 is configured to: perform matrix multiplication and activation on the first channel feature matrix and the second channel feature matrix to obtain a first channel fusion matrix; perform matrix multiplication on the first channel fusion matrix and the third channel feature matrix to obtain a second channel fusion matrix; perform matrix addition on the second channel fusion matrix and the initial features to obtain a channel attention matrix.

[0080] Figure 6Illustrates a schematic diagram of the physical structure of an electronic device, as follows Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call logic instructions in the memory 630 to execute an image segmentation method, which includes: inputting an image to be segmented into an image segmentation model, and obtaining a segmentation result of the image to be segmented output by the image segmentation model; wherein, the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network, and an atrous spatial pyramid pooling (ASPP) network. The preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain initial features of the image to be segmented, and inputs the initial features into the dual attention mechanism network and the ASPP network respectively. The dual attention mechanism network extracts and fuses position features and channel features of the initial features to obtain fused attention features, and the ASPP network extracts context features of the initial features to obtain context features. The fused attention features, the context features, and the initial features are fused to obtain a segmentation result.

[0081] In addition, when the logic instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0082] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image segmentation method provided by each of the above methods. The method includes: inputting an image to be segmented into an image segmentation model, and obtaining a segmentation result of the image to be segmented output by the image segmentation model; wherein, the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network, and an atrous spatial pyramid pooling (ASPP) network. The preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain initial features of the image to be segmented, and inputs the initial features into the dual attention mechanism network and the ASPP network respectively. The dual attention mechanism network extracts and fuses position features and channel features of the initial features to obtain fused attention features, and the ASPP network extracts context features of the initial features to obtain context features. The fused attention features, context features, and initial features are fused to obtain a segmentation result.

[0083] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the image segmentation method provided by each of the above methods. The method includes: inputting an image to be segmented into an image segmentation model, and obtaining a segmentation result of the image to be segmented output by the image segmentation model; wherein, the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network, and an atrous spatial pyramid pooling (ASPP) network. The preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain initial features of the image to be segmented, and inputs the initial features into the dual attention mechanism network and the ASPP network respectively. The dual attention mechanism network extracts and fuses position features and channel features of the initial features to obtain fused attention features, and the ASPP network extracts context features of the initial features to obtain context features. The fused attention features, context features, and initial features are fused to obtain a segmentation result.

[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0085] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image segmentation method, characterized in that: include: The image to be segmented is input into an image segmentation model, and a segmentation result of the image to be segmented output by the image segmentation model is obtained; wherein the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network and a dilated convolutional spatial pyramid ASPP network, the preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain the initial features of the image to be segmented, the initial features are respectively input into the dual attention mechanism network and the ASPP network, the dual attention mechanism network extracts and fuses position features and channel features of the initial features to obtain fused attention features, the ASPP network performs context feature extraction on the initial features to obtain context features, the fused attention features, the context features and the initial features are fused to obtain the segmentation result.

2. The image segmentation method according to claim 1, characterized in that: The preset lightweight convolutional neural network includes a convolutional layer, a first inverted residual layer, and a second inverted residual layer using a dilated convolution, and the preset lightweight convolutional neural network is used to obtain the initial features: Based on the convolution layer, feature extraction is performed on the image to be segmented to obtain a convolution feature map; Based on the first inverted residual layer, extract features from the convolution feature map to obtain an inverted residual feature map; Based on the second inverted residual layer, hole convolution feature extraction is performed on the inverted residual feature map to obtain the initial feature.

3. The image segmentation method according to claim 1, characterized in that: The dual attention mechanism network is used to obtain the fused attention feature: Convolution, first dimension conversion and transposition are sequentially performed on the initial features to obtain a first position feature matrix; Perform the convolution and the second dimension conversion on the initial features in sequence to obtain a second position feature matrix; The initial features are sequentially subjected to the convolution and the third dimension conversion to obtain a third position feature matrix; Performing feature fusion on the first position feature matrix, the second position feature matrix, the third position feature matrix and the initial feature to obtain a position attention matrix; Performing fourth-dimensional transformation and transposition on the initial features in sequence to obtain a first channel feature matrix; Performing the fourth dimension transformation on the initial features to obtain a second channel feature matrix; Performing a fifth-dimensional transformation on the initial features to obtain a third channel feature matrix; Performing feature fusion on the first channel feature matrix, the second channel feature matrix, the third channel feature matrix and the initial feature to obtain a channel attention matrix; The position attention matrix and the channel attention matrix are added to obtain the fused attention feature.

4. The image segmentation method according to claim 1, characterized in that: The ASPP network includes multiple layers of dilated convolutional layers with different dilation factors, and the ASPP network is used to obtain the context features: Based on the multiple layers of atrous convolutional layers with different dilation factors, extracting features of different scales from the initial features to obtain initial context features of multiple different scales; Feature fusion and feature extraction are performed on the multiple initial context features of different scales to obtain the context features.

5. The image segmentation method according to claim 1, characterized in that: The image segmentation model is used to obtain the segmentation result: Performing matrix addition on the fused attention feature and the context feature to fuse the fused attention feature and obtain a first initial feature map; Upsampling the first initial feature map to obtain a second initial feature map; Concatenate the second initial feature map and the initial feature to obtain a third initial feature map; The third initial feature map is convolved and up-sampled to restore the third initial feature map to the size of the image to be segmented, thereby obtaining the segmentation result.

6. The image segmentation method according to claim 3, characterized in that: The step of fusing the first position feature matrix, the second position feature matrix, the third position feature matrix, and the initial feature to obtain a position attention matrix includes: Performing matrix multiplication and activation on the first position feature matrix and the second position feature matrix to obtain a first position fusion matrix; Performing matrix multiplication on the first position fusion matrix and the third position feature matrix to obtain a second position fusion matrix; The second position fusion matrix and the initial feature are added together to obtain the position attention matrix.

7. The image segmentation method according to claim 3, characterized in that: The step of fusing the first channel feature matrix, the second channel feature matrix, the third channel feature matrix, and the initial feature to obtain a channel attention matrix includes: Performing matrix multiplication and activation on the first channel characteristic matrix and the second channel characteristic matrix to obtain a first channel fusion matrix; Performing matrix multiplication on the first channel fusion matrix and the third channel characteristic matrix to obtain a second channel fusion matrix; The second channel fusion matrix and the initial feature are matrix-added to obtain the channel attention matrix.

8. An image segmentation device, characterized in that: include: A segmentation module is used to input the image to be segmented into an image segmentation model, and obtain the segmentation result of the image to be segmented output by the image segmentation model; wherein the image segmentation model includes a preset lightweight convolutional neural network, a dual attention mechanism network and a dilated convolutional spatial pyramid ASPP network, the preset lightweight convolutional neural network performs initial feature extraction on the image to be segmented to obtain the initial features of the image to be segmented, and the initial features are respectively input into the dual attention mechanism network and the ASPP network, the dual attention mechanism network extracts and fuses position features and channel features of the initial features to obtain fused attention features, the ASPP network performs context feature extraction on the initial features to obtain context features, and the fused attention features, the context features and the initial features are fused to obtain the segmentation result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the image segmentation method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image segmentation method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the image segmentation method according to any one of claims 1 to 7 is implemented.