An image semantic segmentation method and system based on a lightweight neural network model
By extracting and fusing spatial and multi-scale feature information of images through the multi-branch structure of a lightweight neural network model, this method solves the problem of insufficient accuracy in real-time and lightweight image semantic segmentation methods, achieving efficient image semantic segmentation suitable for edge devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2026-04-07
AI Technical Summary
Existing image semantic segmentation methods based on neural network models cannot guarantee segmentation accuracy while ensuring real-time performance and lightweight design, and are difficult to deploy to edge devices.
A lightweight neural network model is adopted, including an initialization module, spatial branch, semantic branch and multi-scale feature fusion decoder, to achieve image semantic segmentation by extracting and fusing spatial information and multi-scale feature information of the image.
In a lightweight image semantic segmentation network with a small number of parameters, the segmentation accuracy and precision are improved, and fast semantic segmentation is achieved. It is suitable for edge devices, especially for scenarios such as autonomous driving, security monitoring, medical imaging and remote sensing images.
Smart Images

Figure CN116993987B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, particularly to computer vision, image processing, deep learning, and more specifically, to an image semantic segmentation method and system based on a lightweight neural network model. Background Technology
[0002] Image semantic segmentation methods, used to distinguish and classify different objects in images, have been a key research focus in the field of artificial intelligence in recent years. They can be applied in scenarios such as autonomous driving, security monitoring, medical imaging, facial recognition, and remote sensing images. For example, in autonomous driving scenarios, image semantic segmentation methods can be used to process environmental information. A high-level semantic segmentation method for road scenes can provide intelligent vehicles with fast and accurate road condition information, enabling them to make correct route planning and ensuring the safe driving of autonomous vehicles.
[0003] Traditional image semantic segmentation algorithms, due to their low computational complexity, typically offer fast processing speeds but low segmentation accuracy. Taking autonomous driving scenarios as an example, while these algorithms can meet the real-time requirements of road scene segmentation, they are prone to missegmentation of road targets, which severely impacts vehicles in motion. Existing technologies have proposed image learning-based semantic segmentation methods using deep networks. These methods typically construct neural networks by using a large number of trainable weights to form convolutional operations, and then train the network with a large number of image samples. This allows the network to automatically learn and complete the segmentation task. With advantages such as end-to-end processing and strong fitting ability, they have achieved significant progress in accuracy, compensating for the segmentation accuracy shortcomings of traditional image algorithms. However, they suffer from a large number of parameters and computational demands, placing extremely high hardware and computing resource requirements on edge devices where the algorithms are deployed. These computationally intensive, high-accuracy semantic segmentation methods are difficult to deploy on edge devices. Therefore, achieving lightweight image semantic segmentation methods while ensuring real-time performance has become an important research direction in modern computer science to meet the ever-increasing timeliness requirements of the information age.
[0004] Existing real-time image semantic segmentation methods generally use bottleneck structures as the basic building blocks of the encoder in the image semantic segmentation network to achieve lightweighting, but this leads to the loss and destruction of image features, which in turn affects the segmentation accuracy. Summary of the Invention
[0005] To overcome the shortcomings of the prior art in image semantic segmentation tasks based on neural network models, which cannot guarantee segmentation accuracy while ensuring real-time performance and lightweight design, this invention provides an image semantic segmentation method and system based on a lightweight neural network model.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0007] In a first aspect, there is an image semantic segmentation method based on a lightweight neural network model, wherein the lightweight neural network model includes an initialization module, a spatial branch, a semantic branch, and a multi-scale feature fusion decoder;
[0008] The image semantic segmentation method includes:
[0009] In response to the processing instruction of the image to be processed, the initialization module performs feature extraction on the image to be processed to obtain a first feature map;
[0010] Spatial information of the first feature map is extracted based on the spatial branch;
[0011] Based on the semantic branch, multi-scale feature information of the first feature map is extracted, and the multi-scale feature information and the spatial information are fused to obtain an enhanced feature map;
[0012] Based on the multi-scale feature fusion decoder, the first feature map and the enhanced feature map are fused and decoded, and the image size is restored to obtain the image semantic segmentation result.
[0013] In a second aspect, an image semantic segmentation system, applying the method described in the first aspect, includes:
[0014] The receiving unit is used to acquire the image to be processed;
[0015] The processing unit is used to carry a lightweight neural network model; it is also used to process the image to be processed to obtain image semantic segmentation results; wherein...
[0016] The lightweight neural network model includes:
[0017] An initialization module is used to extract features from the image to be processed to obtain a first feature map;
[0018] Spatial branch, used to extract spatial information from the first feature map;
[0019] The semantic branch is used to extract multi-scale feature information from the first feature map; it is also used to fuse the multi-scale feature information and the spatial information to obtain an enhanced feature map.
[0020] A multi-scale feature fusion decoder is used to fuse and decode the first feature map and the enhanced feature map, and perform image size restoration to obtain the image semantic segmentation result.
[0021] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0022] This invention provides an image semantic segmentation method and system based on a lightweight neural network model. The method utilizes the multi-branch structure of the lightweight neural network model to extract spatial information from the first feature map via the spatial branch and multi-scale feature information from the first feature map via the semantic branch at low cost. This multi-scale feature information is then fused with the spatial information to supplement the spatial information of the multi-scale feature map. Subsequently, a multi-scale fusion decoder fuses the first feature map with the spatially supplemented multi-scale feature information (i.e., the enhanced feature map) and performs image precision restoration (i.e., image size restoration). This method ensures the precision and accuracy of image semantic segmentation results in a lightweight image semantic segmentation network model with relatively few parameters, thereby improving the model's inference speed. Compared to existing technologies, this invention not only improves the segmentation accuracy of targets within an image (measured by mIoU), but also achieves fast semantic segmentation of images while maintaining lightweight characteristics. Ultimately, it achieves a good performance balance between segmentation accuracy and real-time performance, which can simultaneously meet the timeliness and accuracy requirements of practical application scenarios. It is easy to deploy in edge devices and is particularly suitable for application scenarios such as autonomous driving, security monitoring, medical imaging, face recognition, and remote sensing images. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the image semantic segmentation method in Example 1;
[0024] Figure 2 This is a schematic diagram of the image semantic segmentation network model in Example 1;
[0025] Figure 3 This is a schematic diagram of the FEB structure in Example 1;
[0026] Figure 4 This is a schematic diagram of the structure of DAB in Example 1;
[0027] Figure 5 This is a schematic diagram of the multi-scale feature fusion decoder in Example 1;
[0028] Figure 6 This is a comparison chart of experimental results for image semantic segmentation methods based on different image semantic segmentation models in Example 2;
[0029] Figure 7 This is a schematic diagram of the image semantic segmentation system in Example 3;
[0030] The reference numerals in the figures include:
[0031] 101 - Initialization module; 102 - Spatial branch; 103 - Semantic branch; 104 - Multi-scale feature fusion decoder;
[0032] 1021 - First spatial branch; 1022 - Second spatial branch;
[0033] 1031 - First semantic branch; 1032 - Second semantic branch;
[0034] 1041 - Spatial Attention Module. Detailed Implementation
[0035] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0036] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.
[0037] To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions; the same or similar reference numerals correspond to the same or similar parts;
[0038] It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.
[0039] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.
[0040] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0041] Example 1
[0042] This embodiment proposes an image semantic segmentation method based on a lightweight neural network model. (See reference...) Figure 1 The flowchart shown and Figure 2 The schematic diagram shown illustrates that the lightweight neural network model includes an initialization module 101, a spatial branch 102, a semantic branch 103, and a multi-scale feature fusion decoder 104.
[0043] Therefore, the image semantic segmentation method includes:
[0044] S1: In response to the processing instruction of the image to be processed, the initialization module 101 performs feature extraction on the image to be processed to obtain a first feature map;
[0045] S2: Extract spatial information of the first feature map based on the spatial branch 102;
[0046] S3: Extract multi-scale feature information from the first feature map based on the semantic branch 103, and fuse the multi-scale feature information and the spatial information to obtain an enhanced feature map; wherein, the semantic branch 103 extracts the multi-scale feature information based on the bottleneck structure;
[0047] S4: Based on the multi-scale feature fusion decoder (FFD) 104, the first feature map and the enhanced feature map are fused, and the image size is restored to obtain the image semantic segmentation result.
[0048] The lightweight neural network model for image semantic segmentation constructed in this embodiment is a Fast Ultra-lightweight Bilateral Network (FUBNet). This network model adopts a multi-branch structure (i.e., spatial branch 102 and semantic branch 103), extracts spatial information and multi-scale feature information from the first feature map, and fuses them to obtain an enhanced feature map, thus supplementing the spatial information of the multi-scale feature information, enhancing and preserving the spatial information, and ensuring the recovery of spatial features on the encoding side. In addition, on the decoding side, the first feature map and the enhanced feature map are fused by the multi-scale feature fusion decoder 104, and then the image size is restored. This can accurately and quickly improve the spatial detail recovery capability of the network model, thereby improving the segmentation accuracy of the image semantic segmentation result while ensuring real-time performance. Compared with the prior art, it can make full use of the multi-scale feature information of the image to be processed.
[0049] Those skilled in the art should understand that a lightweight neural network is a network with a small number of parameters; specifically, the number of parameters of the lightweight neural network is less than 50M.
[0050] In a preferred embodiment, the spatial branch includes a first spatial branch and a second spatial branch; the step of extracting spatial information of the first feature map based on the spatial branch includes:
[0051] Based on the first spatial branch 1021, a first spatial information compressed feature map about the first feature map is obtained;
[0052] Based on the second spatial branch 1022, a second spatial information compressed feature map about the first feature map is obtained;
[0053] The semantic branches include a first semantic branch 1031, a first superimposed unit, a second semantic branch 1032, and a second superimposed unit; the step of extracting multi-scale feature information from the first feature map based on the semantic branches and fusing the multi-scale feature information and the spatial information includes:
[0054] Based on the first semantic branch 1031, semantic features are extracted from the first feature map to obtain a first-scale semantic feature map;
[0055] The first scale semantic feature map and the first spatial information compressed feature map are fused together using the first overlay device to obtain the first enhanced feature map.
[0056] Based on the second semantic branch 1032, semantic features are extracted from the first enhanced feature map to obtain a second-scale semantic feature map;
[0057] The second superimposed feature map is fused with the second spatial information compressed feature map to obtain the second enhanced feature map.
[0058] In this preferred embodiment, spatial information is extracted from the first feature map through the first spatial branch 1021 and the second spatial branch 1022 to obtain a first spatial information compressed feature map and a second spatial information compressed feature map, respectively. Multi-scale feature information is extracted from the first feature map through the first semantic branch 1031 and the second semantic branch 1032 to obtain a first-scale semantic feature map and a second-scale semantic feature map, respectively. The first-scale semantic feature map is then fused with the first spatial information compressed feature map and the second-scale semantic feature map is fused with the second spatial information compressed feature map through the first superimposed unit and the second superimposed unit, respectively, to obtain enhanced feature maps (i.e., the first enhanced feature map and the second enhanced feature map).
[0059] In an optional embodiment, the first spatial branch 1021 includes a first multi-channel convolutional layer and a first single-channel convolutional layer; the second spatial branch 1022 includes a second multi-channel convolutional layer and a second single-channel convolutional layer.
[0060] The step of obtaining a first spatial information compressed feature map about the first feature map based on the first spatial branch 1021 includes:
[0061] The first spatial information feature map is obtained by using the first multi-channel convolutional layer to extract spatial information from the first feature map.
[0062] The first spatial information feature map is compressed by using the first single-channel convolutional layer to obtain the first spatial information compressed feature map.
[0063] And, obtaining a second spatial information compressed feature map about the first feature map based on the second spatial branch 1022 includes:
[0064] The second spatial information feature map is obtained by using the second multi-channel convolutional layer to extract spatial information from the first spatial information feature map;
[0065] The second spatial information feature map is compressed by using the second single-channel convolutional layer to obtain the second spatial information compressed feature map.
[0066] In this optional implementation, spatial information is extracted through a first multi-channel convolutional layer and a second multi-channel convolutional layer, respectively, to ensure processing efficiency while preventing excessive nonlinear operations from damaging spatial information, thereby obtaining a first spatial information feature map and a second spatial information feature map. Channel compression is performed through a first single-channel convolutional layer and a second single-channel convolutional layer, respectively, to prevent additional parameter consumption, thereby obtaining a first spatial information compressed feature map and a second spatial information compressed feature map.
[0067] It should be understood that the size and / or number of channels of the first multi-channel convolutional layer, the first single-channel convolutional layer, the second multi-channel convolutional layer, and the second single-channel convolutional layer shall be determined by those skilled in the art based on the actual situation, such as based on the size and / or number of channels of the feature map being processed.
[0068] In some examples, the first multi-channel convolutional layer and / or the second multi-channel convolutional layer are standard convolutional layers with a size of 3×3 and 32 channels;
[0069] In some examples, the first single-channel convolutional layer and / or the second single-channel convolutional layer are standard convolutional layers of size 3×3 with 1 channel.
[0070] In an optional embodiment, both the first semantic branch 1031 and the second semantic branch 1032 include a downsampling layer, several feature enhancement layers, a concatenation layer, and a point convolutional layer connected in sequence; the output of the downsampling layer is also connected to the input of the concatenation layer, for fusing feature map information of different depths;
[0071] The step of extracting semantic features from the first feature map based on the first semantic branch 1031 to obtain a first-scale semantic feature map includes:
[0072] The first feature map is downsampled using the downsampling layer to obtain a first downsampled feature map.
[0073] The first downsampled feature map is extracted by using a series of sequentially connected feature enhancement layers to obtain a first-scale feature map.
[0074] The first downsampled feature map and the first scale feature map are spliced and fused using a splicing layer, and the channel information of each channel of the feature map output by the splicing layer is fused and compressed through a point convolutional layer to obtain the first scale semantic feature map.
[0075] And, based on the second semantic branch 1032, the extraction of semantic features from the first enhanced feature map to obtain a second-scale semantic feature map includes:
[0076] The first enhanced feature map is downsampled using a downsampling layer to obtain a second downsampled feature map;
[0077] The second downsampled feature map is used to extract features by using several sequentially connected feature enhancement layers to obtain a second-scale feature map;
[0078] The second downsampled feature map and the second scale feature map are spliced and fused using a concatenation layer (Concat). Then, the channel information of each channel of the feature map output by the concatenation layer is fused and compressed through a point convolutional layer to obtain the second scale semantic feature map.
[0079] This optional embodiment processes the feature maps (i.e., the first feature map and the first enhanced feature map) through parallel dual branches (i.e., the first semantic branch 1031 and the second semantic branch 1032) with different receptive fields, generating feature maps of different scales (i.e., the first scale feature map and the second scale feature map). The spatial information (i.e., the first spatial information compressed feature map and the second spatial information compressed feature map) is fused to the aforementioned feature maps of different scales by an additive operation through the first superimposed unit and the second superimposed unit, respectively. The spatial information of the multi-scale feature maps is supplemented by a spatial mask method, which improves the accuracy of the image semantic segmentation result (considered as the accuracy of a lightweight neural network model).
[0080] It should be noted that the size and / or number of channels of the point convolutional layer in the first semantic branch 1031 and the point convolutional layer in the second semantic branch 1032 are adjusted by those skilled in the art according to the needs of model accuracy and model size in actual scenarios.
[0081] In some examples, the point convolutional layer in the first semantic branch is a standard convolutional layer of size 1×1 with 64 channels;
[0082] In some examples, the point convolutional layer in the second semantic branch is a standard convolutional layer of size 1×1 with 128 channels.
[0083] Furthermore, the feature enhancement layer includes at least one of the following: a DAB (Depth-wise Asymmetric Bottleneck) module, a FEB (Feature Enhancement Bottleneck) module, and a ResNet residual bottleneck module;
[0084] Among them, such as Figure 3 As shown, the FEB module includes:
[0085] The first deep convolutional layer is used to perform convolution operations on each channel of the input feature map using independent two-dimensional convolutional kernels to obtain the first deep feature map;
[0086] The first point convolutional layer is used to perform channel fusion on the first depth feature map through a preset number of point convolutional kernels to obtain the corresponding channel compressed feature map.
[0087] The second deep convolutional layer is used to perform convolution operations on each channel of the channel compressed feature map using independent two-dimensional convolutional kernels to obtain a second deep feature map with the same number of channels as the channel compressed feature map.
[0088] The deep-dilated convolution (DDConv) layer is used to perform depth convolution operations on each channel of the channel compressed feature map using independent two-dimensional convolution kernels and according to a preset dilation rate, so as to obtain a deep-dilated feature map.
[0089] The first fusion layer (Addition) is used to add the corresponding element values of the channel compression feature map, the second depth feature map, and the depth hole feature map to obtain the first semantic feature fusion feature map;
[0090] The first three-dimensional convolutional layer is used to mix the spatial features and channel features of each channel of the first semantic feature fusion feature map using several three-dimensional convolutional kernels, respectively, to obtain a feature map of each output channel; it is also used to combine the feature maps of each output channel to obtain a first combined feature map.
[0091] The second point convolutional layer is used to fuse the channel features of the combined feature map using a preset number of point convolutional kernels to obtain a feature map of each output channel; it is also used to combine the feature maps of each output channel to obtain a second combined feature map.
[0092] The second fusion layer (Addition) is used to add the corresponding elements of the input feature map of the first deep convolutional layer to the second combined feature map to obtain the base feature map output by the current feature enhancement layer.
[0093] It should be noted that the above embodiment achieves lightweighting of the encoding layer of the lightweight neural network model through a bottleneck structure feature enhancement layer, and realizes the extraction of multi-scale feature information. Several sequentially connected feature enhancement layers constitute a bottleneck structure module. It should also be noted that using spatial branching to extract spatial information from the first feature map can solve the problem that the bottleneck structure feature enhancement layer is prone to feature loss and destruction.
[0094] Those skilled in the art will understand that the ResNet residual module employs convolutional layers of varying sizes to form a bottleneck structure to reduce the number of parameters, including using a 1×1 convolutional layer for dimensionality reduction, followed by a 3×3 convolutional layer for convolution, and finally a 1×1 convolutional layer for dimensionality increase. The DAB module in this embodiment combines the ResNet residual module with various convolutional techniques, such as... Figure 4 The schematic diagram of the DAB structure shown employs a bottleneck structure that compresses channels to half through 3×3 convolutions. This structure retains more channels for feature extraction during compression compared to the ResNet residual module. Subsequent feature extraction uses depthwise convolution and asymmetric convolution techniques to limit parameter scaling, through two branches with different dilation rates (one of which includes two sequentially connected asymmetric depthwise convolution layers, i.e.) Figure 4 The other branch consists of two sequentially connected asymmetric depth-hole convolutional layers: 3×1DConv and 1×3DConv. Figure 4 The 3×1DDConv and 1×3DDConv in the model are used to obtain information at different scales.
[0095] In some examples, several sequentially connected ResNet residual modules are combined into a bottleneck structure module;
[0096] In some examples, several sequentially connected DABs are combined to form a bottleneck structure module;
[0097] In some examples, several sequentially connected FEBs are combined to form a bottleneck structure module;
[0098] In other examples, several FEBs and DABs are combined to form a bottleneck structure module.
[0099] For the FEB module, it should be noted that the two-dimensional convolution kernel represents only length and width.
[0100] Those skilled in the art should understand that, in the first semantic branch / second semantic branch, the basic feature map output by the second fusion layer in the last FEB module is the first-scale semantic feature map / second-scale semantic feature map.
[0101] In some examples, the kernel size of the first depthwise convolutional layer is 3×3;
[0102] In some examples, the kernel size of the first point convolutional layer is 1×1;
[0103] In some examples, the kernel size of the second depthwise convolutional layer is 3×3;
[0104] In some examples, the kernel size of the deep-hole convolutional layer is 3×3;
[0105] In some examples, the first 3D convolutional layer is operated using a standard convolutional kernel of size 3×3×C / 2.
[0106] It should be noted that the number of channels in the output feature map can be recovered by controlling the number of point convolution kernels. For example, see [link to relevant documentation]. Figure 3 The kernel size in the second point convolutional layer is 1×1×C, where C is the number of channels in the input feature map of the current FEB module, and the number of kernels C1 is the number of channels in the output feature map. When the number of kernels C1 is controlled to be C / 2, the number of channels C2 in the output feature map of the second point convolutional layer is twice C1 (i.e., C2 = C). The number of channels in the input and output feature maps of the entire FEB module is the same. It should be noted that the FEB module placed at different positions in the neural network can use different numbers of input / output channels, which need to match the number of channels allocated to the downsampling layer of the current network layer.
[0107] It should also be noted that the FEB module follows the bottleneck structure and multi-branch approach, introducing convolutional branches with different receptive fields and different dilation rates (i.e., a second deep convolutional layer and a deep dilated convolutional layer used to process channel compression feature maps) to acquire short-range and long-range features respectively, generating feature maps of different scales. Information is aggregated through a first fusion layer, and then enhanced through a first three-dimensional convolutional layer to obtain a multi-scale enhanced feature map (i.e., a first combined feature map). This effectively improves the multi-scale feature capture capability of the lightweight neural network model and enhances its ability to segment targets of different scales. Furthermore, it should be understood that the kernel size and / or number of channels of the first three-dimensional convolutional layer and the second point convolutional layer are determined by those skilled in the art based on the actual situation.
[0108] In some examples, experiments were conducted using the Cityscapes dataset to evaluate network models based solely on different feature enhancement layers. The results are as follows:
[0109] Table 1 Experimental results of different feature enhancement layers
[0110]
[0111] It can be seen that all three types of feature enhancement layers have high accuracy. Among them, the network model built based on the ResNet residual module only requires 0.23M parameters and achieves an inference speed of 294.6fps, with a segmentation performance of 50.2% mIoU. The network model built based on FEB has a 20.6% higher mIoU than the ResNet residual module. This is because FEB uses the difference in receptive fields of the two branches to capture multi-scale feature information, and uses convolutional windows with different weighting ranges to calculate the same input feature map. The pixel weights calculated by different windows are further fused and feature enhanced, so that FEB can simultaneously consider the relationship between dense feature points in a small range and short distance and sparse feature points in a large range and long distance, thereby effectively improving the multi-scale feature extraction capability of the semantic branch. It should be emphasized that although the FEB channel compression ratio is smaller than that of the ResNet residual module, the number of parameters is slightly increased, and the processing efficiency is reduced to a certain extent, those skilled in the art should understand that, given the continuous improvement of current hardware computing power, sacrificing some processing efficiency in order to achieve a huge increase in accuracy is completely acceptable.
[0112] Furthermore, the DAB-based network model achieves 69.6% mIoU with a parameter count of 0.68M. It's worth noting that FEB and DAB employ different convolutional techniques as the bottleneck structure's entry point: FEB uses depthwise separable convolution, while DAB uses standard convolution. Additionally, FEB utilizes a first 3D convolutional layer in subsequent processing to implement feature-enhancing convolution, further improving semantic extraction capabilities. Compared to DAB, FEB maintains high segmentation accuracy despite using sparser convolutional techniques. The combined parameter count of the first depthwise convolutional layer and the feature-enhancing convolution in the FEB module is lower than that of the standard convolution used in the DAB module at the entry point, achieving a lighter and more efficient result. The semantic branch built based on FEB is characterized by its lightweight and high accuracy.
[0113] Furthermore, in one of the FEB modules, the outputs of the first depthwise convolutional layer, the second depthwise convolutional layer, the depthwise dilated convolutional layer, the first three-dimensional convolutional layer, and the second point convolutional layer, as well as the input of the first depthwise convolutional layer, are all sequentially connected to a BN (Batch Normalization) layer and a PReLU activation layer.
[0114] Those skilled in the art should understand that the BN layer is used to standardize the feature map data, which to some extent suppresses the problems of gradient vanishing and gradient explosion, accelerates the convergence speed of the model, and improves the generalization ability of the model; the PReLU activation layer can avoid the problem of neuron inactivation and can adaptively learn parameters from the feature map data.
[0115] Furthermore, the number of feature enhancement layers in the first semantic branch and the second semantic branch are different; and the feature map space size and number of channels of each feature enhancement layer in the first semantic branch and the second semantic branch are the same.
[0116] It should be noted that the feature map space size and number of channels of each feature enhancement layer in the first semantic branch and the second semantic branch are the same; those skilled in the art should understand that the feature map space scale and number of channels can be adjusted and controlled by the downsampling layer in the corresponding semantic branch (i.e., the first semantic branch or the second semantic branch).
[0117] In an optional embodiment, see Figure 5 The multi-scale feature fusion decoder includes a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, a fifth point convolutional layer, an upsampling layer, and a sixth point convolutional layer.
[0118] The step of fusing and decoding the first feature map and the enhanced feature map based on the multi-scale feature fusion decoder, and performing image size restoration, includes:
[0119] The first feature map is convolved by the first depthwise separable convolutional layer to obtain a first feature map to be decoded; wherein, the first depthwise separable convolutional layer includes a third depthwise convolutional layer and a third pointwise convolutional layer connected in sequence, the number of convolutional kernels of the third depthwise convolutional layer is the same as the number of channels of the first feature map, and the number of convolutional kernels of the third pointwise convolutional layer is the same as the number of channels of the first enhanced feature map.
[0120] The first enhanced feature map is convolved by the second depthwise separable convolutional layer to obtain the second feature map to be decoded; wherein the second depthwise separable convolutional layer includes a fourth depthwise convolutional layer and a fourth pointwise convolutional layer, and the number of convolutional kernels of the fourth depthwise convolutional layer and the fourth pointwise convolutional layer are the same as the number of channels of the first enhanced feature map;
[0121] The second enhanced feature map is merged and compressed by the fifth convolutional layer to restore the number of channels of the second enhanced feature map to the same number of channels as the first enhanced feature map. Then, the spatial size is restored by the upsampling layer to obtain the third feature map to be decoded.
[0122] The first feature map to be decoded, the second feature map to be decoded, and the third feature map to be decoded are added and fused to obtain the final fused feature map.
[0123] The image semantic segmentation result is obtained by performing image size restoration and pixel-level classification on the final fused feature map through the sixth convolutional layer.
[0124] It should be noted that the size and / or number of channels of the first depthwise separable convolutional layer, the second depthwise separable convolutional layer, the fifth point convolutional layer, and the sixth point convolutional layer in the multi-scale feature fusion decoder shall be determined by those skilled in the art based on the actual situation.
[0125] It should also be noted that the number of channels in the sixth convolutional layer is the same as the number of target categories that can be obtained from the image semantic segmentation of the image to be detected. Those skilled in the art should understand that the training set used during the training of the lightweight neural network model is labeled with target category labels, and each target category label corresponds to a type of target object. This means that the number of channels in the sixth convolutional layer is consistent with the total number of target category labels labeled in the training set.
[0126] In some examples, the target categories include, but are not limited to, cars, signs, road lines, pedestrians, bicycles, buildings, and curbs.
[0127] In some examples, the kernel size used in the first depthwise separable convolutional layer for processing the first feature map is 3×3;
[0128] In some examples, the kernel size used in the second depthwise separable convolutional layer for processing the first enhanced feature map is 3×3;
[0129] In some examples, the fifth convolutional layer is a standard convolutional layer with a size of 1×1 and 64 channels;
[0130] In some examples, the sixth convolutional layer is a standard convolutional layer of size 1×1.
[0131] In some examples, the upsampling operation in the upsampling layer is implemented using a bilinear interpolation upsampling method.
[0132] Furthermore, in the multi-scale feature fusion decoder, a spatial attention mechanism is introduced during the addition and fusion operation of the first feature map to be decoded, the second feature map to be decoded, and the third feature map to be decoded, including:
[0133] The first feature map to be decoded is weighted and focused by the spatial attention module 1041 to obtain a spatial attention feature map.
[0134] The spatial attention feature map, the first feature map to be decoded, the second feature map to be decoded, and the third feature map to be decoded are added and fused to obtain the final fused feature map.
[0135] It should be noted that the above embodiments introduce a spatial attention mechanism to give weighted attention to important information in the image spatial region and suppress unimportant information. This effectively improves the decoder's recovery accuracy of objects at different scales with low resource overhead, thereby improving model performance and enhancing the segmentation accuracy of image semantic segmentation results.
[0136] Furthermore, the spatial attention module 1041 includes a seventh-point convolutional layer, a single-channel second three-dimensional convolutional layer, a sigmoid activation layer, and a multiplicative weighted layer;
[0137] The step of weighting the first feature map to be decoded through the spatial attention module 1041 includes:
[0138] The first feature map to be decoded is sequentially processed by the seventh point convolutional layer and the second three-dimensional convolutional layer to obtain a compressed spatial information feature map.
[0139] The Sigmoid activation layer is used to assign weights to each element in the compressed spatial information feature map to obtain a spatial attention mask map.
[0140] The corresponding elements in the first feature map to be decoded and the spatial attention mask map are multiplied and weighted using the multiplication weighting layer to obtain the spatial attention feature map.
[0141] It should be noted that the spatial attention module 1041 follows the general paradigm of spatial attention mechanisms; the size and / or number of channels of the seventh point convolutional layer and the second three-dimensional convolutional layer shall be determined by those skilled in the art based on the actual situation.
[0142] In some examples, the seventh convolutional layer is a standard convolutional layer with a size of 1×1 and 64 channels;
[0143] In some examples, the second 3D convolutional layer is a standard convolutional layer with a size of 3×3 and a channel count of 1.
[0144] In a preferred embodiment, the initialization module performs feature extraction on the image to be processed by sequentially passing the image to be processed through at least three ninth standard convolutional layers to obtain a first feature map.
[0145] In a preferred embodiment, the training strategy for the lightweight neural network model includes:
[0146] Construct an initial lightweight neural network model, train it from scratch using the network parameter initialization method, and use the Stochastic Gradient Descent (SGD) optimizer or the Adam optimizer as the optimization strategy.
[0147] The training strategy also includes at least one of the following:
[0148] A polynomial decay learning rate strategy is adopted;
[0149] Embed an Online Hard Example Mining (OHEM) mechanism into the stochastic gradient descent optimizer;
[0150] The images in the training set are preprocessed, including random training order, random horizontal flipping, mean subtraction, random scaling and / or random cropping.
[0151] Those skilled in the art should understand that neural network models need to be trained and optimized before they can be used, and the termination conditions for training can be set by those skilled in the art according to the actual situation.
[0152] By way of example, the training of the FUBNet described in this invention may be terminated when the loss function converges to a preset threshold or when the training epochs reach a preset value.
[0153] In some examples, the FUBNet is trained on an RTX 3090 GPU with CUDA 11.4 and cuDNN V8 integrated, using the Cityscapes dataset and / or the CamVid dataset as training sets during training.
[0154] It should be noted that the polynomial decay learning rate strategy used in this preferred embodiment can avoid unreasonable learning rate settings.
[0155] As a non-restrictive example, using the SGD optimizer as the optimization strategy for model training, if the learning rate is too large, it is easy to cause excessive gradient descent and excessive loss jitter. If the learning rate is too small, the gradient descent is too slow, which can also make it difficult for the network to converge. To avoid these problems, a multinomial decay learning rate strategy can be used.
[0156] In one specific implementation, the Cityscapes dataset was used as the training set, and a "poly" learning rate decay strategy was adopted, with the initial learning rate set to 4.5e-2 and the power factor to 0.9. To ensure appropriate momentum decay to take into account the weighting of historical gradient information during gradient descent, the momentum and momentum weight decay coefficients were set to 0.9 and 1e-4, respectively.
[0157] As a non-restrictive example, the CamVid dataset is used as the training set, and the Adam optimizer is used as the optimization strategy for model training: the initial learning rate and weight decay coefficient are set to 1e-3 and 1e-4, respectively; the batch size for the dataset is 8, and if there are fewer than 8 remaining images, the batch is skipped; and the maximum number of training rounds is set to 1000.
[0158] In some examples, the preprocessing operations employed include random scaling; where the random scaling factors are set to {0.75, 1.0, 1.25, 1.5, 1.75, 2.0} respectively.
[0159] In some examples, the preprocessing operations employed include random cropping, such as randomly cropping image data from the Cityscapes dataset to a resolution of 512×1024; or randomly cropping image data from the CamVid dataset to two image resolutions: 720×960 and 360×480.
[0160] In some examples, the preprocessing operations used include random training order, random horizontal flipping, mean subtraction, and random scaling.
[0161] It should also be noted that the OHEM mechanism can be used to alleviate the imbalance between easy and difficult samples during training.
[0162] In some examples, a class weighting scheme was also used to mitigate the class imbalance problem in the dataset.
[0163] Example 2
[0164] To verify the present invention, this example conducted an experiment on the method described in Example 1, as follows:
[0165] (I) Experiments evaluating spatial branching in combination with different semantic branches on the Cityscapes dataset
[0166] Experiments were conducted using the Cityscapes dataset. On the model encoding side, using the same network width and structure, three different semantic branches were constructed using ResNet residual modules, DAB, and FEB, respectively. The experimental results for each semantic branch before and after adding the spatial branch are recorded below:
[0167] Table 2. Experimental results of semantic branching combined with spatial branching based on different feature enhancement layers.
[0168]
[0169] As shown in Table 2, for the semantic branch using the ResNet residual module, adding a spatial branch to supplement spatial details improved the mIoU (mean intersection-over-union ratio) by 0.5% and decreased the speed (forward inference speed) by 23 fps. For the semantic branch using DAB, adding a spatial branch increased the total number of params (weight parameters) by 0.02M and decreased the speed by 11 fps, but achieved a 0.6% mIoU accuracy gain. For the semantic branch using FEB, adding a spatial branch, at a similar cost to the DAB module, achieved a 1.2% mIoU accuracy gain. Overall, the additional computational cost of spatial branches is similar for different networks, and their accuracy improvement effect is related to the characteristics of the model's encoding side, with the best effect observed in networks based on the FEB module.
[0170] (II) Evaluation of FFD with different semantic branches on the Cityscapes dataset
[0171] Experiments were conducted using the Cityscapes dataset. Three different semantic branches were constructed using ResNet residual modules, DAB, and FEB, respectively. The experimental results before and after using FFD for each semantic branch are recorded below:
[0172] Table 3. Experimental results of semantic branching combined with FFD based on different feature enhancement layers.
[0173]
[0174] FFD is a three-branch decoder whose channel width is determined by the channel widths of the three input feature maps. These three inputs correspond to the outputs of the three stages of the model's encoding side. The network width of the FFD is the same for different semantic branches; in other words, the additional parameter consumption brought by FFD is the same, 0.04M, as shown in Table 3. Furthermore, since the FFD width is the same, the additional processing time added by different FFDs for the three semantic branches is also the same. Table 3 uses the Speed metric to measure the number of images processed per second. It can be seen that the speed changes caused by FFD for different semantic branches are different. Adding FFD to the ResNet residual module-based encoder reduced the processing speed by 40 fps, the DAB encoder by about 17 fps, and the FEB encoder by 14 fps. Those skilled in the art should understand that this speed sacrifice is worthwhile. After adding FFD, the accuracy improvements for the three semantic branches are 0.7%, 0.6%, and 0.7% mIoU, respectively. In particular, the FEB-based semantic branch achieved a segmentation accuracy of 72.7% mIoU after adding FFD, which is an excellent performance for real-time semantic segmentation tasks. Overall, the experiment demonstrates that FFD has excellent spatial information recovery capabilities and features such as being ultra-lightweight and highly efficient, making it very suitable for ultra-lightweight real-time semantic segmentation scenarios.
[0175] (III) Experiments evaluating image semantic segmentation methods based on different models on the Cityscapes and CamVid datasets
[0176] Experiments were conducted using the Cityscapes and CamVid datasets to test image semantic segmentation methods based on different lightweight neural network models, including FUBNet, SegNet, ENet, SQNet, ESPNet, ESPNetV2, CGNet, EDANet, LEDNet, DABNet, ESNet, DFANet, MiniNet-v2, AGLNet, and MSCFNet. The experimental results are recorded below:
[0177] Table 4 Comparison of experimental results of different image semantic segmentation methods on the Cityscapes dataset.
[0178]
[0179] As shown in Table 4, when the input image resolution is 512×1024, FUBNet achieves a mIoU of 72.7% on the Cityscapes validation set, performing best among all network models. On the Cityscapes test set, FUBNet's mIoU reaches 72.4%, also the best result. ENet and ESPNet are the networks with the smallest number of parameters, with only 0.36M parameters, 0.21M fewer than the proposed FUBNet, almost half its size. Their mIoU is more than 10% lower than FUBNet's. Those skilled in the art should understand that in practical applications, for ultra-lightweight models with less than 1M parameters, a 0.21M parameter difference is negligible for deployment, while a 10% difference in mIoU can produce very significant performance differences. For networks of similar size to FUBNet, such as CGNet (0.50M) and MiniNet-v2, the increased expressive power due to the larger network size leads to higher segmentation accuracy, but they still lag behind FUBNet by 7.6% and 1.9% in mIoU, respectively. For larger models, such as DABNet (0.76M) and LEDNet (0.95M), FUBNet still achieves an accuracy advantage of over 2.3%. For networks around 1M in size, such as AGLNet and MSCFNet, they achieve segmentation accuracy of over 71% mIoU. These real-time semantic segmentation networks strike a balance between accuracy and network size, but FUBNet clearly achieves a superior balance. Overall, FUBNet achieves moderate performance in terms of computational complexity and forward inference speed, but leads in parameter size and accuracy. These results demonstrate that FUBNet has achieved a new balance among the three metrics and performs exceptionally well among numerous real-time semantic segmentation models.
[0180] Table 5 Comparison of experimental results of different image semantic segmentation methods on the CamVid dataset.
[0181]
[0182] As shown in Table 5, on the CamVid dataset, when the input resolution is 360×480, FUBNet's mIoU is 1.1% higher than LEDNet, and its network size is only 60% of LEDNet's. Compared to ENet and ESPNet, which have the fewest parameters, FUBNet's segmentation accuracy is more than 10% higher. When the image input resolution is 720×960, FUBNet can still achieve an accuracy as high as 71.3% mIoU, which is 2.3% and 2.0% higher than MiniNet-v2 and MSCFNet, respectively. In terms of computation, FUBNet's FLOPs are only 19.6 GFLOPs, which is lower than most lightweight neural network models. For some low-computational-cost networks such as FPENet and ESPNet, although these networks have achieved certain results in reducing computational cost, they often sacrifice segmentation performance. In contrast, FUBNet performs better in balancing efficiency and accuracy. In summary, even on the CamVid dataset, FUBNet still performs excellently. It not only rivals most existing real-time semantic segmentation methods in terms of efficiency, but also achieves higher segmentation accuracy. Its superior performance can provide higher segmentation quality assurance for practical application scenarios.
[0183] (iv) Visual Evaluation: Experiments were conducted using the Cityscapes dataset to evaluate image semantic segmentation methods based on different models, including FUBNet, LEDNet, ERFNet, and DABNet. The resulting image semantic segmentation results are shown below. Figure 6 As shown; the boxed portion represents the differences in image semantic segmentation results obtained based on different lightweight neural networks.
[0184] It can be seen that, Figure 6The first column of images shows a distant lawn, a very small target with limited pixel information. FUBNet correctly identified and represented this grassy area relatively completely, while other networks either failed to identify it or only identified a very small area, demonstrating FUBNet's strong ability to recognize small-scale objects. The second column of images primarily observes the ability of different lightweight neural network models to distinguish the true categories. The box outlines an overlapping area of a bus and a car. The first two lightweight neural network models misidentified the bus as a car or truck, and the car as a bus, due to the complexity of the information within the large receptive field, making accurate segmentation difficult. LEDNet achieved high accuracy in center segmentation, but its contour segmentation of the car was severely distorted. FUBNet, however, performed well in distinguishing between the two vehicles, although the car's contour was slightly distorted. The third column shows a relatively simple wall. All networks misclassified the wall as lawn, and the boundaries also showed some undesirable characteristics. FUBNet, however, still performed exceptionally well. The object marked in the fourth column is a fence. Its small size and occlusion make it highly susceptible to misidentification; only LEDNet and FUBNet can identify it relatively accurately and completely. The fifth column shows a complete sign, consisting of a triangular sign and a matching slender pole. ERFNet segmented the triangular sign the worst; DABNet and LEDNet segmented the triangular sign relatively completely, but neither handled the accompanying pole well; while FUBNet was able to completely segment the entire sign and pole.
[0185] Overall, FUBNet demonstrates excellent accuracy, precision, and inference speed, indicating that the FUBNet-based image semantic segmentation method can guarantee segmentation accuracy while maintaining real-time performance and lightweight design.
[0186] Example 3
[0187] This embodiment proposes an image semantic segmentation system, see reference. Figure 7 The image semantic segmentation method described in Example 1 includes:
[0188] The receiving unit is used to acquire the image to be processed;
[0189] The processing unit is used to carry a lightweight neural network model; it is also used to process the image to be processed to obtain image semantic segmentation results; wherein...
[0190] The lightweight neural network model includes:
[0191] An initialization module is used to extract features from the image to be processed to obtain a first feature map;
[0192] Spatial branch, used to extract spatial information from the first feature map;
[0193] The semantic branch is used to extract multi-scale feature information from the first feature map; it is also used to fuse the multi-scale feature information and the spatial information to obtain an enhanced feature map.
[0194] A multi-scale feature fusion decoder is used to fuse and decode the first feature map and the enhanced feature map, and perform image size restoration to obtain the image semantic segmentation result.
[0195] It is understood that the system in this embodiment corresponds to the method in Embodiment 1 above, and the options in Embodiment 1 above are also applicable to this embodiment, so they will not be described again here.
[0196] Preferably, the image semantic segmentation system is configured with Figure 2 The neural network model shown.
[0197] Example 4
[0198] This embodiment proposes a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor, causing the processor to perform some or all of the steps of the method described in Embodiment 1.
[0199] It is understood that the storage medium can be transient or non-transient. Exemplarily, the storage medium includes, but is not limited to, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0200] By way of example, the processor may be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0201] In some examples, a computer program product is provided, which can be implemented by hardware, software, or a combination thereof. As a non-limiting example, the computer program product can be embodied in the storage medium, or it can be embodied in a software product, such as an SDK (Software Development Kit).
[0202] In some examples, a computer program is provided, including computer-readable code, wherein, when the computer-readable code is run in a computer device, a processor in the computer device performs some or all of the steps for implementing the method.
[0203] This embodiment also proposes an electronic device, including a memory and a processor. The memory stores at least one instruction, at least one program, code set, or instruction set. When the processor executes the at least one instruction, at least one program, code set, or instruction set, it implements some or all of the steps of the method described in Embodiment 1.
[0204] In some examples, a hardware entity of the electronic device is provided, including: a processor, a memory, and a communication interface; wherein the processor typically controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers via a network; the memory is configured to store instructions and applications executable by the processor, and may also cache data to be processed or already processed (including but not limited to image data, audio data, voice communication data, and video communication data) to be processed by the processor and various modules in the electronic device, and may be implemented using flash memory or random access memory (RAM).
[0205] Furthermore, data can be transferred between the processor, communication interface, and memory via a bus, which can include any number of interconnected buses and bridges, connecting various circuits of one or more processors and memories together.
[0206] It is understood that the options in Embodiment 1 above also apply to this embodiment, so they will not be described again here.
[0207] The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent.
[0208] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. It should be understood that in the various embodiments of this disclosure, the sequence number of each step / process does not imply the order of execution. The execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments. It should also be understood that the device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. For those skilled in the art, other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the claims of the present invention.
Claims
1. An image semantic segmentation method based on a lightweight neural network model, characterized in that, The lightweight neural network model includes an initialization module, a spatial branch, a semantic branch, and a multi-scale feature fusion decoder; The image semantic segmentation method includes: In response to the processing instruction of the image to be processed, the initialization module performs feature extraction on the image to be processed to obtain a first feature map; Spatial information of the first feature map is extracted based on the spatial branch; Based on the semantic branch, multi-scale feature information of the first feature map is extracted, and the multi-scale feature information and the spatial information are fused to obtain an enhanced feature map; Based on the multi-scale feature fusion decoder, the first feature map and the enhanced feature map are fused and decoded, and the image size is restored to obtain the image semantic segmentation result; The spatial branch includes a first spatial branch and a second spatial branch; the extraction of spatial information from the first feature map based on the spatial branch includes: Based on the first spatial branch, a first spatial information compressed feature map about the first feature map is obtained; Based on the second spatial branch, a second spatial information compressed feature map about the first feature map is obtained; The semantic branch includes a first semantic branch, a first superimposed unit, a second semantic branch, and a second superimposed unit; the step of extracting multi-scale feature information from the first feature map based on the semantic branch and fusing the multi-scale feature information and the spatial information includes: Based on the first semantic branch, semantic features are extracted from the first feature map to obtain a first-scale semantic feature map; The first scale semantic feature map and the first spatial information compressed feature map are fused together using the first overlay device to obtain the first enhanced feature map. Based on the second semantic branch, semantic features are extracted from the first enhanced feature map to obtain a second-scale semantic feature map; The second scale semantic feature map and the second spatial information compressed feature map are fused together by the second overlay to obtain the second enhanced feature map. The multi-scale feature fusion decoder includes a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, a fifth point convolutional layer, an upsampling layer, and a sixth point convolutional layer; The step of fusing and decoding the first feature map and the enhanced feature map based on the multi-scale feature fusion decoder, and performing image size restoration, includes: The first feature map is convolved by the first depthwise separable convolutional layer to obtain a first feature map to be decoded; wherein, the first depthwise separable convolutional layer includes a third depthwise convolutional layer and a third pointwise convolutional layer connected in sequence, the number of convolutional kernels of the third depthwise convolutional layer is the same as the number of channels of the first feature map, and the number of convolutional kernels of the third pointwise convolutional layer is the same as the number of channels of the first enhanced feature map. The first enhanced feature map is convolved by the second depthwise separable convolutional layer to obtain the second feature map to be decoded; wherein the second depthwise separable convolutional layer includes a fourth depthwise convolutional layer and a fourth pointwise convolutional layer, and the number of convolutional kernels of the fourth depthwise convolutional layer and the fourth pointwise convolutional layer are the same as the number of channels of the first enhanced feature map; The second enhanced feature map is merged and compressed by the fifth convolutional layer to restore the number of channels of the second enhanced feature map to the same number of channels as the first enhanced feature map. Then, the spatial size is restored by the upsampling layer to obtain the third feature map to be decoded. The first feature map to be decoded, the second feature map to be decoded, and the third feature map to be decoded are added and fused to obtain the final fused feature map. The image semantic segmentation result is obtained by performing image size restoration and pixel-level classification on the final fused feature map through the sixth convolutional layer.
2. The image semantic segmentation method based on a lightweight neural network model according to claim 1, characterized in that, The first spatial branch includes a first multi-channel convolutional layer and a first single-channel convolutional layer; the second spatial branch includes a second multi-channel convolutional layer and a second single-channel convolutional layer. The step of obtaining a first spatial information compressed feature map based on the first spatial branch and the first feature map includes: The first spatial information feature map is obtained by using the first multi-channel convolutional layer to extract spatial information from the first feature map. The first spatial information feature map is compressed by using the first single-channel convolutional layer to obtain the first spatial information compressed feature map. And, obtaining a second spatial information compressed feature map about the first feature map based on the second spatial branch includes: The second spatial information feature map is obtained by using the second multi-channel convolutional layer to extract spatial information from the first spatial information feature map; The second spatial information feature map is compressed by using the second single-channel convolutional layer to obtain the second spatial information compressed feature map.
3. The image semantic segmentation method based on a lightweight neural network model according to claim 1, characterized in that, Both the first semantic branch and the second semantic branch include a downsampling layer, several feature enhancement layers, a concatenation layer, and a point convolutional layer connected in sequence; the output of the downsampling layer is also connected to the input of the concatenation layer, which is used to fuse feature map information of different depths; The step of extracting semantic features from the first feature map based on the first semantic branch to obtain a first-scale semantic feature map includes: The first feature map is downsampled using the downsampling layer to obtain a first downsampled feature map. The first downsampled feature map is extracted by using a series of sequentially connected feature enhancement layers to obtain a first-scale feature map. The first downsampled feature map and the first scale feature map are spliced and fused using a splicing layer, and the channel information of each channel of the feature map output by the splicing layer is fused and compressed through a point convolutional layer to obtain the first scale semantic feature map. And, the step of extracting semantic features from the first enhanced feature map based on the second semantic branch to obtain a second-scale semantic feature map includes: The first enhanced feature map is downsampled using a downsampling layer to obtain a second downsampled feature map; The second downsampled feature map is used to extract features by using several sequentially connected feature enhancement layers to obtain a second-scale feature map; The second downsampled feature map and the second scale feature map are spliced and fused using a splicing layer, and the channel information of each channel of the feature map output by the splicing layer is fused and compressed through a point convolutional layer to obtain the second scale semantic feature map.
4. The image semantic segmentation method based on a lightweight neural network model according to claim 3, characterized in that, The feature enhancement layer includes at least one of the DAB module, the FEB module, and the ResNet residual module; The FEB module includes: The first deep convolutional layer is used to perform convolution operations on each channel of the input feature map using independent two-dimensional convolutional kernels to obtain the first deep feature map; The first point convolutional layer is used to perform channel fusion on the first depth feature map through a preset number of point convolutional kernels to obtain the corresponding channel compressed feature map. The second deep convolutional layer is used to perform convolution operations on each channel of the channel compressed feature map using independent two-dimensional convolutional kernels to obtain a second deep feature map with the same number of channels as the channel compressed feature map. The deep-dilated convolutional layer is used to perform deep convolution operations on each channel of the channel compressed feature map using independent two-dimensional convolutional kernels and according to a preset dilation rate, so as to obtain a deep-dilated feature map. The first fusion layer is used to add the corresponding element values of the channel compression feature map, the second depth feature map, and the depth hole feature map to obtain the first semantic feature fusion feature map. The first three-dimensional convolutional layer is used to mix the spatial features and channel features of each channel of the first semantic feature fusion feature map using several three-dimensional convolutional kernels, respectively, to obtain a feature map of each output channel; it is also used to combine the feature maps of each output channel to obtain a first combined feature map. The second point convolutional layer is used to fuse the channel features of the first combined feature map using a preset number of point convolutional kernels to obtain a feature map of each output channel; it is also used to combine the feature maps of each output channel to obtain the second combined feature map. The second fusion layer is used to add the corresponding elements of the input feature map of the first deep convolutional layer to the second combined feature map to obtain the feature map output by the current feature enhancement layer.
5. The image semantic segmentation method based on a lightweight neural network model according to claim 3, characterized in that, The number of feature enhancement layers in the first semantic branch and the second semantic branch are different; and the spatial size and number of channels of the feature maps output by each feature enhancement layer in the first semantic branch and the second semantic branch are the same.
6. The image semantic segmentation method based on a lightweight neural network model according to claim 1, characterized in that, In the multi-scale feature fusion decoder, a spatial attention mechanism is introduced during the addition and fusion operation of the first feature map to be decoded, the second feature map to be decoded, and the third feature map to be decoded, including: The first feature map to be decoded is weighted and focused using a spatial attention module to obtain a spatial attention feature map. The spatial attention feature map, the first feature map to be decoded, the second feature map to be decoded, and the third feature map to be decoded are added and fused to obtain the final fused feature map.
7. The image semantic segmentation method based on a lightweight neural network model according to claim 6, characterized in that, The spatial attention module includes a seventh-point convolutional layer, a single-channel second three-dimensional convolutional layer, a sigmoid activation layer, and a multiplicative weighted layer. The step of weighting the first feature map to be decoded through a spatial attention module includes: The first feature map to be decoded is sequentially processed by the seventh point convolutional layer and the second three-dimensional convolutional layer to obtain a compressed spatial information feature map. The Sigmoid activation layer is used to assign weights to each element in the compressed spatial information feature map to obtain a spatial attention mask map. The corresponding elements in the first feature map to be decoded and the spatial attention mask map are multiplied and weighted using the multiplication weighting layer to obtain the spatial attention feature map.
8. An image semantic segmentation system, employing the image semantic segmentation method based on a lightweight neural network model as described in any one of claims 1-7, characterized in that, include: The receiving unit is used to acquire the image to be processed; Processing unit, used to carry lightweight neural network models; It is also used to process the image to be processed to obtain image semantic segmentation results; wherein, the lightweight neural network model includes: An initialization module is used to extract features from the image to be processed to obtain a first feature map; Spatial branch, used to extract spatial information from the first feature map; The semantic branch is used to extract multi-scale feature information from the first feature map; it is also used to fuse the multi-scale feature information and the spatial information to obtain an enhanced feature map. A multi-scale feature fusion decoder is used to fuse and decode the first feature map and the enhanced feature map, and perform image size restoration to obtain the image semantic segmentation result.
Citation Information
Patent Citations
Lightweight multi-scale feature fusion real-time image semantic segmentation method and system
CN114445430A
Joint boundary detection image segmentation and object recognition using deep learning
EP3171297A1