Semantic segmentation method and device based on lightweight coding-decoding network, equipment and medium
Through the lightweight encoding-decoding network combined with the multi-scale attention mechanism, the problem of resource limitation of deep learning methods on mobile terminals is solved, real-time semantic segmentation on mobile devices is realized, and category accuracy and meticulousness of object contour prediction are improved.
Patent Information
- Application Number
- CN202510626122.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-12
AI Technical Summary
Existing deep learning methods have high computing resource requirements, complex models, and long training time in semantic segmentation tasks, making it difficult to implement real-time inference on mobile devices, and traditional methods have limited effects in complex image and semantic segmentation tasks.
A lightweight encoding-decoding network is adopted to perform low-resolution feature extraction through the encoder and feature details recovery through the decoder. Combining multi-scale spatial attention and channel attention mechanisms, multi-layer perceptron and Softmax functions are used for pixel-by-pixel classification prediction.
Real-time semantic segmentation on mobile devices is realized, class accuracy and object profile prediction are improved, and the optimization balance between speed and accuracy is achieved.
Smart Images

Figure CN120472171A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a semantic segmentation method, device, equipment and medium based on a lightweight encoding-decoding network, and relates to the field of image processing technology. Background Art
[0002] Semantic segmentation is a key task in computer vision. It aims to label each pixel in an image as one of predefined categories, thereby achieving pixel-level understanding and segmentation of the image. Semantic segmentation has been widely used in fields such as healthcare, remote sensing, and autonomous driving, and is particularly important in autonomous driving. Autonomous driving is a complex task for wheeled robots, requiring intelligent algorithms to assist mobile robots in autonomous motion planning. It requires environmental perception, path planning, and decision-making in dynamic environments. Because safety is paramount in autonomous driving, the highest performance and accuracy are required. Computer vision plays a vital role in autonomous driving, and semantic segmentation, in particular, is crucial for scene recognition and perception. In autonomous driving, semantic segmentation can help vehicles accurately identify key elements such as roads, pedestrians, and vehicles, thereby improving environmental perception and supporting intelligent decision-making. By segmenting images into distinct semantic regions, vehicles can better understand their surroundings, enhance their ability to perceive and respond to traffic situations, and thus enhance driving safety and efficiency. In summary, the application of semantic segmentation in autonomous driving is crucial, providing vehicles with detailed environmental perception capabilities and contributing to a safe and efficient autonomous driving experience.
[0003] Traditional image segmentation methods, such as those based on thresholding, edge detection, region growing, and graph theory, generally perform well in specific scenarios, but have limited effectiveness when dealing with complex images and semantic segmentation tasks. They usually rely on manually designed features and rules and are difficult to generalize to various types of data sets. However, although deep learning methods have achieved great success in image semantic segmentation tasks, they also have some shortcomings, including model complexity, high computing resource requirements, and long training time. Most current semantic segmentation networks based on deep learning have problems such as high computing and storage costs, large scale, long training time, and high computing resource requirements. At the same time, the low efficiency of feature extraction also results in an inference speed that is insufficient to meet the computing power requirements of mobile devices (such as small unmanned vehicles).
[0004] In autonomous driving scenarios, traditional image segmentation methods usually perform well in specific scenarios, but have limited effectiveness in complex image and semantic segmentation tasks. These methods rely on manually designed features and are difficult to adapt to different datasets. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To address the above-mentioned problems, the present invention aims to provide a semantic segmentation method, apparatus, device, and medium based on a lightweight encoding-decoding network, which can achieve higher classification accuracy and more detailed object contour prediction, thereby achieving a better balance between speed and accuracy.
[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0007] In a first aspect, the present invention provides a semantic segmentation method based on a lightweight encoding-decoding network, comprising:
[0008] Collect image data;
[0009] Preprocessing the collected image data;
[0010] The preprocessed image is fed into the encoder for low-resolution feature extraction;
[0011] The extracted low-resolution features are input into the decoder to restore feature details and obtain a feature map with the same resolution as the original image;
[0012] The feature map with the same resolution as the original image is input into the semantic segmentation classifier to obtain pixel-by-pixel classification predictions.
[0013] In some possible implementations, the encoder includes a feature extraction module, and the specific process implemented by the feature extraction module is as follows:
[0014] First, two ordinary convolutions are used in sequence to extract features. The number of channels of the first convolution layer is set to 1 / 2 of the target number of channels, and the second convolution layer is set to 1 / 4 of the total number of channels. The formula is:
[0015]
[0016] Where I is the input image, is the feature output by the first ordinary convolution, is the feature of the second ordinary convolution output, F is the ordinary convolution operation, w 3×3 Indicates that the convolution kernel size is 3*3, and N is the number of target channels;
[0017] Then, a deep convolution with a dilation rate is used to expand the receptive field to generate features with 1 / 4 of the total number of channels. The number of channels of the feature map output by this layer is 1 / 4 of the target number of channels. The formula is:
[0018]
[0019] in, is the feature output by the deep convolution with a dilation rate, FDW It is a depth convolution with dilation rate;
[0020] Next, skip connections and concatenation operations are used to concatenate the three feature maps mentioned above to obtain a feature map with a dimension equal to the total number of target channels. The formula is:
[0021]
[0022] Among them, F concat represents the feature concatenation operation, is the feature output after splicing;
[0023] Finally, point convolution is used to perform channel information interaction on features, and the formula is:
[0024]
[0025] Among them, w 1×1 Indicates that the convolution kernel size is 1*1, and F represents the point convolution operation.
[0026] In some possible implementations, the encoder further includes an attention mechanism module, where the attention mechanism module includes a spatial attention mechanism. The spatial attention mechanism adopts a multi-scale pyramid spatial attention mechanism, and the specific implementation process is as follows:
[0027] First, the input feature map is processed using a max pooling layer to obtain the maximum value of each channel. In parallel, the same image feature map is processed using an average pooling layer to extract the mean of each channel.
[0028] Secondly, the obtained maximum pooling features and average pooling features are spliced to form a new feature map;
[0029] Then, the concatenated feature map is input into the point convolution to adjust the number of feature channels;
[0030] Finally, the Sigmoid activation function is used to calculate the weight map of spatial pixels, and its formula is:
[0031]
[0032] in, and Represents the average pooling and maximum pooling of the Nth layer, w 1×1 Indicates that the convolution kernel size is 1*1, σ indicates the Sigmoid activation function, F concat Represents feature concatenation operation, SA N represents the spatial attention of the Nth layer, and F is the normal convolution operation.
[0033] In some possible implementations, the encoder further includes a channel attention mechanism, wherein the specific implementation process of the channel attention mechanism is as follows:
[0034] First, the global pooling layer is used to pool the entire input feature map to extract the global features of each channel. At the same time, the adaptive pooling layer is used to process the input feature map to adaptively output a fixed-size feature map according to the size of the input feature map.
[0035] Secondly, for the features of the global pooling branch, 1*1 convolution is used for channel information interaction, and for the features of the adaptive pooling branch, 2*2 convolution is used for adaptive weight learning of spatial positions;
[0036] Then, the multi-scale channel attention weight map is calculated through the Sigmoid activation function;
[0037] Finally, based on the multi-scale channel attention weight map, the features of the global pooling branch and the adaptive pooling branch are added in the channel dimension. The formula is:
[0038] CA=σ(F(f global ,w 1×1 ))+σ(F DW (f adapt ,w 2×2 ));
[0039] Among them, σ represents the Sigmoid activation function, CA is the multi-scale channel attention weight map, F is the ordinary convolution operation, F DW is the depthwise convolution operation, w 1×1 ,w 2×2 They indicate that the convolution kernel sizes are 1*1 and 2*2 respectively.
[0040] In some possible implementations, the decoder includes an upsampling module, which uses bilinear interpolation or deconvolution technology to restore the low-resolution feature map to a feature map of a target size.
[0041] In some possible implementations, the decoder further includes a multi-scale fusion module. The specific implementation process of the multi-scale fusion module is as follows:
[0042] First, low-resolution feature maps of different scales are obtained. The low-resolution feature maps are restored to the same size as the higher-resolution feature maps through the upsampling module. The formula is:
[0043] f 22 =F(f 21 ,w 1×1 ),f 22 =F up (F(f 31 ,w1×1 ));
[0044] Among them, F is the ordinary convolution operation, F up is the upsampling operation, w 1×1 They represent the convolution kernel size of 1*1 ordinary convolution, f 31 、f 21 The features are expressed as 1 / 8 image resolution and 1 / 4 image resolution respectively;
[0045] Then, the above feature maps are spliced, depth-wise convolution and point-wise convolution are performed, and the corresponding attention weights are calculated using the Sigmoid function to generate a global attention weight map. The formula is:
[0046] M msa =σ(F(F DW (F concat (f 22 ,f 33 ),w 3×3 ),w 1×1 ));
[0047] Among them, M msa is the global attention weight, F DW ,F are depth and ordinary convolution respectively, σ is the Sigmoid function;
[0048] Finally, the global attention weight map is upsampled and point-by-point multiplied with the feature map of 1 / 2 image resolution to achieve effective fusion of feature maps of different scales and generate a comprehensive feature map.
[0049] In some possible implementations, the semantic segmentation classifier includes a multi-layer perceptron and a Softmax function, wherein the training process of the semantic segmentation classifier is:
[0050] First, the features output by the decoder are input into the multi-layer perceptron to convert them into the number of output channels corresponding to the number of predicted categories;
[0051] Then, the Softmax function is used to convert the feature vector of each pixel into a probability distribution and calculate the probability of each pixel belonging to each category;
[0052] Finally, by comparing the probabilities of each category, the category with the highest probability is selected as the final classification result of the pixel.
[0053] In a second aspect, the present invention further provides a semantic segmentation device based on a lightweight encoding-decoding network, the device comprising:
[0054] a data acquisition unit configured to acquire image data;
[0055] a preprocessing unit, configured to preprocess the collected image data;
[0056] a feature extraction unit configured to input the preprocessed image into the encoder and perform low-resolution feature extraction;
[0057] A feature recovery unit is configured to input the low-resolution features extracted by the encoder into the decoder to recover the feature details and obtain a feature map with the same resolution as the original image;
[0058] The classification prediction unit is configured to input the feature map with the same resolution as the original image into the semantic segmentation classifier to obtain pixel-by-pixel classification prediction.
[0059] In a third aspect, the present invention also provides an electronic device comprising: at least one processor; and a memory communicatively connected to the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to execute the described method.
[0060] In a fourth aspect, the present invention further provides a computer-readable storage medium storing one or more programs, wherein the one or more programs include computer instructions, and the computer instructions are used to enable a computer to execute the described method.
[0061] The present invention adopts the above technical solution, which has the following characteristics:
[0062] 1. This paper proposes a lightweight semantic encoder and decoder, which can achieve higher category accuracy and more detailed object contour prediction while ensuring model lightweight and real-time inference, thereby achieving a better balance between speed and accuracy.
[0063] 2. The encoder proposed in the present invention can realize efficient coding feature extraction, and can efficiently extract the texture, color and contour information of image features while reducing the amount of calculation.
[0064] 3. The present invention proposes dual-pyramid spatial attention and channel attention, which can effectively enhance the feature representation ability in the image encoding process.
[0065] 4. During the decoding process, the present invention proposes a decoder based on multi-scale feature fusion, which can better restore image feature details and thus improve the final prediction ability of the model.
[0066] In summary, the present invention can be widely applied to autonomous driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. Throughout the drawings, the same reference numerals are used to denote the same components. In the drawings:
[0068] Figure 1 This is a flow chart of a semantic segmentation method according to an embodiment of the present invention;
[0069] Figure 2 This is a flowchart of image preprocessing according to an embodiment of the present invention;
[0070] Figure 3 This is a flowchart of encoder feature extraction according to an embodiment of the present invention;
[0071] Figure 4 This is a schematic diagram of the attention mechanism module of an encoder according to one embodiment of the present invention;
[0072] Figure 5 This is a schematic diagram of a decoder based on multi-scale feature fusion according to an embodiment of the present invention;
[0073] Figure 6 is a schematic diagram of a semantic segmentation predictor according to an embodiment of the present invention;
[0074] Figure 7 FIG. 1 is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0075] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0076] Although the terms first, second, third, etc. can be used in the text to describe multiple elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms can only be used to distinguish an element, component, region, layer or section from another region, layer or section. Unless the context clearly indicates otherwise, terms such as "first", "second" and other numerical terms do not imply order or sequence when used in the text. Therefore, the first element, component, region, layer or section discussed below can be referred to as the second element, component, region, layer or section without departing from the teaching of the example embodiments.
[0077] For ease of description, spatially relative terms may be used herein to describe the relationship of one element or feature relative to another element or feature as shown in the figures, such as "inside," "outside," "inner side," "outer side," "lower," "upper," etc. Such spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures.
[0078] Developing lightweight and efficient semantic segmentation methods is crucial for practical application, as small-scale models can not only improve inference speed and computational efficiency, but also reduce the cost of computing equipment. Lightweight networks are more user-friendly for mobile or embedded devices and have strong application scalability. Lightweight encoding-decoding network structures can reduce the number of model parameters and computational complexity while maintaining high accuracy, thereby achieving efficient deployment on resource-constrained mobile devices. This method can effectively address the challenges of traditional deep learning methods in mobile applications and improve the robustness and practicality of the model. Due to the problem that existing semantic segmentation algorithms cannot achieve a balance between accuracy and speed in autonomous driving scenarios, the semantic segmentation method, device, equipment and medium based on a lightweight encoding-decoding network provided by the present invention include: collecting image data; preprocessing the collected image data; inputting the preprocessed image into an encoder for low-resolution feature extraction; inputting the extracted low-resolution features into a decoder for feature detail recovery to obtain a feature map with the same resolution as the original image; and inputting the feature map with the same resolution as the original image into a semantic segmentation classifier to obtain pixel-by-pixel classification predictions. Therefore, the present invention can achieve higher category accuracy prediction and more detailed object contour prediction while ensuring model lightweight and real-time inference prediction of semantic segmentation tasks, achieving a better trade-off between speed and accuracy.
[0079] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0080] Example 1: Figure 1 As shown, the semantic segmentation method based on the lightweight encoding-decoding network provided in this embodiment includes:
[0081] S1. Use RGB camera to collect image data.
[0082] S2. Perform image preprocessing on the collected RGB image data.
[0083] In this embodiment, Figure 2 As shown in the figure, image preprocessing includes cropping, scaling, padding, flipping, and adjusting brightness and contrast of the image in sequence. The dataset is enhanced through the above operations in order to improve the fitting ability and robustness of the model.
[0084] Furthermore, the specific process of image preprocessing includes:
[0085] Randomly crop the input image to change its scale and perspective. This operation not only increases data diversity but also helps the model better learn features from different regions. The crop size can be adjusted based on the model's input requirements, ensuring that the cropped image still contains important semantic information.
[0086] Randomly scale the image down or up. This process provides the model with information on multiple scales, allowing it to learn features of objects of different sizes during training. This scaling can be achieved by defining a range of scaling factors and randomly selecting the factor to scale.
[0087] For images smaller than the target resolution, padding techniques can be used to ensure that all images are of the same size. Padding can be done by zero padding (filling the image with zeros) or by padding with pixel values similar to the image edges. This process helps maintain the shape of the image and avoids training issues caused by inconsistent sizes.
[0088] Perform random flips in the horizontal dimension to further increase data diversity. By flipping the image horizontally, the model learns the characteristics of left-right symmetric objects, which is crucial for improving semantic segmentation accuracy. Flipping is typically performed with a certain probability to ensure diversity in the training set.
[0089] Randomly brighten and dim the image. This process adjusts the overall brightness of the image by selecting a random brightness factor. This brightness variation allows the model to adapt to scenes with varying lighting conditions, improving its robustness in practical applications.
[0090] This method increases the discernibility of features by randomly adjusting the contrast of an image. This contrast variation can be applied to the brightness and color of an image by selecting a random contrast factor. The increased contrast helps highlight important details in the image, ensuring that the model can better segment different objects.
[0091] S3. Input the preprocessed image into the encoder to extract relevant features such as color, texture, and contour.
[0092] In this embodiment, the encoder includes a feature extraction module and an attention mechanism module. The attention mechanism module further enhances the feature representation capability based on the feature extraction module.
[0093] Further, if Figure 3 As shown in the figure, the feature extraction module includes ordinary convolution, depth convolution with dilation rate, and point convolution. The specific process includes:
[0094] First, two ordinary convolutions are used in sequence to extract features. The number of channels of the first convolution layer can be set to 1 / 2 of the target number of channels, and the second convolution layer can be set to 1 / 4 of the total number of channels. The formula is:
[0095]
[0096] Where I is the input image, is the feature output by the first ordinary convolution, is the feature of the second ordinary convolution output, F is the ordinary convolution operation, w 3×3 Indicates that the convolution kernel size is 3*3, and N is the number of target channels.
[0097] Then, a deep convolution with a dilation rate is used to expand the receptive field to generate features with 1 / 4 of the total number of channels. The number of channels of the feature map output by this layer is 1 / 4 of the target number of channels. The formula is:
[0098]
[0099] in, is the feature output by the deep convolution with a dilation rate, F DW It is a depthwise convolution with dilation rate.
[0100] Next, the three feature maps are concatenated using skip connections and concatenation operations to obtain a feature map with a dimension equal to the total number of target channels. The formula is:
[0101]
[0102] Among them, F concat represents the feature concatenation operation, is the feature output after splicing.
[0103] Finally, point convolution is used to interact channel information between features to further improve feature expression capabilities. The formula is:
[0104]
[0105] Among them, f 1 is the feature output by the first-layer feature extraction module, w 1×1 Indicates that the convolution kernel size is 1*1.
[0106] Further, if Figure 4 As shown in Figure 3, the attention mechanism module includes spatial attention and channel attention. Spatial attention uses a maximum pooling layer and an average pooling layer in parallel to obtain the maximum channel feature and mean channel feature of the image channel dimension. The parallel pooled features are then concatenated and input into a point convolution to adjust the number of feature channels. Finally, a sigmoid function is used to calculate the spatial pixel weight map.
[0107] Specifically, the present invention adopts a multi-scale pyramid spatial attention mechanism to obtain richer spatial information. The specific implementation process is as follows:
[0108] First, the input feature map is processed using a max pooling layer to obtain the maximum value for each channel. This operation can highlight the most significant features in the image, emphasizing those pixels with the highest response in a specific area. In parallel, the same image feature map is processed using an average pooling layer to extract the mean value for each channel. This process helps capture the overall information of the image and reduce sensitivity to noise. The formula is:
[0109]
[0110] Among them, f mean ,f max ∈R 1×h×w , respectively represent average pooling and maximum pooling in the channel dimension, C is the number of categories, u ij is the feature point corresponding to each spatial position.
[0111] Secondly, the maximum pooling features and average pooling features obtained in parallel are concatenated to form a new feature map. The concatenation operation can effectively combine two different feature information, thereby improving the expressive power of the feature.
[0112] The concatenated feature map is then fed into a point convolution to adjust the number of feature channels. This operation helps reduce the dimension of the channel and makes the features more compact and effective.
[0113] Finally, the Sigmoid activation function is used to calculate the weight map of spatial pixels. Through this calculation, a weight value can be assigned to each pixel to reflect its importance in the feature map. This step is a key step in improving the model's ability to perceive key features. Its formula is:
[0114]
[0115] in, and Represents the average pooling and maximum pooling of the Nth layer, w 1×1 Indicates that the convolution kernel size is 1*1, σ indicates the Sigmoid activation function, F concat Represents feature concatenation operation, SA N represents the spatial attention of the Nth layer, and F is the normal convolution operation.
[0116] In summary, this paper adopts a multi-scale pyramid spatial attention mechanism to further enrich the acquisition of spatial information by processing features at different scales. This mechanism enables the model to more comprehensively understand the context of the image and improve its performance in complex scenes. Its formula is:
[0117] SA=P(SA 1 ,SA 2 ,...,SA N );
[0118] Among them, P represents the pooling and summing operation, and SA is the spatial attention weight map.
[0119] Furthermore, the channel attention mechanism of the present invention includes: first, performing global pooling and adaptive pooling in parallel on the spatial position of the feature map, using 1*1 convolution for channel information interaction for the features of the global pooling branch, and using 2*2 convolution for adaptive weight learning of the spatial position for the features of the adaptive pooling branch, and then using the Sigmoid function to calculate the channel weight, and finally performing channel dimension addition operation on the features of the two branches.
[0120] Specifically, the implementation process of the channel attention mechanism of the present invention is:
[0121] First, the global pooling layer performs a pooling operation on the entire feature map to extract the global features of each channel. This operation compresses each channel of the feature map into a single value, representing the importance of the channel in the global scope. At the same time, the feature map is processed using an adaptive pooling layer to adaptively output a fixed-size feature map based on the size of the feature map. This method can flexibly handle feature maps of different input sizes and ensure the effectiveness of subsequent operations. Its formula is:
[0122]
[0123] Among them, f global ∈R 1×1×c is the global pooling of image space dimension, f adapt ∈R 2×2×c Image spatial dimension adaptive pooling, this paper adopts 2*2 adaptive pooling, H, W are the length and width of the image respectively, C is the number of categories, i, j represents the information of the spatial position, u c (i, j) represents the feature point corresponding to a certain category channel.
[0124] Secondly, for the features of the global pooling branch, 1*1 convolution is used for channel information interaction, and for the features of the adaptive pooling branch, 2*2 convolution is used for adaptive weight learning of spatial positions;
[0125] Then, the channel weight is calculated through the Sigmoid activation function. This weight map can reflect the importance of each channel in the feature map, thereby helping the model to better focus on key features.
[0126] Finally, the features of the global pooling branch and the adaptive pooling branch are added in the channel dimension. This step can effectively fuse the features from two different sources and enhance the expressive power of the channel. The formula is:
[0127] CA=σ(F(f global ,w 1×1 ))+σ(F DW (f adapt ,w 2×2 ));
[0128] Among them, σ represents the Sigmoid activation function, CA is the multi-scale channel attention weight map, F is the ordinary convolution operation, F DW is the depthwise convolution operation, w 1×1 ,w 2×2 They indicate that the convolution kernel sizes are 1*1 and 2*2 respectively.
[0129] In this embodiment, after the image is input to the encoder, the final output multi-scale feature f = {f 1 ,f 2,f 3 ,...,f n}, the above features represent features of different depths, and each feature represents the feature output after passing through the feature extraction module of each layer and the attention mechanism module.
[0130] S4. The low-resolution features extracted by the encoder are input into the decoder to restore the feature details, which includes the use of multi-scale fusion and linear interpolation upsampling. Multi-scale fusion can fuse feature information of different network depths and obtain a feature map with the same resolution as the original image through upsampling.
[0131] In this embodiment, the decoder includes an upsampling module and a multi-scale fusion module. The upsampling module is used to restore the resolution of deep image features, thereby providing the same resolution for the fusion of shallow and deep features in the multi-scale fusion module.
[0132] Furthermore, the upsampling module uses bilinear interpolation or deconvolution techniques to restore the low-resolution feature map to the target size feature map. Specifically, the upsampling module uses bilinear interpolation to restore the low-resolution feature map to the target size feature map. The bilinear interpolation formula is:
[0133] f(P)=(1-u)(1-v)f(A)+u(1-v)f(B)+(1-u)vf(C)+uvf(D);
[0134] Among them, (A, f(A)), (B, f(B)), (C, f(C)), and (D, f(D)) are the coordinates and pixel values of four adjacent pixels respectively, and u and v are the ratios of the distances from point P to points A and B in the AB direction and in the CD direction respectively.
[0135] Further, if Figure 5 As shown in the figure, the multi-scale fusion module effectively fuses feature maps from different scales through weighted summation, concatenation or attention mechanism to generate a comprehensive feature map, including:
[0136] First, feature maps of different scales are input into the multi-scale fusion module. The low-resolution feature map (for example, 1 / 8 image resolution) is restored to the same size as the higher-resolution feature map (for example, 1 / 4 image resolution) through upsampling technology. The formula is:
[0137] f 22 =F(f 21 ,w 1×1 ),f 22 =F up (F(f 31 ,w 1×1 ));
[0138] Among them, F is the ordinary convolution operation, F up is the upsampling operation, w 1×1 They represent the convolution kernel size of 1*1 ordinary convolution, f 31 、f 21 Represented as features of 1 / 8 image resolution and 1 / 4 image resolution respectively.
[0139] Next, these feature maps will be spliced, and then through depth convolution and point convolution, and finally the Sigmoid function will be used to calculate the corresponding attention weights to generate a global attention weight map. The formula is:
[0140] M msa =σ(F(F DW (F concat (f 22 ,f 33 ),w 3×3 ),w 1×1 ));
[0141] Among them, M msa is the global attention weight, F DW ,F are depth and ordinary convolution respectively, and σ is the Sigmoid function.
[0142] Finally, the global attention weight map is upsampled and point-by-point multiplied with the feature map of 1 / 2 image resolution to achieve more accurate information fusion. This process not only improves the richness of feature representation, but also enhances the adaptability of the model at different scales. The formula is:
[0143] f out =F up (F SA (F(f 11 ,w 1×1 ))*F up (M msa ));
[0144] Among them, f out is the output feature of the final decoder, F SA is the spatial attention mechanism operation, * is the feature point-by-point multiplication, f 11 They are represented as features of 1 / 2 image resolution respectively.
[0145] S5. Input the feature map with the same resolution as the original image into the semantic segmentation classifier and obtain the pixel-by-pixel classification prediction through the maximum probability.
[0146] In this embodiment, Figure 6As shown in the figure, the features output by the decoder are input into the semantic segmentation classifier for pixel-by-pixel semantic prediction. In the figure, H and W are the length and width of the image respectively, C is the number of categories, and N is the number of feature channels.
[0147] Furthermore, the semantic segmentation classifier is composed of a multi-layer perceptron and a Softmax function. The specific training process is as follows:
[0148] First, the features output by the decoder are fed into a multi-layer perceptron (MLP), which consists of multiple fully connected layers. The main function of these layers is to process the input features and transform them. During this process, the MLP gradually reduces the number of channels in the feature map, ultimately converting it to a number of output channels corresponding to the number of predicted categories.
[0149] Then, the Softmax function is used to convert the feature vector of each pixel into a probability distribution and calculate the probability of each pixel belonging to each category;
[0150] Finally, by comparing the probabilities of each category, the model selects the category with the highest probability as the final classification result for the pixel.
[0151] Furthermore, the present invention adopts cross entropy and focal loss loss functions in the training process of the entire model, and the total loss function formula is:
[0152] L=λ1L c +λ2L f ;
[0153] Among them, λ1 and λ2 are weight coefficients of different loss functions, L c is the cross entropy loss function, L f is the Focal loss function, where:
[0154] The cross entropy loss function formula is as follows: Where y is the actual label (0 or 1), is the probability value predicted by the model.
[0155] The formula of Focal loss function is as follows: L f (p t )=-α t (1-p t ) γ log(p t ), where p t is the model’s predicted probability for the true category, α t is a balancing factor used to adjust the category weights, and γ is a regulating factor that controls the degree of loss reduction of easy-to-classify samples.
[0156] Furthermore, the present invention adopts a difficult example learning (OHEM) strategy during the training process. The formula of the difficult example learning strategy is as follows:
[0157]
[0158] Among them, η is a hyperparameter, P c It is the contribution degree of the category samples, generally expressed as the proportion of the category to the total number.
[0159] Furthermore, the learning rate decay strategy adopted by the present invention during the training process is the ploy exponential decay strategy, and its formula is:
[0160]
[0161] Among them, β is the learning rate at iteration i, β0 is the initial learning rate, i is the number of iterations, i m is the maximum number of iterations.
[0162] Furthermore, the present invention also proposes a trade-off coefficient for evaluating model speed and accuracy to determine the degree of model lightweighting, and the formula is:
[0163]
[0164] Where m i 、s i 、p i are the model parameters corresponding to mIoU, inference speed (FPS) and the i-th iteration of the Bayesian optimization algorithm, w1, w2, w3 are the weight coefficients of mIoU, FPS and Params, ε is the parameter sensitivity suppression coefficient, m base and s base is the baseline value of mIoU and FPS. It can be seen that the smaller the denominator corresponds to the model parameter in the formula, the larger the mIoU and inference speed corresponding to the numerator, and the corresponding F b The larger the value, the better the balance of the model.
[0165] Example 2: The above-mentioned Example 1 provides a semantic segmentation method based on a lightweight encoding-decoding network. Correspondingly, this embodiment provides a semantic segmentation device based on a lightweight encoding-decoding network. The device provided in this embodiment can implement the semantic segmentation method based on a lightweight encoding-decoding network of Example 1, and the device can be implemented by software, hardware, or a combination of software and hardware. For the convenience of description, this embodiment is described by dividing the functions into various units and describing them separately. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware during implementation. For example, the device may include integrated or separate functional modules or functional units to perform the corresponding steps in each method of Example 1. Since the device of this embodiment is basically similar to the method embodiment, the description process of this embodiment is relatively simple. For relevant points, please refer to the partial description of Example 1. The embodiment of the semantic segmentation device based on a lightweight encoding-decoding network provided by the present invention is only schematic.
[0166] Specifically, the present invention also provides a semantic segmentation device based on a lightweight encoding-decoding network, the device comprising:
[0167] a data acquisition unit configured to acquire image data;
[0168] a preprocessing unit, configured to preprocess the collected image data;
[0169] a feature extraction unit configured to input the preprocessed image into the encoder and perform low-resolution feature extraction;
[0170] A feature recovery unit is configured to input the low-resolution features extracted by the encoder into the decoder to recover the feature details and obtain a feature map with the same resolution as the original image;
[0171] The classification prediction unit is configured to input the feature map with the same resolution as the original image into the semantic segmentation classifier to obtain pixel-by-pixel classification prediction.
[0172] Example 3: This example provides an electronic device corresponding to the semantic segmentation method based on a lightweight encoding-decoding network provided in Example 1. The electronic device may be an electronic device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the method of Example 1.
[0173] like Figure 7As shown, the electronic device includes a processor, a memory, a communication interface and a bus. The processor, the memory and the communication interface are connected via the bus to complete communication between them. The memory stores a computer program that can be run on the processor. When the processor runs the computer program, it executes the method of embodiment 1. Its implementation principle and technical effect are similar to those of embodiment 1 and will not be repeated here. It can be understood by those skilled in the art that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computing device to which the solution of the present application is applied. The specific computing device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0174] In a preferred embodiment, the logic instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), optical disk and other media that can store program code.
[0175] In a preferred embodiment, the processor may be a central processing unit (CPU), a digital signal processor (DSP), or other general-purpose processors of various types, which are not limited herein.
[0176] Embodiment 4: This embodiment provides a computer-readable storage medium storing one or more programs, wherein the one or more programs include computer instructions. When the computer instructions are executed by a computer, the computer executes the method provided in the above embodiment 1.
[0177] In a preferred embodiment, a computer-readable storage medium may be a tangible device that retains and stores instructions executed by the computer, such as, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof. The computer-readable storage medium stores computer program instructions that cause a computer to execute the method provided in the first embodiment.
[0178] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (apparatus), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0179] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0180] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0181] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In the description of this specification, the reference terms "a preferred embodiment", "further", "specifically", "in the present embodiment", etc. mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of this specification. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A semantic segmentation method based on a lightweight encoder-decoder network, characterized in that: include: Collect image data; Preprocessing the collected image data; The preprocessed image is fed into the encoder for low-resolution feature extraction; The extracted low-resolution features are input into the decoder to restore feature details and obtain a feature map with the same resolution as the original image; The feature map with the same resolution as the original image is input into the semantic segmentation classifier to obtain pixel-by-pixel classification predictions.
2. The semantic segmentation method based on a lightweight encoding-decoding network according to claim 1, characterized in that The encoder includes a feature extraction module. The specific process implemented by the feature extraction module is as follows: First, two ordinary convolutions are used in sequence to extract features. The number of channels of the first convolution layer is set to 1 / 2 of the target number of channels, and the second convolution layer is set to 1 / 4 of the total number of channels. The formula is: Among them, I is the input image, f1 1 is the feature output by the first ordinary convolution, is the feature of the second ordinary convolution output, F is the ordinary convolution operation, w 3×3 Indicates that the convolution kernel size is 3*3, and N is the number of target channels; Then, a deep convolution with a dilation rate is used to expand the receptive field to generate features with 1 / 4 of the total number of channels. The number of channels of the feature map output by this layer is 1 / 4 of the target number of channels. The formula is: in, is the feature output by the deep convolution with a dilation rate, F DW It is a depth convolution with dilation rate; Next, skip connections and concatenation operations are used to concatenate the three feature maps mentioned above to obtain a feature map with a dimension equal to the total number of target channels. The formula is: Among them, F concat represents the feature concatenation operation, is the feature output after splicing; Finally, point convolution is used to perform channel information interaction on features, and the formula is: Among them, w 1×1 Indicates that the convolution kernel size is 1*1, and F represents the point convolution operation.
3. The semantic segmentation method based on a lightweight encoding-decoding network according to claim 2, characterized in that The encoder also includes an attention mechanism module, which includes a spatial attention mechanism. The spatial attention mechanism adopts a multi-scale pyramid spatial attention mechanism. The specific implementation process is as follows: First, the input feature map is processed using a max pooling layer to obtain the maximum value of each channel. In parallel, the same image feature map is processed using an average pooling layer to extract the mean of each channel. Secondly, the obtained maximum pooling features and average pooling features are concatenated to form a new feature map; Then, the concatenated feature map is input into the point convolution to adjust the number of feature channels; Finally, the Sigmoid activation function is used to calculate the weight map of spatial pixels, and its formula is: in, and Represents the average pooling and maximum pooling of the Nth layer, w 1×1 Indicates that the convolution kernel size is 1*1, σ indicates the Sigmoid activation function, F concat Represents feature concatenation operation, SA N represents the spatial attention of the Nth layer, and F is the normal convolution operation.
4. The semantic segmentation method based on a lightweight encoding-decoding network according to claim 2, characterized in that The encoder also includes a channel attention mechanism, where The specific implementation process of the channel attention mechanism is: First, the global pooling layer is used to pool the entire input feature map to extract the global features of each channel. At the same time, the adaptive pooling layer is used to process the input feature map to adaptively output a fixed-size feature map according to the size of the input feature map. Secondly, for the features of the global pooling branch, 1*1 convolution is used for channel information interaction, and for the features of the adaptive pooling branch, 2*2 convolution is used for adaptive weight learning of spatial positions; Then, the multi-scale channel attention weight map is calculated through the Sigmoid activation function; Finally, based on the multi-scale channel attention weight map, the features of the global pooling branch and the adaptive pooling branch are added in the channel dimension. The formula is: CA=σ(F(f global ,w 1×1 ))+σ(F DW (f adapt ,w 2×2 )); Among them, σ represents the Sigmoid activation function, CA is the multi-scale channel attention weight map, F is the ordinary convolution operation, F DW is the depthwise convolution operation, w 1×1 ,w 2×2 They indicate that the convolution kernel sizes are 1*1 and 2*2 respectively.
5. The semantic segmentation method based on a lightweight encoding-decoding network according to claim 2, characterized in that The decoder includes an upsampling module, which uses bilinear interpolation or deconvolution technology to restore the low-resolution feature map to the feature map of the target size.
6. The semantic segmentation method based on a lightweight encoding-decoding network according to claim 5, characterized in that The decoder also includes a multi-scale fusion module. The specific implementation process of the multi-scale fusion module is as follows: First, low-resolution feature maps of different scales are obtained. The low-resolution feature maps are restored to the same size as the higher-resolution feature maps through the upsampling module. The formula is: f 22 =F(f 21 ,w 1×1 ),f 22 =F up (F(f 31 ,w 1×1 )); Among them, F is the ordinary convolution operation, F up is the upsampling operation, w 1×1 They represent the convolution kernel size of 1*1 ordinary convolution, f 31 、f 21 The features are expressed as 1 / 8 image resolution and 1 / 4 image resolution respectively; Then, the above feature maps are spliced, depth-wise convolution and point-wise convolution are performed, and the corresponding attention weights are calculated using the Sigmoid function to generate a global attention weight map. The formula is: M msa =σ(F(F DW (F concat (f 22 ,f 33 ),w 3×3 ),w 1×1 )); Among them, M msa is the global attention weight, F DW ,F are depth and ordinary convolution respectively, σ is the Sigmoid function; Finally, the global attention weight map is upsampled and point-by-point multiplied with the feature map of 1 / 2 image resolution to achieve effective fusion of feature maps of different scales and generate a comprehensive feature map.
7. The semantic segmentation method based on a lightweight encoding-decoding network according to claim 1, characterized in that The semantic segmentation classifier includes a multi-layer perceptron and a Softmax function. The training process of the semantic segmentation classifier is as follows: First, the features output by the decoder are input into the multi-layer perceptron to convert them into the number of output channels corresponding to the number of predicted categories; Then, the Softmax function is used to convert the feature vector of each pixel into a probability distribution and calculate the probability of each pixel belonging to each category; Finally, by comparing the probabilities of each category, the category with the highest probability is selected as the final classification result of the pixel.
8. A semantic segmentation device based on a lightweight encoding-decoding network, characterized in that: The device includes: a data acquisition unit configured to acquire image data; a preprocessing unit, configured to preprocess the collected image data; a feature extraction unit configured to input the preprocessed image into the encoder and perform low-resolution feature extraction; A feature recovery unit is configured to input the low-resolution features extracted by the encoder into the decoder to recover the feature details and obtain a feature map with the same resolution as the original image; The classification prediction unit is configured to input the feature map with the same resolution as the original image into the semantic segmentation classifier and obtain the pixel-by-pixel classification prediction through the maximum probability.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to perform the method according to any one of claims 1-7.
10. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include computer instructions for causing a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
On-satellite application-oriented lightweight remote sensing image learning type compression method and system
CN121711490A