A road segmentation method with contour enhancement via multi-path compact transmission network
By adopting the encoder-decoder structure of a multi-channel compact transmission network in road segmentation, the problems of inaccuracy and noise interference of road segmentation results in the prior art are solved, and the accuracy of road profile feature extraction and segmentation results are achieved.
Patent Information
- Application Number
- CN202111204700.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-10-15
AI Technical Summary
The existing deep learning-based road segmentation method has problems of high error rates and noise interference when dealing with road feature fragmentation caused by spectral value anomalies, occlusion, and intra-class inconsistencies.
Using a multi-channel compact transmission network, an encoder-decoder structure is designed, including a multi-path encoding module (MPEM), a densely connected fusion block (DCFB) and a noise suppression block (NSB) to enhance road profile characteristics and reduce noise.
Through the integration of multipath encoding modules and the design of densely connected fusion blocks, the network can better extract and restore road profile features, reduce error rates and improve the accuracy of segmentation results.
Smart Images

Figure CN114119616B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of image segmentation, and in particular to a road segmentation method for contour enhancement through a multi-path compact transmission network. Background Art
[0002] Road segmentation is a basic technology in remote sensing image (RSI) processing. This technology has a large number of applications in intelligent transportation systems (ITS), road monitoring, GPS vehicle navigation, infrared search and tracking, etc. For example, in ITS, mapping has a huge demand for accurate road information, so a faster and more accurate road segmentation method is needed. However, although the research on road segmentation has been carried out for decades, its solution has not yet reached a satisfactory level. In traditional methods based on classification, morphology and dynamic programming, the selection of parameters directly affects the segmentation results of the model, resulting in insufficient generalization ability of the model.
[0003] With the great achievements of convolutional neural networks (CNN) in the field of image segmentation, CNN-based methods have gradually occupied a mainstream position in the field of road segmentation since 2015. The research on segmentation tasks using deep learning methods originated from the fully convolutional network (FCN). The FCN composed of convolutional layers and pooling layers can achieve input of any size. It also implements end-to-end and pixel-to-pixel training, and learns the mapping from pixel to pixel to predict each pixel of an image. However, the segmentation results of FCN cannot determine the specific contours of the object well. Therefore, there are some other improvements, such as generative adversarial networks, shallow convolutional neural networks, deep convolutional neural networks, and recursive neural networks. Although these methods have innovated on the original basis and solved some practical problems, new problems have also emerged.
[0004] At present, there are still some problems with the road segmentation method based on deep learning: spectral value anomalies, road feature fragmentation caused by occlusion, and intra-class inconsistency. For example, the most common phenomenon of spectral value anomalies is that objects of the same type have different spectral values. Specifically, roads in RSI show different spectral characteristics when weather or lighting conditions change. This problem can lead to errors in road recognition. In addition, occlusion can also cause the loss of road structure information. In indirect occlusion images, roads are only misidentified. In direct occlusion images, the structural information of roads is more destroyed, which will cause a high error rate. In addition, intra-class inconsistency often appears as noise in the segmentation results, which is caused by the significant increase in RSI resolution. In high-resolution RSI, low-resolution unrecognized areas will be clearer. For example, high-brightness areas, edges, and pixel-level noise contained in unrecognized areas will interfere with the road recognition process. Summary of the invention
[0005] The present invention provides a road segmentation method for contour enhancement through a multi-channel compact transmission network. The proposed network is an efficient encoder-decoder structure. In order to enhance the relevance of semantic details, two modules, DCFB and NSB, are designed in it. In addition, in order to maintain refined spatial details, U-Net is selected as the basic network.
[0006] In order to achieve the above purpose, the technical solution of this application is as follows:
[0007] A road segmentation method with contour enhancement via a multi-path compact transmission network, comprising:
[0008] Use a public dataset and preprocess the images in the dataset to obtain a training dataset;
[0009] An image segmentation model constructed by training with a training data set, the image segmentation model includes an encoder, a context transmission path and a decoder, the encoder includes a plurality of stage multi-path encoding modules connected in sequence, the context transmission path includes dense connection fusion blocks respectively connected to the multi-path encoding modules, the decoder includes noise suppression blocks respectively connected to the dense connection fusion blocks, and a feature fusion module connected to the noise suppression block, the output of the multi-path encoding module of the last stage of the encoder is also connected to the noise suppression block and feature fusion module of the same stage after global pooling, and the output of the feature fusion module of each stage is connected to the noise suppression block and feature fusion module of the previous stage;
[0010] Input the image to be segmented into the trained image segmentation model and output the segmentation result.
[0011] Furthermore, the input of the multi-path encoding module passes through the first path, the second path and the third path. The outputs of the three paths are fused and then output after 1×1 convolution, BatchNorm, and LeakyReLu activation functions. The first path is used to maintain the position invariance and rotation invariance of road spectral features, the second path is used to prevent the network gradient from exploding or disappearing, and the third path is used to associate spectral features with other features.
[0012] Furthermore, the input of the multi-path encoding module of the last stage of the encoder passes through the first path and the second path, and the outputs of the two paths are fused and then output after passing through 1×1 convolution, BatchNorm, and LeakyReLu activation functions. The first path is used to maintain the position invariance and rotation invariance of the road spectral features, and the second path is used to prevent the network gradient from exploding or disappearing.
[0013] Furthermore, the first path is PoolPath, which includes 1×1 convolution, BatchNorm, LeakyReLu activation function and Maxpool in sequence; the second path is ResPath, which includes a Resnet_18 module; and the third path is ConvPath, which includes a convolution kernel and BatchNorm.
[0014] Furthermore, the convolution kernel of the third path decreases with the stage.
[0015] Furthermore, the densely connected fusion block includes a first fusion module FP1, a DP module and a second fusion module FP2; the first fusion module FP1 merges the information of the current stage with the information of the previous stage; the DP module includes three identical components, each component includes a convolution kernel, a BatchNorm and a ReLu activation function, and short paths are used to connect the components; the second fusion module FP2 is used for feature output, outputs Ouput1 and Ouput2, and outputs them to the noise suppression block of the current stage and the densely connected fusion block of the next stage respectively.
[0016] Furthermore, the noise suppression block includes a spatial attention stage and a channel noise suppression stage. The spatial attention stage first performs 1*1 convolution, LeakyRelu activation function, 3*3 convolution and Sigmoid processing on the output of the context transmission path in sequence to generate a spatial attention map, and multiplies and fuses the spatial attention map with the output of the noise suppression block in the previous stage, and outputs it to the channel noise suppression stage; the channel noise suppression stage processes the input features through AvgPool and Sigmoid processing, and then multiplies and fuses them with the input features, and then adds and fuses them with the input features as the final output.
[0017] The present invention proposes a road segmentation method with contour enhancement through a multi-path compact transmission network. First, the encoder is enabled to aggregate internal information at each stage and extract road spectral features, and MPEM is introduced to overcome the shortcomings of the encoder. In addition, an implicitly guided DCFB is proposed, which compensates for the loss of contour features through short paths and restores the occluded area using the guidance of shallow information. Finally, NSB establishes feature channel and feature space relationships for noise removal. Through these three modules, the network makes up for the shortcomings of semantic transmission, further refines the road contour features, and improves the accuracy of road segmentation results. In order to overcome the shortcomings of single-path feature extraction, a multi-path encoding module (MPEM) is used to integrate the multi-path encoding results, while improving the robustness of the encoder and the fault tolerance of the algorithm. A compact transmission dense connection fusion block (DCFB) is designed to fuse the semantic and spatial information of each stage to improve the obstacle recognition ability of the network. In order to reduce the generation of noise, channel and spatial global information are further mapped to select the optimal features, and a noise suppression block (NSB) is designed. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flow chart of a road segmentation method for contour enhancement by a multi-channel compact transmission network in this application;
[0019] Figure 2 This is a schematic diagram of the image segmentation model structure of this application;
[0020] Figure 3 It is a structural diagram of the multi-path encoding module;
[0021] Figure 4 It is a structural diagram of densely connected fusion blocks;
[0022] Figure 5 Figure 2 is the structural diagram of the noise suppression block. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0024] This application provides a road segmentation method for contour enhancement through a multi-path compact transmission network, such as Figure 1 As shown, including:
[0025] Step S1: Use a public data set and preprocess the images in the data set to obtain a training data set.
[0026] The training images come from publicly available datasets for road segmentation: Massachusetts Road dataset and DeepGlobe dataset. Before training, the images in the dataset are augmented by data. The specific data augmentation methods used include random cutting, vertical flipping, horizontal flipping, padding, and image size normalization. Finally, the images are divided into training and test sets.
[0027] Step S2, using a training data set to train and construct an image segmentation model, the image segmentation model includes an encoder, a context transmission path and a decoder, the encoder includes a plurality of stage multi-path encoding modules connected in sequence, the context transmission path includes densely connected fusion blocks respectively connected to the multi-path encoding modules, the decoder includes noise suppression blocks respectively connected to the densely connected fusion blocks, and a feature fusion module connected to the noise suppression blocks, the output of the multi-path encoding module of the last stage of the encoder is also globally pooled and connected to the noise suppression block and feature fusion module of the same stage, and the output of the feature fusion module of each stage is connected to the noise suppression block and feature fusion module of the previous stage.
[0028] The backbone structure of the image segmentation model network in this application is an encoder, a context transfer path, and a decoder structure. The input image will undergo encoding at each stage. In addition to being used for encoding in the next stage, the encoded features are also sent to the decoder through the context transfer path to guide the network's decoding. Finally, the encoded image undergoes global pooling to fuse deep overall features, and as the decoder decodes, the segmented image is output.
[0029] like Figure 2 As shown, the top layer is the encoder, the middle layer (Context Transmission Pathway) is the context transmission path, and the bottom layer is the decoder.
[0030] The encoder includes a plurality of stages of multi-path encoding modules MPEM connected in sequence. In a preferred embodiment, the multi-path encoding module of the last stage is a simple multi-path encoding module MPEM_D. Figure 2 In the embodiment, the encoder includes a multi-path encoding module of 5 stages, where S1 to S5 represent each stage, from a low-level stage to a high-level stage.
[0031] The encoder described in the present application adopts a multi-path encoding module (Multi-path Encoding Module, MPEM), and the MPEM is located at the beginning of each stage of the network encoder, and MPEM_D is its derivative.
[0032] like Figure 3As shown, the image segmentation model of the present application adopts MPEM as the main encoder, and the input passes through the first path, the second path and the third path. The outputs of the three paths are fused and then output after 1×1 convolution, BatchNorm, and LeakyReLu activation functions. The first path is used to maintain the position invariance and rotation invariance of road spectral features, the second path is used to prevent network gradients from exploding or disappearing, and the third path is used to associate spectral features with other features. The input of MPEM_D passes through the first path and the second path. The outputs of the two paths are fused and then output after 1×1 convolution, BatchNorm, and LeakyReLu activation functions. The first path is used to maintain the position invariance and rotation invariance of road spectral features, and the second path is used to prevent network gradients from exploding or disappearing.
[0033] Specifically, the MPEM module includes three paths: PoolPath, ResPath, and ConvPath. PoolPath enables the network to maintain the position invariance and rotation invariance of road spectral features when it runs to a higher stage. It uses a convolutional layer to modify the number of feature channels, performs a maximum pooling operation on it, and then outputs the spectral position features. The pooling path enables the network to maintain the position invariance and rotation invariance of road spectral features when it runs to a higher stage. ConvPath associates spectral features with other features. For feature maps of different sizes after encoding, it chooses to use receptive fields of different sizes to extract features. The final output features contain more detailed global information of spectral features and other features. ResPath is used to prevent network gradients from exploding or disappearing. It uses the network architecture of Resnet_18, whose parameter weights are pre-trained by ImageNet, to achieve faster convergence and give the network a certain degree of fault tolerance.
[0034] PoolPath includes: 1×1 convolution, BatchNorm, LeakyReLu activation function and Maxpool. A 1×1 convolution kernel is used to increase the image channels. This operation increases the proportion of road spectral features in channel fusion and compensates for the spectral features lost after encoding. BatchNorm normalizes the training parameters. In order to prevent the weight and bias parameters from being updated this time, this application uses the LeakyReLu activation function. Finally, Maxpool is used for downsampling to obtain a thumbnail of the input image, retaining the position of the road spectral features.
[0035] ConvPath includes: convolution kernel and BatchNorm. The convolution kernel decreases with the stage. That is, in the low-level stage, the convolution kernel uses a large receptive field convolution to obtain more comprehensive global feature information; in the high-level stage, a small receptive field convolution is used to obtain more local and detailed features. Figure 2 As shown, this path uses different convolutions from low-level to high-level with kernels 7×7, 5×5, 3×3, 2×2, and before the end of this path, BatchNorm is used. The purpose of this path is to extract more detailed global information. For the fifth stage of MPEM_D, this ConvPath path is removed in pursuit of stability.
[0036] ResPath includes: Resnet_18 module, which uses the Resnet_18 network architecture, and its parameter weights are pre-trained by ImageNet, which can achieve faster convergence speed. Therefore, the pre-trained parameters give the network a certain degree of fault tolerance.
[0037] The outputs of the three paths PoolPath, ResPath, and ConvPath are fused and then output after 1×1 convolution, BatchNorm, and LeakyReLu activation functions.
[0038] Since the maximum pooling will reduce the size of the feature map, but the use of different kernels in ConvPath and ResPath will also reduce the size of the feature map to the same extent. Therefore, the output of each path can be fused together through the "C" cascade. The cascaded features are processed by 1×1 convolution, and the same BatchNorm and LeakyReLu are used before output. This operation realizes cross-channel interaction and fusion of information, reducing the number of convolution channels and subsequent calculations.
[0039] In the encoder, the model of this application uses MPEM, which combines the results of multiple paths and collects the encoding features of each stage. MPEM not only makes up for the shortcomings of the UNetPPL method that ignores the internal semantic features of each stage and makes these features relatively independent, but also makes the internal semantic feature information of each stage correlated with each other, and improves the recognition of road branches that require more internal correlation features.
[0040] The final encoded image undergoes global pooling (Global Model) to fuse deep overall features, and as the decoder decodes, the segmented image is output.
[0041] The U-Net in the prior art only transmits coding characteristics to the decoder without any processing of the coding and decoding characteristics. This not only ignores the correlation of spatial information between stages, but also fails to accurately extract detailed information on the edges of the road. These are possible reasons for the breakage of the road segmentation results. In A-DenseUNet, the context transfer path is redesigned to build an adaptive network and pass the lost features to the decoder. The context transfer path of this application adopts a densely connected fusion block (Densely Connected Fusion Block, DCFB), which is located between the encoder and the decoder of this stage. The DCFB corresponds to the NSB connecting the decoder and the DCFB of the next stage, and DCB and DCFB_1 are both derivatives thereof.
[0042] like Figure 4 As shown, the image segmentation model of the present application adopts DCFB as the main context transmission path part, and fuses the context information to provide guidance for the decoder: the densely connected fusion block includes a first fusion module FP1, a DP module and a second fusion module FP2; the first fusion module FP1 merges the information of the current stage with the information of the previous stage; the DP module includes three identical components, each component includes a convolution kernel, a BatchNorm and a ReLu activation function, and short paths are used to connect the components; the second fusion module FP2 is used for feature output, outputs Ouput1 and Ouput2, and outputs them to the noise suppression block of the current stage and the densely connected fusion block of the next stage respectively.
[0043] The dense part (DP) of DCFB adopts a short path to enhance the contour features of the road by compensating for the loss of contour features in the input source. The fusion part (FP1, FP2) uses deep semantic information to compensate for the road feature information lost in the shallow layer, restore the occluded area, and enhance the integrity of the road. In order to compactly transmit context features, the feature map is output to the DCFB of the next stage and the NSB of the current stage.
[0044] DP consists of three identical components, and the features in each component are extracted by an internal 3×3 convolution, and the components also contain BatchNorm and ReLu for correction. The input features of DP are also reused between components through short paths. The short path enables compact feature transmission, and each component directly accesses the gradient from the original input signal to generate implicit guidance.
[0045] Taking into account the global image features, DCFB also uses shallow information to compensate for the missing spatial information in deep information, so that the network can better identify the spatial information around the road and optimize the defects of road connections. The FP1 module merges the information of the current stage with the information of the previous stage. FP1 performs information fusion at different stages from the perspective of input. FP2 is used for feature output, and Ouput2 uses 1×1 convolution and Maxpool before output to ensure that the size and number of channels of Output2 are the same as the next DCFB input feature. Ouput1 and Ouput2 are used for NSB in the current stage and DCFB in the next stage respectively.
[0046] In the input part, DCB only has the current stage feature, which is located in the context transfer path of the first stage, while the output part of DCFB_1 only has ouput1, which is the context transfer path of the fifth stage.
[0047] The existing decoder part, such as RIC-Unet, uses a channel attention mechanism to combine features of different resolutions. This mechanism only focuses on channel information and ignores spatial information, resulting in some residual noise in the road segmentation result. The encoder of this application uses a noise suppression block (NSB), which is located after the decoding operation of the previous stage.
[0048] like Figure 5 As shown, the image segmentation model of the present application adopts the noise suppression block NSB as a decoding auxiliary part to suppress the noise of the deconvolved image. It can fully suppress the noise through the channel and spatial suppression functions, selectively enhance the required features, and ignore the interference information.
[0049] The noise suppression block includes a spatial attention stage and a channel noise suppression stage. The spatial attention stage first performs 1*1 convolution, LeakyRelu activation function, 3*3 convolution and Sigmoid processing on the output of the context transmission path in sequence to generate a spatial attention map, and multiplies and fuses the spatial attention map with the output of the noise suppression block in the previous stage, and outputs it to the channel noise suppression stage; the channel noise suppression stage performs AvgPool (pooling) and Sigmoid processing on the input features, and then multiplies and fuses them with the input features, and then adds and fuses them with the input features as the final output.
[0050] NSB’s spatial noise suppression uses the spatial feature information transmitted in shallow layers, recovers the missing feature map space after decoding, strengthens the connection between road spatial information, and eliminates discrete noise. Then NSB’s channel noise suppression mechanism maps the global features in the channel and selects the channel with the best features. Combining these two parts, NSB not only suppresses image noise but also enhances intra-class consistency. Spatial noise suppression uses contextual transmission features to generate spatial attention maps, which helps the network focus on important spatial locations, suppress spatial noise, and highlight road information. 1×1 convolution flattens the contextual transmission features and 3×3 with LeakyReLu for correction. Sigmoid activation function generates spatial attention maps. Spatial noise is suppressed based on spatial context information, and then waits for channel noise suppression processing.
[0051] Channel noise suppression computes a channel attention map and assigns higher weights to more useful channels. This establishes a relationship between feature maps and channel information. Global pooling is used to encode a single channel into a single feature. Sigmoid activation function obtains the channel attention map. Channel noise suppression combines the channel attention map with the features output by spatial noise suppression to ensure that the desired channels have higher weights.
[0052] The image segmentation model of this application adopts a multi-path encoding module (MPEM) with a parallel multi-path encoding structure to solve the problem of spectral value anomalies, and proposes a dense connection fusion block (DCFB) that is tightly transmitted throughout the entire network to solve the occlusion problem. A noise suppression block (NSB) is established to further map channel and spatial global information to select the optimal features and reduce the generation of noise to solve the problem of intra-class inconsistency.
[0053] In this application, road segmentation can be viewed as a binary classification problem. Therefore, the loss function uses binary cross entropy, which describes the difference between the ground truth and the model output. The binary cross entropy is defined as:
[0054]
[0055] Where n represents the number of images, y i represents ground truth, is the actual output of the network.
[0056] Step S3: input the image to be segmented into the trained image segmentation model and output the segmentation result.
[0057] After the image segmentation model is trained, the picture to be segmented can be input into the trained image segmentation model, and the image output by the decoder is the segmentation result.
[0058]
[0059] Table 1 shows the experimental results of different road extraction methods on the DeepGlobe dataset.
[0060] By comparing the experimental data of the DeepGlobe dataset, the network structure proposed in this application has achieved good results in most metrics. More intuitively, ours's Precision, F1-score and IoU are 0.14%, 1.95% and 1.6% higher than the latest ScRoadExtractor, respectively. In the road segmentation task, we can clearly see that the metrics of the network with a dual decoder structure are not better than those with a single decoder structure. In addition, the Precision, F1-score and IoU of MCTC-Net are 8.53%, 6.54% and 7.03% higher than those of weaklyOS, respectively. In addition, the network structure proposed in this application outperforms ScribbleSup, BPG and WSOD in terms of Precision, F1-score and IoU. By comparing the encoder-decoder structure (ScRoadExtractor and WeaklyOS) with the non-encoder-decoder structure (ScribbleSup, BPG and WSOD), it is clear that the encoder-decoder structure performs better in the road segmentation task, which also proves that full image semantics is more conducive to road segmentation.
[0061]
[0062] Table 2 shows the experimental results of different road extraction methods on the Massachusetts dataset.
[0063] In order to verify the generalization of the network, the road extraction results are shown on the Massachusetts dataset. The Recall, Precision and F1-score of this application are 2.28%, 1.05% and 1.13% higher than those of DenseUNet, respectively. It is believed that the DCFB module may play a role. In addition, the Recall, Precision and F1-score of this network are 0.23%, 1.9% and 0.88% higher than those of the latest CT-UNet, respectively. It can be considered that multi-path encoding MEPM is more suitable for road segmentation tasks than single-path encoders. Multi-path encoding in MEPM can collect internal semantic features at each stage, while single-path encoders may ignore some spectral information. In fact, the simplest U-Net can only reach 64.28%, 77.98% and 69.62% in Recall, Precision and F1-score, respectively. MCTC-Net performed the best, and was 8.41%, 1.32%, and 5.58% higher than U-Net in these three indicators, respectively. This means that the network in this application has strong generalization ability and can handle not only simple scenarios, but also complex scenarios.
[0064] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A road segmentation method for contour enhancement through a multi-channel compact transmission network, characterized in that, the road segmentation method for contour enhancement through a multi-channel compact transmission network includes: using a publicly available dataset and preprocessing the pictures in the dataset to obtain a training dataset; training an image segmentation model constructed with the training dataset, the image segmentation model includes an encoder, a context transmission path and a decoder, the encoder includes multiple stage multi-path encoding modules connected in sequence, the context transmission path includes dense connection fusion blocks respectively connected to the multi-path encoding modules, the decoder includes noise suppression blocks respectively connected to the dense connection fusion blocks, and a feature fusion module connected to the noise suppression blocks, the output of the multi-path encoding module of the last stage of the encoder also passes through global pooling and then accesses the noise suppression block and the feature fusion module of the same stage, and the output of the feature fusion module of each stage accesses the noise suppression block and the feature fusion module of the previous stage; inputting the picture to be segmented into the trained image segmentation model and outputting a segmentation result; wherein, the dense connection fusion block includes a first fusion module FP1, a DP module and a second fusion module FP2; the first fusion module FP1 merges the information of the current stage with the information of the previous stage; the DP module includes three identical components, each component includes a convolution kernel, BatchNorm and a ReLu activation function, and short-path connections are adopted between the components; the second fusion module FP2 is used for feature output, outputting Ouput1 and Ouput2, and respectively outputting to the noise suppression block of the current stage and the dense connection fusion block of the next stage; the noise suppression block includes a spatial attention stage and a channel noise suppression stage, the spatial attention stage first performs 1*1 convolution, LeakyRelu activation function, 3*3 convolution and Sigmoid processing on the output of the context transmission path in sequence to generate a spatial attention map, multiplies and fuses the spatial attention map with the output of the noise suppression block of the previous stage, and outputs to the channel noise suppression stage; the channel noise suppression stage performs AvgPool and Sigmoid processing on the input features, then multiplies and fuses them with the input features, and then adds and fuses them with the input features as the final output.
2. The road segmentation method for contour enhancement through a multi-channel compact transmission network according to claim 1, characterized in that, the input of the multi-path encoding module passes through a first path, a second path and a third path, and after the outputs of the three paths are fused, they are output after 1×1 convolution, BatchNorm and LeakyReLu activation function. The first path is used to maintain the position invariance and rotation invariance of the road spectral features, the second path is used to prevent the network gradient from exploding or vanishing, and the third path is used to associate the spectral features with other features.
3. The road segmentation method for contour enhancement through a multi-channel compact transmission network according to claim 2, characterized in that, The input of the multi-path encoding module of the last stage of the encoder passes through the first path and the second path. The outputs of the two paths are fused and then output after passing through 1×1 convolution, BatchNorm, and LeakyReLu activation functions. The first path is used to maintain the position invariance and rotation invariance of the road spectral features, and the second path is used to prevent the network gradient from exploding or disappearing.
4. The road segmentation method for contour enhancement by multi-path compact transmission network according to claim 3, It is characterized in that The first path is PoolPath, which includes 1×1 convolution, BatchNorm, LeakyReLu activation function and Maxpool in sequence; the second path is ResPath, which includes a Resnet_18 module; the third path is ConvPath, which includes a convolution kernel and BatchNorm.
5. The road segmentation method for contour enhancement by multi-path compact transmission network according to claim 3, It is characterized in that The convolution kernel of the third path decreases with the stage.
Citation Information
Patent Citations
Medical image automatic segmentation method based on multi-path attention fusion
CN111681252A
Remote sensing image road segmentation method based on context information and attention mechanism
CN112183258A