Road extraction method based on pooling enhancement
By introducing the MSASPP module and post-processing algorithm in the SegFormer network, the problem of insufficient feature expression and discontinuity of extraction during road extraction in the prior art is solved, and higher road extraction accuracy and connectivity are achieved.
Patent Information
- Application Number
- CN202510016630.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-03
AI Technical Summary
When the existing road extraction methods deal with complex and multi-scale linear road features, there are problems of insufficient feature expression and discontinuous extraction, resulting in unsatisfactory recognition accuracy.
Based on the SegFormer network, a multi-scale stripe space hollow pooling module (MSASPP) was introduced to enhance the feature map through pooling, improve the model's extraction ability of multi-scale line-shaped targets, and combine post-processing algorithms to optimize the connectivity and integrity of the road.
It effectively improves the accuracy and connectivity of road extraction, enhances the model's processing ability of complex linear structures, and improves the robustness and reliability of road extraction tasks.
Smart Images

Figure CN120088641A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of road extraction, and more specifically, to a road extraction method based on pooling enhancement. Background Art
[0002] The road extraction technology in remote sensing images has important application values in urban planning, traffic monitoring, disaster emergency response, etc. However, due to the complex road morphology, especially in high-resolution remote sensing images, linear features are easily interfered by noise, occlusion, and other complex backgrounds, making it difficult for existing road extraction methods to achieve satisfactory results. These roads have significant linear features, usually presenting as horizontal or vertical straight lines or curves. The existing remote sensing image segmentation methods have limited effects in extracting these linear targets. Especially when facing complex and multi-scale linear targets, traditional segmentation networks are prone to ignoring the linear features of the targets, resulting in unsatisfactory segmentation results.
[0003] Currently, the mainstream road extraction methods are mostly based on deep learning semantic segmentation models, such as U-Net, DeepLab, etc. However, although these models can extract road information to a certain extent, they often cannot handle slender, tortuous, and multi-scale road features well, resulting in less than ideal road recognition accuracy.
[0004] The SegFormer network combines the advantages of Transformer and traditional convolutional neural network (CNN), and can provide accurate segmentation results without using complex post-processing. As a lightweight semantic segmentation model with a Transformer architecture, it uses a hierarchical Transformer encoder for feature extraction and combines a multi-layer perceptron (MLP) for segmentation prediction, having good feature extraction capabilities. However, in road extraction tasks, especially when dealing with the complex linear structure of roads, there are still problems such as insufficient feature expression and discontinuous extraction. Therefore, how to further improve the extraction accuracy of road features on the basis of the SegFormer model, especially to effectively optimize the connectivity and integrity of roads, has become the focus of technological innovation. Summary of the Invention
[0005] The present invention aims at the technical problems existing in the prior art, and provides a road extraction method based on pooling enhancement, which can overcome the problems of insufficient feature expression and discontinuous extraction that still exist in existing road extraction tasks, especially when dealing with the complex linear structure of roads.
[0006] According to a first aspect of the present invention, there is provided a road extraction method based on pooling enhancement, including:
[0007] Converting the bit depth of the remote sensing road image into the bit depth of the deep learning semantic segmentation model;
[0008] Input the remotely sensed road image after bit-depth conversion into the deep learning semantic segmentation model to obtain the road segmentation result of the remotely sensed road image; where:
[0009] The deep learning semantic segmentation model is a SegFormer network, and the SegFormer network includes an encoder, multiple Atrous Spatial Pyramid Pooling (MSASPP) modules, and a decoder;
[0010] Extract feature maps of different scales of the input remotely sensed road image through multiple feature extraction modules in the encoder;
[0011] Perform pooling enhancement on the feature maps of each scale through each MSASPP module, and the number of MSASPP modules is the same as the number of feature extraction modules;
[0012] Fuse the feature maps of different scales after pooling enhancement through the decoder to obtain a fused feature map, and segment the fused feature map to obtain the road segmentation result.
[0013] A road extraction method based on pooling enhancement provided by the present invention effectively enhances the model's ability to extract multi-scale linear targets in remotely sensed images, especially the accurate segmentation of roads, by introducing a Multi-Scale Atrous Strip Pooling (MSASPP) module into the skip structure between the encoder and the decoder. Description of the Drawings
[0014] Figure 1 It is a flowchart of a road extraction method based on pooling enhancement provided by the present invention;
[0015] Figure 2 It is a schematic structural diagram of the SegFormer network;
[0016] Figure 3 It is a schematic structural diagram of the MSASPP module;
[0017] Figure 4 It is a schematic working diagram of the MSP;
[0018] Figure 5 It is a schematic diagram of stripe pooling of an MSP module for a feature sub-map;
[0019] Figure 6 It is a schematic diagram of the road segmentation result extracted from the remotely sensed image. Detailed Embodiments
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or individual embodiment provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the order of steps and / or the structural composition mode, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions is contradictory or cannot be achieved, it should be considered that such combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0021] Figure 1 The following is a flowchart of a road extraction method based on pooling enhancement provided by the present invention. As Figure 1 shown, the method includes:
[0022] Step 1, convert the bit depth of the remote sensing road image to the bit depth of the deep learning semantic segmentation model.
[0023] It can be understood that the bit depth of the remote sensing image is usually 16 bits, and it needs to be converted to 8 bits, which is suitable for the deep learning semantic segmentation model. The present invention uses the method of percentage truncation stretching to convert the bit depth of the remote sensing image.
[0024] In a possible embodiment of the present invention, the method of converting the 16-bit bit depth of the remote sensing road image to 8-bit bit depth based on the percentage truncation stretching method includes:
[0025] Sort the pixel values of each band in the remote sensing road image according to their magnitudes, and determine the low threshold and high threshold of the pixel values based on the sorting information.
[0026] Based on the low threshold and high threshold of the pixel values, perform truncation processing on the pixel points in the remote sensing road image to obtain the remote sensing road image after truncation processing;
[0027] Stretch the truncated remote sensing road image to the target dynamic range.
[0028] Specifically, the basic principle of percentage truncation stretching is to truncate the extreme pixel values in the image (i.e., a certain percentage of the lowest and highest pixel values) and stretch or adjust the remaining pixel values. The general steps include:
[0029] (1) Determine the percentage range: After sorting the pixel values of each band in the image by size, the pixel value at the α% position is set as the low threshold Plow, and the pixel value at the (100% - β%) position is the high threshold Phigh, so that extreme dark or bright invalid pixels can be ignored.
[0030] P low = Percentile(I, α)
[0031] P high = Percentile(I, 100 - β)
[0032] Among them, Percentile(I, q) represents the q% speed limit value in the image I.
[0033] (2) Truncate pixels: Set the pixel values in the image I that are less than P low to P low , and set the pixel values greater than P high to P high to eliminate the influence of extreme values on the image. Define the processed image as I clipped :
[0034]
[0035] (3) Stretch to the full dynamic range: Stretch the truncated image I clipped to the target dynamic range [0, 255] to enhance the contrast. Perform linear stretching on each pixel value I clipped (x, y):
[0036]
[0037] Through these formulas, the entire percentage truncation and stretching process truncates from the extreme values of the image to the stretching mapping, enhancing the light and dark contrast of the image and making the visual effect clearer. When it is 0 and 255, it is the pixel range of an 8-bit depth image.
[0038] Step 2, input the remotely sensed road image after bit depth conversion into the deep learning semantic segmentation model to obtain the road segmentation result of the remotely sensed road image.
[0039] It should be understood that the remotely sensed road image after bit depth conversion is input into the deep learning semantic segmentation model to obtain the road extraction result.
[0040] Among them, the deep learning semantic segmentation model in the present invention is the SegFormer network, and the SegFormer network includes an encoder, multiple Atrous Spatial Pyramid Pooling (ASPP) modules, and a decoder.
[0041] Extract feature maps of different scales of the input remote sensing road image through multiple feature extraction modules in the encoder; perform pooling enhancement on the feature maps of each scale through each MSASPP module, where the number of MSASPP modules is the same as the number of feature extraction modules; and fuse the feature maps of different scales after pooling enhancement through the decoder to obtain a fused feature map, and segment the fused feature map to obtain the road segmentation result.
[0042] It can be understood that the SegFormer network is a Transformer-based image segmentation network with strong global feature capture ability. However, for the large number of linear objects (such as roads) in remote sensing images, traditional convolutional or standard pooling operations are difficult to capture their unique geometric features. Therefore, a multi-scale stripe pooling (MSP) module and an improved atrous spatial pyramid pooling (MSASPP) module are introduced into the SegFormer network. At the same time, in order to make more full use of the feature information output by the encoders of each scale, referring to the decoder of the UNet network, the feature maps are fused and upsampled step by step from small to large, reducing the loss of feature information caused by direct upsampling and fusing the feature information of different scales, so as to strengthen the propagation and fusion of features.
[0043] See Figure 2 , which is a schematic structural diagram of the improved SegFormer network. The remote sensing image after bit-depth conversion is input into the input layer of the encoder, and feature maps of different scales of the remote sensing image are extracted through multiple feature extraction modules. As Figure 2 shown, the encoder includes four feature extraction modules, which respectively extract feature maps of four different scales in the remote sensing image. Then, the feature maps of the four scales extracted are enhanced through the corresponding MSASPP modules respectively. The decoder fuses the feature maps of the four scales after enhancement, and extracts the road segmentation result according to the fused feature map.
[0044] Among them, inspired by the traditional atrous spatial pyramid pooling (ASPP), the present invention introduces MSP into the ASPP module to form multi-scale atrous stripe pooling ASPP (Multi-ScaleAtrousStrip Pooling, MSASPP).
[0045] The design goal of the MSASPP module is to improve the capture ability of road features of different scales. The specific approach is to replace the global average pooling in the original ASPP module with three MSP modules of different scales, and the number of divided parts of the feature maps of the three MSPs are [4, 2, 1] respectively.
[0046] This structure allows the model to capture linear features at multiple scales and fuse them, which helps the model understand the shape of the road from different scales and improve the segmentation accuracy of the road.
[0047] See Figure 3 , which is a schematic diagram of the structure of the multi-scale atrous spatial pyramid pooling module MSASPP with multi-scale stripe pooling. Each MSASPP module includes three MSP modules, denoted as the first MSP module, the second MSP module, and the third MSP module. The working principle of each MSASPP module for pooling and enhancing the feature map at each scale is as follows:
[0048] The feature maps output from the feature extraction module are respectively input into the three MSP modules. The three MSP modules respectively divide the feature maps into multiple sub-feature maps, perform horizontal and vertical stripe pooling on each sub-feature map, and splice the pooled sub-feature maps to capture linear features at different scales.
[0049] In the present invention, the first MSP module divides the feature map into four feature sub-maps, the second MSP module divides the feature map into two feature sub-maps, and the third MSP module does not divide the feature map.
[0050] In a possible implementation manner of the present invention, the first MSP module divides the feature map into a first number of first feature sub-maps in the length and width directions, performs horizontal and vertical stripe pooling on each first feature sub-map respectively, obtains each first feature sub-map after stripe pooling, and splices all the first feature sub-maps after stripe pooling at the original positions to obtain a fused linear feature map of different sizes.
[0051] The second MSP module divides the feature map into a second number of second feature sub-maps in the length and width directions, performs horizontal and vertical stripe pooling on each second feature sub-map respectively, obtains each second feature sub-map after stripe pooling, and splices all the second feature sub-maps after stripe pooling at the original positions to obtain a fused linear feature map of different sizes.
[0052] The third MSP module divides the feature map into a third number of third feature sub-maps in the length and width directions, performs horizontal and vertical stripe pooling on each third feature sub-map respectively, obtains each third feature sub-map after stripe pooling, and splices all the third feature sub-maps after stripe pooling at the original positions to obtain a fused linear feature map of different sizes.
[0053] Each MSASPP module fuses the linearly feature maps pooled by three MSP modules to obtain the pooled enhanced feature maps at the corresponding scales. Feature enhancement is performed on the feature maps at four scales through four MSASPP modules to obtain four feature-enhanced feature maps respectively.
[0054] See Figure 4 , which is a schematic diagram of the working principle of an MSP module. First, the input feature map (trans block) is divided into multiple regions in the length and width directions respectively (for example, divided into 2 blocks, 4 blocks, etc., Figure 4 divided into four blocks in ), and each block of feature sub-map is subjected to horizontal and vertical stripe pooling respectively. The feature maps generated during the stripe pooling process are feature maps of 1×W and H×1 respectively, retaining the long-distance stripe features. Finally, the pooled feature maps are stitched together at the original positions to obtain the fused linear feature information of different sizes.
[0055] See Figure 5 , which is a schematic diagram of the processing of each feature sub-map by an MSP module. The MSP module performs horizontal stripe pooling and vertical stripe pooling on the feature sub-maps respectively, and then performs linear transformation and dilation after processing to obtain the processed feature sub-maps.
[0056] Compared with the original stripe pooling, this MSP module can flexibly adjust the scale of stripe pooling according to various scale changes of the feature map, adapt to multi-scale linear targets in remote sensing images, and thus improve the extraction ability of roads at different scales.
[0057] Among them, the decoder includes multiple multi-layer perceptron MLP Layers, a fusion layer and a semantic segmentation model. The decoder fuses the feature maps at different scales after pooling enhancement to obtain a fused feature map, and segments the fused feature map to obtain the road segmentation result, including:
[0058] The feature maps at different scales output by the multiple MSASPP modules are fused and upsampled step by step in ascending order of scale through multiple multi-layer perceptron MLP Layers and a fusion layer to obtain the feature map after fusing different scales;
[0059] The road segmentation is performed on the feature map after fusing different scales through a semantic segmentation module to obtain a road segmentation image.
[0060] Among them, after extracting the road segmentation result from the remote sensing image through the improved SegFormer network, the road segmentation result is a binary segmentation map of the road. There are cases where the road is broken or disconnected in the segmentation result. Therefore, post-processing is performed on the segmentation result.
[0061] In a possible implementation manner of the present invention, after the feature map after fusing different scales is segmented by the semantic segmentation module to obtain the road segmentation result, it further includes:
[0062] Detecting multiple line segments from the road segmentation image based on the Hough line transform;
[0063] Calculating the angle information and distance information between two adjacent line segments, and connecting the two adjacent line segments whose angle information and distance information are respectively within the set angle threshold and distance threshold ranges to obtain a complete road structure.
[0064] It can be understood that after the road segmentation is completed by the SegFormer network, the present invention also introduces a post-processing algorithm based on direction and distance constraints to solve the problem of disconnection of roads in the segmentation result. This algorithm is mainly implemented through the following steps:
[0065] ① Detecting line segments from the segmented image obtained by inference using the Hough line transform.
[0066] ② According to the starting point, ending point and angle information of the line segments, applying the set distance and angle thresholds to connect adjacent line segments, that is, when the included angle and distance between two adjacent line segments are both within the set range, the two adjacent line segments are connected. Finally, a complete road structure is formed.
[0067] This post-processing method can effectively make up for the situation of road breakage or disconnection in the segmentation result, and improve the connectivity and integrity of the final extraction result.
[0068] Among them, for the training process of the deep learning semantic segmentation model, in the model training stage, the improved multi-scale stripe pooling atrous spatial pyramid pooling (MSASPP) module is integrated into the SegFormer network model to improve the feature extraction ability for linear road targets of different scales. The training adopts the standard data set reading method of the original SegFormer network model, and sets hyperparameters such as the optimizer, initial learning rate, weight decay rate, and warm-up training times to ensure the stability and efficiency of the learning process.
[0069] In addition, to avoid overfitting, an early stopping strategy is adopted to monitor the performance of the validation set. If the model performance does not improve within a certain number of rounds, the training is stopped in advance to obtain the optimal model. Introducing the early stopping strategy in the model training not only accelerates the training convergence, but also effectively prevents the overfitting phenomenon caused by too long training, thereby further improving the generalization ability of the model in the road extraction task.
[0070] After the model training is completed, we use the trained weights to perform inference evaluation on the test set of the remote sensing road data set to verify the effectiveness of the newly added module. During the evaluation process, the following evaluation indicators are adopted:
[0071] (1) Accuracy: It is used to measure the classification accuracy of the model for target road pixels and reflects the performance of the model in the global range.
[0072] (2) Loss: By calculating the cross-entropy loss, it evaluates the error size of the model for road segmentation. A lower loss value indicates that the model can predict the road area more accurately.
[0073] (3) Mean Intersection over Union (MIOU): MIOU is an important evaluation criterion for the segmentation model. It calculates the ratio of the intersection and union of the model prediction result and the true annotation. This index can comprehensively reflect the segmentation effect of the model on various road targets and is especially suitable for evaluating the accuracy and consistency of road extraction.
[0074] To further verify the road extraction effect of the model in the actual remote sensing scenario, the present invention introduces manual interpretation and evaluation. The specific steps are as follows:
[0075] (1) Comparative analysis: Compare the inference result map generated by the model with the test image one by one, and manually observe the matching degree of the roads extracted by the model and the roads in the real image to evaluate the recognition accuracy of the model for roads under complex backgrounds (such as buildings, vegetation, etc.).
[0076] (2) Effect evaluation: Based on the manual comparison results, objectively evaluate the continuity, integrity and boundary clarity of the roads extracted by the model, and especially pay attention to whether the performance of the model on details and small-scale roads meets the actual requirements.
[0077] (3) Problem summary: On the basis of comparison and evaluation, summarize the problems existing in the model, such as the phenomenon of some roads being broken, mis-extracted or missed. Analyze the possible reasons for the problems and provide directional guidance for model improvement.
[0078] Through the analysis of the above indicators, comprehensively evaluate the improvement effect of the newly added modules (including multi-scale stripe pooling MSP and MSASPP) on the model performance.
[0079] The following uses a specific embodiment to illustrate the road extraction method based on pooling enhancement provided by the present invention, which mainly includes the following steps:
[0080] 1. Image preprocessing
[0081] Select the GF-2 remote sensing image of a certain city in the first half of 2022. After preprocessing such as radiometric correction, orthorectification and fusion, a cloud-free synthetic image with a spatial resolution of 0.8 meters is obtained, and the bands are the blue band, the green band and the red band. The data type is set to 8 bits through percentage truncation stretching.
[0082] For the remote sensing image, through pre - processing such as radiometric correction, orthorectification, and image fusion, a cloud - free composite image with a spatial resolution of 0.8 meters was generated, which contains three bands: blue, green, and red. To adapt to deep - learning processing, the image data was stretched to an 8 - bit data type through percentage truncation.
[0083] 2. Road dataset production
[0084] After registering and overlaying the remote sensing image with the corresponding labels, the image was cropped and segmented into blocks with a window size of 256×256 pixels at an overlap rate of 30%. Subsequently, to improve the generalization ability of the deep - learning semantic segmentation model, various data augmentation techniques such as rotation and flipping were applied to these cropped patches to generate a diverse sample dataset. This process not only enriched the sample size but also enhanced the robustness of the data, providing high - quality input data for model training.
[0085] 3. Model training
[0086] On the improved SegFormer model, 23,923 pairs of labeled remote - sensing road sample images were trained. During training, the batch size was set to 16, the initial learning rate was 6e - 5, and the AdamW optimizer was selected to improve learning stability and efficiency. Throughout the training process, the model reached a preliminary convergence after 23 epochs of iteration, and an additional 10 epochs were set as the observation period for early stopping; if the loss function did not decrease during this observation period, the training was terminated early, and the weights of the model at the 24th epoch were retained as the final result.
[0087] 4. Model inference and post - processing
[0088] Using the model weights obtained from training, the test remote - sensing image was input into the model for inference. First, an overlapping - block strategy was adopted to perform block - based inference on the test image to ensure the continuity and accuracy of the edge regions. Then, to address the problem of discontinuous roads in the preliminary segmentation results, a post - processing algorithm based on direction and distance constraints was used for correction. This algorithm connects adjacent line segments by detecting line segments in the segmentation map and based on the direction and distance of the line segments, further improving the integrity of the roads. Finally, the processed results of each block were stitched together to form a complete segmentation map and geographical coordinate information was assigned. Figure 6 It is a schematic diagram of the road segmentation result.
[0089] A road extraction method based on pooling enhancement provided by the present invention can accurately and completely extract multi-scale linear roads from satellite remote sensing images through deep learning technology. By adding a multi-scale stripe spatial hole pooling layer module (MSP) between the encoder and decoder of the SegFormer model, the model's ability to extract linear targets (such as roads) is greatly improved. This MSP module is improved based on the principle of the atrous spatial pyramid pooling module (ASPP), replacing the global average pooling in ASPP with MSP, so that it can more flexibly capture strip feature information in different directions and at multiple scales. This module is applicable to road extraction in complex scenarios, such as urban road networks and small roads in natural environments. In addition, combined with a post-processing algorithm based on direction and distance constraints, it effectively solves common problems in the road extraction process, such as discontinuous and broken roads, and further enhances the connectivity and accuracy of the final output result. This solution greatly improves the robustness and reliability of road extraction and is applicable to a wide range of remote sensing image analysis tasks.
[0090] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0091] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0092] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0093] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the function specified in one process or a plurality of processes and / or one block or a plurality of blocks in the flow. Figure 1 one process or a plurality of processes and / or Figure 1 one block or a plurality of blocks.
[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one process or a plurality of processes and / or one block or a plurality of blocks in the flow. Figure 1 one process or a plurality of processes and / or Figure 1 one block or a plurality of blocks.
[0095] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0096] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A road extraction method based on pooling enhancement, characterized in that: include: Convert the bit depth of remote sensing road images to the bit depth of deep learning semantic segmentation models; The remote sensing road image after bit depth conversion is input into the deep learning semantic segmentation model to obtain the road segmentation result of the remote sensing road image; wherein: The deep learning semantic segmentation model is a SegFormer network, which includes an encoder, multiple dilated spatial pyramid pooling MSASPP modules and a decoder; Extracting feature maps of different scales of the input remote sensing road image through multiple feature extraction modules in the encoder; Pooling enhancement is performed on the feature map of each scale through each of the MSASPP modules, and the number of the MSASPP modules is the same as the number of feature extraction modules; The decoder is used to fuse the feature maps of different scales after pooling enhancement to obtain a fused feature map, and the fused feature map is segmented to obtain a road segmentation result.
2. The road extraction method based on pooling enhancement according to claim 1 is characterized in that: The step of converting the bit depth of the remote sensing road image into the bit depth of the deep learning semantic segmentation model includes: The remote sensing road image with a bit depth of 16 bits is converted to a bit depth of 8 bits based on a percentage truncation and stretching method.
3. The road extraction method based on pooling enhancement according to claim 2 is characterized in that: The percentage-based truncation stretching method converts the remote sensing road image with a 16-bit bit depth into an 8-bit bit depth, including: Sorting the pixel values of each band in the remote sensing road image according to size, and determining a low threshold and a high threshold of the pixel value according to the sorting information; Based on a low threshold and a high threshold of pixel values, truncating the pixel points in the remote sensing road image to obtain a truncated remote sensing road image; The truncated remote sensing road image is stretched to a target dynamic range.
4. The road extraction method based on pooling enhancement according to claim 3 is characterized in that: The step of determining the low threshold and the high threshold of the pixel value according to the sorting information includes: The pixel value at the α%th position in the sorted remote sensing road image is set as the low threshold value Plow, and the pixel value at the (100%-β%)th position is set as the high threshold value Phigh, wherein: P low =Percentile(I,α) P high =Percentile(I,100-β) Where Percentile(I,q) represents the qth percentile speed limit value in the remote sensing road image I; The step of truncating the pixel points in the remote sensing road image based on the low threshold and the high threshold of the pixel value to obtain the truncated remote sensing road image includes: The remote sensing road image I is smaller than P low The pixel value is set to a low threshold value P low , greater than P high The pixel value is set to the high threshold P high , the remote sensing road image after truncating is: The step of stretching the truncated remote sensing road image to a target dynamic range includes: The truncated remote sensing road image I clipped Each pixel value in is linearly stretched to the target dynamic range [0,255]:
5. The road extraction method based on pooling enhancement according to claim 1 is characterized in that: Each of the MSASPP modules includes three MSP modules, which are recorded as a first MSP module, a second MSP module, and a third MSP module. The feature map of each scale is pooled and enhanced by each of the MSASPP modules, including: Input the feature graphs output from the feature extraction module into three MSP modules respectively; The feature map is divided into a first number of first feature sub-maps in the length and width directions by the first MSP module, and each of the first feature sub-maps is subjected to horizontal and vertical stripe pooling to obtain each first feature sub-map after stripe pooling, and all the first feature sub-maps after stripe pooling are spliced at the original positions to obtain fused linear feature maps of different sizes; The feature map is divided into a second number of second feature submaps in the length and width directions by the second MSP module, and each of the second feature submaps is subjected to horizontal and vertical stripe pooling to obtain each second feature submap after stripe pooling, and all the second feature submaps after stripe pooling are spliced at the original positions to obtain fused linear feature maps of different sizes; Dividing the feature map into a third number of third feature submaps in length and width directions by the third MSP module, performing horizontal and vertical stripe pooling on each of the third feature submaps to obtain each third feature submap after stripe pooling, and splicing all the third feature submaps after stripe pooling at their original positions to obtain fused linear feature maps of different sizes; Each of the MSASPP modules fuses the linear feature maps after pooling of the three MSP modules to obtain a pooled enhanced feature map of the corresponding scale.
6. The road extraction method based on pooling enhancement according to claim 5 is characterized in that: The first MSP module divides the feature graph into four feature sub-graphs, the second MSP module divides the feature graph into two feature sub-graphs, and the third MSP module does not divide the feature graph.
7. The road extraction method based on pooling enhancement according to claim 1 is characterized in that: The decoder includes multiple multi-layer perceptron MLP layers, fusion layers and semantic segmentation models. The decoder fuses the feature maps of different scales after pooling enhancement to obtain a fused feature map, and the fused feature map is segmented to obtain a road segmentation result, including: The feature maps of different scales output by the multiple MSASPP modules are sequentially fused and upsampled step by step from small to large scales through multiple multi-layer perceptron MLP layers and fusion layers to obtain feature maps after fusion of different scales; The semantic segmentation module performs road segmentation on the feature maps after fusion of different scales to obtain a road segmentation image.
8. The road extraction method based on pooling enhancement according to claim 7 is characterized in that: The performing road segmentation on the feature maps after fusion of different scales by the semantic segmentation module to obtain the road segmentation result also includes: Detecting a plurality of line segments from the road segmentation image based on Hough line transform; The angle information and the distance information between two adjacent line segments are calculated, and the two adjacent line segments whose angle information and the distance information are respectively within the set angle threshold and the distance threshold are connected to obtain a complete road structure.