Semantic segmentation-based road extraction method and system
By introducing a hybrid loss function and adaptive multi-scale feature fusion into the semantic segmentation model, the connectivity and integrity issues of road extraction in remote sensing images are solved, achieving higher quality road extraction results.
Patent Information
- Application Number
- CN202510833322.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-21
AI Technical Summary
Existing semantic segmentation methods based on deep learning have difficulty in effectively capturing the continuity characteristics of roads in remote sensing image road extraction, resulting in poor integrity and connectivity of the extraction results, especially in complex scenarios.
By introducing a hybrid loss function into the semantic segmentation model, combining pixel classification loss and connectivity loss, the connectivity loss is calculated using the predicted road centerline and the real road centerline. Furthermore, the separability of road features is enhanced through adaptive multi-scale feature fusion and connectivity enhancement models, thereby improving the connectivity and completeness of road extraction.
It enhances the connectivity and completeness of road extraction results, enabling it to better handle road extraction tasks in complex scenarios.
Smart Images

Figure CN120823384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a road extraction method and system based on semantic segmentation. Background Art
[0002] Remote sensing imagery, with its wide coverage and rapid update cycles, has become an important data source for road extraction. However, its complexity also presents challenges for road extraction. These challenges include inter-class similarity issues: roads and non-road features such as building roofs and farmland ridges have high similarities in spectral and textural characteristics; and morphological diversity issues: roads of different grades and types exhibit significant differences in geometric features and material properties.
[0003] Currently, deep learning-based semantic segmentation methods have become the mainstream technology for road extraction. These methods utilize convolutional neural networks and encoder-decoder architectures to achieve pixel-level road classification. However, these methods still have significant shortcomings in addressing the aforementioned challenges: relying solely on pixel-level classification loss functions makes it difficult to effectively capture the continuity characteristics of roads, affecting the completeness and connectivity of the extracted results.
[0004] This problem makes the existing methods perform poorly in terms of the completeness and connectivity of road extraction, especially when dealing with complex scenes. How to solve this problem to achieve more complete and connected road extraction is a key issue that needs to be addressed in remote sensing image road extraction systems. Summary of the Invention
[0005] In view of this, it is necessary to provide a road extraction method and system based on semantic segmentation to solve the problems of poor integrity and connectivity of road extraction in existing remote sensing images.
[0006] In order to solve the above problems, in a first aspect, the present invention provides a road extraction method based on semantic segmentation, comprising: Input the remote sensing road image to be identified into the trained semantic segmentation model to obtain the first road segmentation result; Extracting a road according to the first road segmentation result; In which, the semantic segmentation model performs reverse model parameter update according to the hybrid loss during training; the hybrid loss includes pixel classification loss and connectivity loss; the connectivity loss is determined based on the predicted road centerline and the real road centerline; the predicted road centerline is obtained by extracting the road centerline of the second road segmentation result, the second road segmentation result is obtained by semantic segmentation of the sample remote sensing road image by the semantic segmentation model, and the real road centerline is obtained by extracting the road centerline of the real road segmentation result corresponding to the second road segmentation result.
[0007] In a possible implementation, the connectivity loss is calculated based on the predicted road centerline and the actual road centerline in combination with a dice coefficient loss function.
[0008] In a possible implementation, the pixel classification loss is calculated by combining the second road segmentation result and the real road segmentation result with a binary cross entropy loss function.
[0009] In one possible implementation, the semantic segmentation model includes a feature encoding module, a feature fusion module, and a feature decoding module; during the training process, the feature encoding module is used to extract features from the sample remote sensing road image to obtain a feature map; the feature fusion module is used to perform dilated convolution on the feature map according to multiple different dilation rates to obtain multiple dilated feature maps, and perform weighted fusion on the multiple dilated feature maps according to the weights of each dilated feature map to obtain a target fused feature map; the feature decoding module is used to decode the target fused feature map to obtain the second road segmentation result.
[0010] In one possible implementation, the feature fusion module is used to add the multiple expanded feature maps to obtain a sub-fusion feature map, perform channel feature extraction on the sub-fusion feature map to obtain multiple channel features, and determine the weights of multiple channels based on the multiple channel features, and use the weights of the multiple channels as the weights of the multiple expanded feature maps; wherein the number of channels is the same as the number of expanded feature maps.
[0011] In one possible implementation, the number of channels is 3; the feature fusion module is used to perform global maximum pooling and global average pooling operations in the spatial dimension on the sub-fusion feature map to obtain two feature maps of length; the two feature maps of length are input into a multilayer perceptron to obtain two feature maps of length; the two feature maps of length are added element by element and decomposed into 3 feature maps, and the Softmax operation is performed on the 3 feature maps respectively to obtain the weight of each channel.
[0012] In a possible implementation, extracting a road according to the first road segmentation result includes: Inputting the first road segmentation result and the remote sensing road image to be identified into a trained connectivity enhancement model to obtain a third road segmentation result; Extracting a road according to the third road segmentation result; The connectivity enhancement model is obtained by training the real road segmentation results and the damaged road segmentation results.
[0013] In one possible implementation, the module structure of the connectivity enhancement model and the loss function used during training are the same as those of the semantic segmentation model.
[0014] In a possible implementation, the damaged road segmentation result is obtained by randomly removing corner points of the real road segmentation result and / or performing random road masking.
[0015] In a second aspect, the present invention further provides a remote sensing image road extraction system, comprising: A semantic segmentation module is used to input the remote sensing road image to be identified into the trained semantic segmentation model to obtain a first road segmentation result; A road extraction module, configured to extract roads based on the first road segmentation result; In which, the semantic segmentation model performs reverse model parameter update according to the hybrid loss during training; the hybrid loss includes pixel classification loss and connectivity loss; the connectivity loss is determined based on the predicted road centerline and the real road centerline; the predicted road centerline is obtained by extracting the road centerline of the second road segmentation result, the second road segmentation result is obtained by semantic segmentation of the sample remote sensing road image by the semantic segmentation model, and the real road centerline is obtained by extracting the road centerline of the real road segmentation result corresponding to the second road segmentation result.
[0016] The beneficial effects of the present invention are: Considering that road objects that are dissimilar in spectral texture features all have uniform connectivity features, while non-road objects that are similar to road objects in spectral texture do not have connectivity features. Therefore, before using a semantic segmentation model to perform semantic segmentation and road extraction on a remote sensing road image to be identified, the present invention first trains the semantic segmentation model using sample remote sensing road images, and while calculating the pixel classification loss for the second road segmentation result output by the training, also performs road centerline extraction on the second road segmentation result to obtain a predicted road centerline, calculates the connectivity loss based on the predicted road centerline and the actual road centerline, and determines the hybrid loss by combining the pixel classification loss and the connectivity loss. Then, the parameters of the semantic segmentation model are updated based on the hybrid loss, so that the trained semantic segmentation model can comprehensively consider the spectral texture features and geometric features of the road to enhance the separability of the road features and improve the connectivity and integrity of the road extraction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A schematic diagram of a flow chart of an embodiment of a road extraction method based on semantic segmentation provided by the present invention; Figure 2 A schematic diagram of the road extraction network model structure provided by the present invention; Figure 3 A schematic diagram of an adaptive multi-scale feature fusion module provided by the present invention; Figure 4 A flowchart of semantic segmentation model training provided by the present invention; Figure 5 For the present invention Figure 1 A schematic flow chart of an embodiment of S102; Figure 6 A flow chart of connectivity enhancement model training provided by the present invention; Figure 7 A system principle diagram provided by the present invention; Figure 8 A schematic diagram of a system framework provided by the present invention; Figure 9 A flow chart of a system method provided by the present invention; Figure 10 This is a structural diagram of an embodiment of the remote sensing image road extraction system provided by the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] In the description of the embodiments of the present invention, unless otherwise specified, "multiple" means two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0021] The terms "first," "second," and the like used in the embodiments of the present invention are used to distinguish similar objects, and are not used to describe a specific order or precedence, nor are they used to indicate or imply relative importance or implicitly specify the number of technical features indicated. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of a class, and do not limit the number of objects. For example, the first object can be one or more.
[0022] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0023] Reference Figure 1 , which shows a flow chart of an embodiment of a road extraction method based on semantic segmentation provided by the present invention, the method comprising: S101: Input the remote sensing road image to be identified into a trained semantic segmentation model to obtain a first road segmentation result.
[0024] The remote sensing road image to be identified may be a remote sensing road image uploaded by a user for road identification. The remote sensing road image to be identified may be uploaded in slices, and the first road segmentation result may be a road segmentation result of each slice.
[0025] The semantic segmentation model can be a model with a road extraction network model as the network architecture. Figure 2 FIG2 is a schematic diagram showing the structure of a road extraction network model provided by the present invention. The road extraction network model includes a feature encoding module, an adaptive multi-scale feature fusion module, and a feature decoding module.
[0026] The feature encoding module consists of a 7×7 convolution layer, a global maximum pooling layer and four encoding layers. After the remote sensing image (the remote sensing road image used in the model application process or the remote sensing road image used in the model training process) is input into the feature encoding module, it will perform convolution operation, pooling operation and four downsampling operations in sequence to obtain the features of the remote sensing image. Figure X .
[0027] The adaptive multi-scale feature fusion module is used to Figure X Perform multi-scale convolution expansion, channel feature extraction, and channel feature fusion based on the attention mechanism to obtain the target fusion feature map V.
[0028] The feature decoding module consists of four decoding layers, a 3×3 transposed convolution layer, a 3×3 convolution layer, a 3×3 transposed convolution layer, and a sigmoid activation function. The target fused feature map V is sequentially upsampled four times through the four decoding layers. It then undergoes one convolution, two transposed convolutions, and an activation operation to produce the road prediction result (either the first road extraction result or the second road extraction result).
[0029] In this embodiment, the semantic segmentation model is trained on sample remote sensing road images before use. During training, the model parameters are updated inversely based on a hybrid loss until the hybrid loss meets the required training requirements. The hybrid loss includes pixel classification loss and connectivity loss; the connectivity loss is determined based on the predicted road centerline and the actual road centerline; the predicted road centerline is obtained by extracting the road centerline from the second road segmentation result (the second road segmentation result is obtained by semantic segmentation of the sample remote sensing road image using the semantic segmentation model), and the actual road centerline is obtained by extracting the road centerline from the actual road segmentation result corresponding to the second road segmentation result.
[0030] The road centerline extraction process can be as follows: using the Zhang-Suen image skeleton extraction algorithm to iteratively delete the edge pixels of the road segmentation result, and then converting the connected areas in the binary image into a single-pixel wide skeleton to obtain the road centerline.
[0031] S102: Extract roads based on the first road segmentation result.
[0032] The first road segmentation result may be a pixel classification result in the remote sensing road image to be identified. According to the pixel classification result, the road in the remote sensing road image to be identified may be determined.
[0033] The semantic segmentation-based road extraction method provided in this embodiment can be applied to a semantic segmentation-based road extraction software system. This semantic segmentation-based road extraction software system can be a software system running on a terminal device. The terminal device can be a tablet computer, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a mobile phone, or other terminal device. This embodiment does not impose any restrictions on the specific type of terminal device.
[0034] In summary, before performing semantic segmentation and road extraction on the remote sensing road image to be identified through the semantic segmentation model, this embodiment first trains the semantic segmentation model through sample remote sensing road images, and calculates the pixel classification loss for the second road segmentation result output by the training. At the same time, the road centerline is extracted from the second road segmentation result to obtain a predicted road centerline, and the connectivity loss is calculated based on the predicted road centerline and the actual road centerline. The mixed loss is determined by combining the pixel classification loss and the connectivity loss, and the parameters of the semantic segmentation model are updated based on the mixed loss, so that the trained semantic segmentation model can comprehensively consider the spectral texture characteristics and geometric characteristics of the road to enhance the separability of the road features and improve the connectivity and integrity of the road extraction results.
[0035] In some embodiments of the present invention, the pixel classification loss can be calculated by a binary cross entropy loss function, and the calculation process is as follows:
[0036] Where: Indicates the number of pixels involved in the calculation, represents the true label of the pixel (the true road segmentation result), Represents the classification probability value of the predicted pixel.
[0037] The connectivity loss can be calculated using the dice coefficient loss function. The calculation process is as follows: First, the prediction results and the true label Skeletonize the predicted road centerline With the real road centerline , then use and Calculating topological accuracy ,use and Calculating topological sensitivity :
[0038]
[0039] To maximize precision and sensitivity, the harmonic mean of these two metrics is calculated:
[0040] Then, calculate the ratio of the intersection area between the prediction result and the true label to the total area:
[0041] Finally, the dice coefficient loss function is used to calculate the connectivity loss of the road prediction results:
[0042] The hybrid loss can be calculated by combining the above-mentioned binary cross entropy loss function and the centerline dice coefficient loss function. The calculation process is as follows:
[0043] In summary, this embodiment combines the binary cross entropy loss function and the dice coefficient loss function to design a hybrid loss function. The binary cross entropy loss is used to calculate the difference between the predicted value and the true value in the pixel binary classification problem, so as to optimize the parameters of the network model to achieve complete road segmentation; the dice coefficient loss is used to calculate the difference in the road topology structure between the predicted result and the true result, thereby improving the sensitivity of the network model to road breakage and enhancing the connectivity of the road segmentation results.
[0044] Traditional road extraction methods based on semantic segmentation models involve a relatively simple multi-scale feature fusion method, which usually fuses features by direct addition. This fusion method cannot effectively distinguish the importance of features at different scales.
[0045] In view of this, in some embodiments of the present invention, during the training process, the feature fusion module in the semantic segmentation model is used to perform dilated convolution on the feature map output by the feature encoding module according to multiple different dilation rates to obtain multiple dilated feature maps, and perform weighted fusion on the multiple dilated feature maps according to the weights of each dilated feature map to obtain a target fused feature map; the feature decoding module in the semantic segmentation model is used to decode the target fused feature map to obtain a second road segmentation result.
[0046] This embodiment changes the receptive field of the convolution kernel by using dilated convolutions with different dilation rates, and then performs weighted fusion on the dilated feature maps obtained by the dilated convolution, so that the semantic segmentation model can effectively distinguish the importance of features at different scales and capture feature information at different scales.
[0047] In some embodiments of the present invention, during the training process, the feature fusion module is used to add multiple expanded feature maps to obtain a sub-fusion feature map, and perform channel feature extraction on the sub-fusion feature map to obtain multiple channel features, and determine the weights of multiple channels based on the multiple channel features, and use the weights of the multiple channels as the weights of multiple expanded feature maps; wherein the number of channels is the same as the number of expanded feature maps.
[0048] In some embodiments of the present invention, the number of channels is 3; the feature fusion module is used to perform a global maximum pooling operation and a global average pooling operation on the sub-fusion feature map in the spatial dimension to obtain two sub-fusion feature maps of length The feature map of The feature map is input into the multi-layer perceptron to obtain two features of length The feature map of The feature map is decomposed into three elements after element-by-element addition , and respectively for 3 The Softmax operation is performed on the feature map to obtain the weight of each channel.
[0049] Reference Figure 3 , which shows a schematic diagram of an adaptive multi-scale feature fusion provided by the present invention. Feature fusion includes three stages: splitting stage, fusion stage and selection stage.
[0050] In the splitting stage, three dilated convolutions with a convolution kernel size of 3×3 and dilation rates of 1, 2, and 4 are used to perform the dilation on the input feature maps. Perform convolution operation to split into 3 dilated feature maps that can extract feature information of different scales 、 、 , we can obtain feature information with receptive fields of 3×3, 7×7, and 15×15 respectively; In the fusion stage, the dilated feature map 、 、 Directly add to get the sub-fusion feature map :
[0051] Then, yes Perform global maximum pooling and global average pooling in the spatial dimension to obtain two lengths of The two feature maps are fed into a shared multilayer perceptron (MLP) to learn the importance of each channel, and two feature maps of length Finally, the two results are added element by element to obtain the global information :
[0052] Among them, the number of neurons in the first layer of MLP is , the number of neurons in the second layer is .
[0053] In the selection phase, the vector Depend on The feature map becomes The feature map, that is, 3 Feature map 、 、 Next, perform Softmax operation on each column to obtain the channel feature weight 、 、 :
[0054]
[0055]
[0056] Then, separate them with 、 、 Multiply and add the results to get the result of adaptive multi-scale feature information fusion :
[0057] In some embodiments of the present invention, the sample remote sensing road images used in the semantic segmentation model training process can be obtained by: Obtain a publicly available remote sensing road image dataset, perform data augmentation on the remote sensing images in the dataset, and use the resulting remote sensing images as sample remote sensing road images. Data augmentation can include the following three aspects: First, perform color dithering on the input remote sensing image, randomly adjusting the image's hue, saturation, and brightness with a certain probability to simulate image characteristics under different lighting conditions and increase data diversity; Second, perform geometric transformations on the input remote sensing image and mask, randomly shifting, scaling, rotating, and performing affine transformations with a certain probability to change the image's spatial characteristics and enable the model to adapt to images of different perspectives and scales; Third, perform flipping on the input remote sensing image and mask, horizontally flipping, vertically flipping, and rotating them 90 degrees with a certain probability to enable the model to adapt to images of different orientations.
[0058] Reference Figure 4 , shows a flow chart of semantic segmentation model training provided by the present invention. The semantic segmentation model is trained by remote sensing images after data enhancement in the public dataset. During the training process, the input feature encoding module performs feature encoding on the remote sensing image to obtain a feature map. X , through the adaptive multi-scale feature fusion module to fusion the feature map X Perform multi-scale feature fusion, and finally calculate the hybrid loss, and perform reverse parameter adjustment on the semantic segmentation model based on the hybrid loss.
[0059] In some embodiments of the present invention, the use process of the semantic segmentation model is similar to the training process and will not be described in detail here.
[0060] Traditional road extraction methods also face the following challenges: Due to ground occlusion, roads are often blocked by trees, buildings, etc., resulting in incomplete road information in remote sensing images. Although traditional methods use global context modeling mechanisms to infer occluded roads, this method often ignores local details, resulting in poor small-scale road recognition.
[0061] In view of this, in some embodiments of the present invention, as Figure 5 As shown, S102 includes: S501: Input the first road segmentation result and the remote sensing road image to be identified into a trained connectivity enhancement model to obtain a third road segmentation result.
[0062] The remote sensing road image to be identified can also be sliced and input into the trained connectivity enhancement model.
[0063] The connectivity enhancement model is trained using real road segmentation results and damaged road segmentation results. The damaged road segmentation results can be obtained by randomly removing corner points and / or performing random road masking on the real damaged road segmentation results.
[0064] Reference Figure 6 , shows a flow chart of a connectivity enhancement model training provided by the present invention. The module structure of the connectivity enhancement model and the loss function used in training can be the same as those of the semantic segmentation model.
[0065] S502: Extract roads based on the third road segmentation result.
[0066] This embodiment uses a semantic segmentation model to classify each pixel in the remote sensing road image to obtain a preliminary road segmentation result. The connectivity enhancement model can further refine the preliminary road segmentation result, repair and connect broken road sections, and thus enhance road connectivity.
[0067] In some embodiments of the present invention, the damaged road segmentation results used in the connectivity enhancement model training process can be obtained by: A public remote sensing road image dataset was obtained and corner extraction was performed on the road centerlines in the remote sensing images. The corner extraction operation is as follows: First, the Sobel operator is used to calculate the horizontal and vertical gradients of each pixel in the road centerline image. Then, the autocorrelation matrix within the local window of each pixel is calculated. Next, the two eigenvalues of this matrix are calculated to reflect the gradient strength in the two main directions within the local window. The smaller eigenvalue is used as the corner response value. Finally, a threshold is determined. If the response value exceeds the set threshold, the pixel is determined to be a corner point.
[0068] Then, these corner points are randomly removed, and the pixel points representing the road area in the real road segmentation result are randomly selected as the center point, and a mask is randomly generated based on the center point to cover the road pixels.
[0069] Reference Figure 7 , showing a schematic diagram of a system provided by the present invention. A user uploads a remote sensing image to the road extraction system, which slices the uploaded remote sensing image and uniformly crops it into 512×512 image slices with a 10% overlap. The remote sensing image road extraction system uses a trained semantic segmentation model to perform semantic segmentation on the remote sensing image slices, obtaining road segmentation results for each slice. The trained connectivity enhancement model is then used to perform connectivity enhancement on the road segmentation results for each slice based on the user-uploaded remote sensing image, obtaining connectivity-enhanced road segmentation results for each slice. Finally, the connectivity-enhanced road segmentation results are spliced and restored to the same size as the original remote sensing image, completing road extraction from the remote sensing image.
[0070] Reference Figure 8 , which shows a schematic diagram of a system framework provided by the present invention. The remote sensing image road extraction system includes a data preprocessing module, a model training module, a semantic segmentation module, and a connectivity enhancement module.
[0071] Reference Figure 9 , which shows a flow chart of a system method provided by the present invention. The user uploads remote sensing images to the road extraction system, which extracts roads from the remote sensing images using a semantic segmentation model and a connectivity enhancement model.
[0072] Reference Figure 10 , which shows a schematic structural diagram of an embodiment of a remote sensing image road extraction system provided by the present invention, wherein the system 10 includes: The semantic segmentation module 110 is used to input the remote sensing road image to be identified into the trained semantic segmentation model to obtain a first road segmentation result; A road extraction module 120, configured to extract roads according to the first road segmentation result; In which, the semantic segmentation model performs reverse model parameter update according to the hybrid loss during training; the hybrid loss includes pixel classification loss and connectivity loss; the connectivity loss is determined based on the predicted road centerline and the real road centerline; the predicted road centerline is obtained by extracting the road centerline of the second road segmentation result, the second road segmentation result is obtained by semantic segmentation of the sample remote sensing road image by the semantic segmentation model, and the real road centerline is obtained by extracting the road centerline of the real road segmentation result corresponding to the second road segmentation result.
[0073] It should be noted that the implementation principles or implementation processes of the above modules can refer to the above-mentioned embodiment of the road extraction method based on semantic segmentation, and will not be described one by one here.
[0074] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0075] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.
Claims
1. A road extraction method based on semantic segmentation, characterized in that: include: Input the remote sensing road image to be identified into the trained semantic segmentation model to obtain the first road segmentation result; Extracting a road according to the first road segmentation result; In which, the semantic segmentation model performs reverse model parameter update according to the hybrid loss during training; the hybrid loss includes pixel classification loss and connectivity loss; the connectivity loss is determined based on the predicted road centerline and the real road centerline; the predicted road centerline is obtained by extracting the road centerline of the second road segmentation result, the second road segmentation result is obtained by semantic segmentation of the sample remote sensing road image by the semantic segmentation model, and the real road centerline is obtained by extracting the road centerline of the real road segmentation result corresponding to the second road segmentation result.
2. The road extraction method based on semantic segmentation according to claim 1, characterized in that: The connectivity loss is calculated based on the predicted road centerline and the actual road centerline in combination with the dice coefficient loss function.
3. The road extraction method based on semantic segmentation according to claim 1, characterized in that: The pixel classification loss is calculated by combining the second road segmentation result and the real road segmentation result with a binary classification cross entropy loss function.
4. The road extraction method based on semantic segmentation according to claim 1, characterized in that: The semantic segmentation model includes a feature encoding module, a feature fusion module and a feature decoding module; during the training process, the feature encoding module is used to extract features from the sample remote sensing road image to obtain a feature map; the feature fusion module is used to perform dilated convolution on the feature map according to multiple different dilation rates to obtain multiple dilated feature maps, and perform weighted fusion on the multiple dilated feature maps according to the weights of each dilated feature map to obtain a target fused feature map; the feature decoding module is used to decode the target fused feature map to obtain the second road segmentation result.
5. The road extraction method based on semantic segmentation according to claim 4, characterized in that: The feature fusion module is used to add the multiple expanded feature maps to obtain a sub-fusion feature map, and perform channel feature extraction on the sub-fusion feature map to obtain multiple channel features, and determine the weights of multiple channels based on the multiple channel features, and use the weights of the multiple channels as the weights of the multiple expanded feature maps; wherein the number of channels is the same as the number of expanded feature maps.
6. The road extraction method based on semantic segmentation according to claim 5, characterized in that: The number of channels is 3; the feature fusion module is used to perform global maximum pooling and global average pooling operations on the sub-fusion feature map in the spatial dimension to obtain two sub-fusion feature maps of length The feature map of the two lengths is The feature map is input into the multi-layer perceptron to obtain two features of length The feature map of The feature map is decomposed into three elements after element-by-element addition , and respectively for the three The Softmax operation is performed on the feature map to obtain the weight of each channel; among them, is the number of channels.
7. The road extraction method based on semantic segmentation according to claim 1, characterized in that: The extracting a road according to the first road segmentation result includes: Inputting the first road segmentation result and the remote sensing road image to be identified into a trained connectivity enhancement model to obtain a third road segmentation result; Extracting a road according to the third road segmentation result; The connectivity enhancement model is obtained by training the real road segmentation results and the damaged road segmentation results.
8. The road extraction method based on semantic segmentation according to claim 7, characterized in that: The module structure of the connectivity enhancement model and the loss function used during training are the same as those of the semantic segmentation model.
9. The road extraction method based on semantic segmentation according to claim 7, characterized in that: The damaged road segmentation result is obtained by randomly removing corner points of the real road segmentation result and / or performing random road masking.
10. A remote sensing image road extraction system, characterized in that: include: A semantic segmentation module is used to input the remote sensing road image to be identified into the trained semantic segmentation model to obtain a first road segmentation result; A road extraction module, configured to extract roads based on the first road segmentation result; In which, the semantic segmentation model performs reverse model parameter update according to the hybrid loss during training; the hybrid loss includes pixel classification loss and connectivity loss; the connectivity loss is determined based on the predicted road centerline and the real road centerline; the predicted road centerline is obtained by extracting the road centerline of the second road segmentation result, the second road segmentation result is obtained by semantic segmentation of the sample remote sensing road image by the semantic segmentation model, and the real road centerline is obtained by extracting the road centerline of the real road segmentation result corresponding to the second road segmentation result.
Citation Information
Cited By
Road geometric physical supervision training method, device and equipment based on binary mask
CN121544659A
Road geometry physics supervision training method, device and equipment based on binary mask
CN121544659B