A method and device for constructing a road segmentation model of remote sensing images
Through the combination of dynamic road detail matcher, cross-context adaptive attention and multi-scale global information integrator, the problem of missing road extraction caused by occlusion or density in remote sensing images is solved, and higher road extraction accuracy and continuity are achieved.
Patent Information
- Application Number
- CN202411602625.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-11-11
AI Technical Summary
The existing deep learning road extraction technology has not effectively solved the problem of missing road images due to being blocked by obstacles or too dense roads in remote sensing images.
The dynamic road detail matcher, the cross-context adaptive attention and the multi-scale global information integrator are used to build a road segmentation model through synchronous hierarchical learning and cross-level feature interaction, and use the geometric shape of the shallow feature map, the context information of the middle-level feature map and the global information of the deep feature map to improve the continuity and integrity of road extraction.
In occlusion and dense road scenarios, the accuracy and continuity of road extraction are significantly improved, and the problem of missing road parts caused by occlusion or dense roads is solved.
Smart Images

Figure CN119360348B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the technical field of remote sensing image technology, and in particular to a method and apparatus for constructing a road segmentation model for remote sensing images. Background Art
[0002] Currently, deep learning road extraction technology can automatically and accurately identify and extract road networks from large-scale remote sensing image data, greatly improving the efficiency of urban planning and traffic control, and providing sufficient data support for the construction and sustainable development of smart cities.
[0003] However, the problem that the road images extracted by the deep learning road model may be partially missing due to being blocked by obstacles or overly dense roads is a technical problem that urgently needs to be solved at present. Summary of the Invention
[0004] This application describes a method and apparatus for constructing a road segmentation model for remote sensing images, which can solve the above technical problems.
[0005] According to a first aspect, there is provided a method for constructing a road segmentation model for remote sensing images, the method comprising:
[0006] Obtain a remote sensing image training sample set, where any set of remote sensing image training samples includes a sample image with a road partially covered and its corresponding road label image;
[0007] Input the sample images of the remote sensing image training sample set into an initial road segmentation model to obtain a road segmentation prediction image of the sample images;
[0008] Based on the difference between the road label image of the sample image and the road segmentation prediction image, perform parameter iteration on the initial model to obtain a road segmentation model;
[0009] The initial road segmentation model includes a dynamic road detail matcher, a cross-context adaptive attention mechanism, and a multi-scale global information integrator, where:
[0010] The dynamic road detail matcher is configured to obtain a shallow feature map through the feature image of the sample image, where the shallow feature map contains the geometric shape of the road;
[0011] The cross-context adaptive attention mechanism is configured to obtain a middle feature map according to the feature image and the shallow feature map, where the middle feature map contains the road surrounding environment information of the blocked road;
[0012] The multi-scale global information integrator is used to obtain a deep feature map based on the feature image and the middle-level feature map. The deep feature map contains the deep semantic information and global information of the road. The deep semantic information of the road includes the mutual relationship and attributes between the occluded road and surrounding objects, and the global information includes the overall layout and road structure information of the occluded road.
[0013] According to a second aspect, there is provided an apparatus for constructing a road segmentation model of a remote sensing image, the apparatus including:
[0014] An acquisition module, configured to obtain a remote sensing image training sample set, and any set of remote sensing image training samples includes a sample image with a partially occluded road and its corresponding road label image;
[0015] An input module, configured to input the sample image of the remote sensing image training sample set into an initial road segmentation model to obtain a road segmentation prediction image of the sample image;
[0016] A training module, configured to perform parameter iteration on the initial model based on the difference between the road label image of the sample image and the road segmentation prediction image to obtain a road segmentation model;
[0017] The initial road segmentation model includes a dynamic road detail matcher, a cross-context adaptive attention mechanism, and a multi-scale global information integrator, where:
[0018] The dynamic road detail matcher is used to obtain a shallow feature map through the feature image of the sample image. The shallow feature map contains the geometric shape of the road;
[0019] The cross-context adaptive attention mechanism is used to obtain a middle-level feature map according to the feature image and the shallow feature map. The middle-level feature map contains the road surrounding environment information of the occluded road;
[0020] The multi-scale global information integrator is used to obtain a deep feature map based on the feature image and the middle-level feature map. The deep feature map contains the deep semantic information and global information of the road. The deep semantic information of the road includes the mutual relationship and attributes between the occluded road and surrounding objects, and the global information includes the overall layout and road structure information of the occluded road.
[0021] According to a third aspect, there is provided a method for road segmentation of a remote sensing image, the method including:
[0022] Obtain a target remote sensing image and a road segmentation model constructed according to the method for constructing the road segmentation model in the above technical solution. The target remote sensing image includes a target road with partial occlusion;
[0023] Input the target remote sensing image into the road segmentation model to obtain the target road image that is not covered in the target remote sensing image.
[0024] According to a fourth aspect, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the method for constructing the road segmentation model of the remote sensing image as described in the above technical solution.
[0025] According to a fifth aspect, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the road segmentation method of the remote sensing image as described in the above technical solution.
[0026] In the above systems and methods provided in the embodiments of this specification, a dynamic road detail matcher suitable for extracting a large amount of road detail information in the shallow feature map, a cross-context adaptive attention suitable for extracting sufficient context information and cross-level feature interaction, and a multi-scale global information integrator capable of efficiently enhancing the global information of the road and improving the road continuity are proposed. The present invention is applicable to the field of high-resolution remote sensing image road segmentation, and is particularly suitable for a model that solves the problem of missing road parts caused by road occlusion or dense distribution. Description of the Drawings
[0027] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a schematic diagram of the network structure of a road segmentation model of a remote sensing image provided by the present invention;
[0029] Figure 2 It is a schematic diagram of the structure of FusedMBConv in a road segmentation model of a remote sensing image provided by the present invention;
[0030] Figure 3 It is a schematic diagram of the structure of MBConv in a road segmentation model of a remote sensing image provided by the present invention;
[0031] Figure 4 It is a schematic diagram of the structure of the dynamic road detail matcher in a road segmentation model of a remote sensing image provided by the present invention;
[0032] Figure 5It is a schematic diagram of the cross-context adaptive attention in a road segmentation model for remote sensing images provided by the present invention;
[0033] Figure 6 It is a schematic diagram of the multi-scale global information integrator in a road segmentation model for remote sensing images provided by the present invention;
[0034] Figure 7 It is an example of a remote sensing image of the DeepGlobe dataset used in an embodiment of the present invention;
[0035] Figure 8 It is an example of a remote sensing image of the label in the DeepGlobe dataset used in an embodiment of the present invention;
[0036] Figure 9 It is a schematic diagram of the road result extracted from the DeepGlobe dataset by using the road segmentation method of the present invention;
[0037] Figure 10 It is an example of a remote sensing image of the Massachusetts dataset used in an embodiment of the present invention;
[0038] Figure 11 It is an example of a remote sensing image of the label in the Massachusetts dataset used in an embodiment of the present invention;
[0039] Figure 12 It is the road result extracted from the Massachusetts dataset by using the road segmentation method of the present invention;
[0040] Figure 13 It is a schematic flow diagram of a method for constructing a road segmentation model for remote sensing images provided by the present invention;
[0041] Figure 14 It is a schematic diagram of the structure of a device for constructing a road segmentation model for remote sensing images provided by the present invention. Detailed implementation manners
[0042] Next, the solutions provided in this specification will be described with reference to the accompanying drawings.
[0043] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0044] When the road surface is blocked by objects such as trees, complex environmental occlusion factors greatly reduce the recognizability of road features in remote sensing images. The surrounding context information and global information can be utilized to assist the model in making inferences. For example, the existence of a road nearby can be inferred through context information such as crosswalks, buildings along the street, and traffic signs. In terms of global information, factors such as regional boundaries, urban layouts, and traffic networks can be analyzed to assist in the inference. For example, if adjacent farmland and woods can be identified, it is very likely that there is a road at the junction of the two; at the same time, these roads beside the woods are also very likely to face the problem of being blocked.
[0045] When the road distribution is too dense, it is a current processing difficulty to accurately process and distinguish closely connected road structures within a very small spatial interval among highly crowded and visually intertwined road elements.
[0046] Based on the U-Net architecture, the present invention implements the strategies of synchronous hierarchical learning and cross-level feature interaction to comprehensively utilize information at each level for accurate and complete road extraction. Specifically, synchronous hierarchical learning means that for the feature maps at different levels and the information with different granularities they possess in the encoder and skip connection parts of the U-Net architecture, they are learned separately: the shallow feature maps are larger in size and have less loss of detailed features, which are suitable for extracting small-scale road details; the middle-level feature maps are balanced in size and content and are rich in key context information, which helps the model infer the road path; while the deep feature maps possess sufficient global information and high-level semantic information, which are suitable for learning the global dependencies of road structures and improving the model's understanding of coherent road structures.
[0047] In response to these hierarchical characteristics, the present invention designs corresponding modules at the corresponding positions of the model to fully learn these three types of information that are crucial for solving the missing problem. And cross-level feature interaction uses context information as a medium to fuse road details, global features, and context information respectively. By integrating features from different levels, the model can more fully understand the road structure, ensuring that the model can maintain the continuity and integrity of the structure even in a highly complex road network and solve the problem of missing road parts.
[0048] As Figures 1 to 13 shown, the present invention provides a method for constructing a road segmentation model for remote sensing images, including the following steps:
[0049] Collect data from two high-resolution remote sensing image road datasets, DeepGlobe and Massachusetts. These datasets contain sample pairs of 1024×1024 pixels and 1500×1500 pixels respectively, that is, the corresponding high-resolution remote sensing images and high-quality road labels.
[0050] Divide these sample pairs into a training set and a test set at a ratio of approximately 4:1. During the training process, preprocess these images and implement multiple data augmentation techniques, such as image flipping and random rotation, to enhance the diversity and richness of the training data. These operations are applied not only to the images but also to the corresponding label data.
[0051] Input the preprocessed and data-augmented training set images into the constructed road extraction network, set appropriate training parameters, train the road extraction network, and save the optimal training weights during the training process for testing.
[0052] Input the randomly divided test set images into the road segmentation model, load the optimal weights saved during the training process, and output the accurate segmentation results of the model for the road data in high-resolution remote sensing images.
[0053] Specifically, as Figure 1 shown, the remote sensing image road extraction network model includes a five-layer structure. In order from top to bottom, the first layer includes the first stage of the encoder and the upsampling stage of the decoder. The second layer includes the second stage of the encoder, the dynamic road detail matcher, and the fourth stage of the decoder. The third layer includes the third stage of the encoder, the dynamic road detail matcher, and the third stage of the decoder. The fourth layer includes the fourth stage of the encoder, the cross-context adaptive attention, and the second stage of the decoder. The fifth layer includes the fifth stage of the encoder, the cross-context adaptive attention, the multi-scale global information integrator, and the first stage of the decoder. The processing results of each module in each stage are simultaneously output to other modules in the same stage and the same module in the next stage, forming the remote sensing image road extraction network model.
[0054] The encoder uses the network parameters of the EfficientNet V2 network pre-trained on the ImageNet-1K dataset to initialize the encoder module, while other parameters of the network are initialized randomly.
[0055] The EfficientNet V2 network is a highly optimized neural network architecture optimized through compound scaling techniques, which not only improves the processing efficiency of the network but also ensures high accuracy.
[0056] The encoder module mainly performs downsampling through the backbone network. The backbone network has five downsampling stages, and each stage outputs feature maps of different dimensions for transmission to other parts or modules for learning. The structure of each stage is as follows:
[0057] The first stage consists of a two-layer structure. The first layer is configured with a standard convolutional kernel of size 3×3, a convolutional stride of 2, followed by batch normalization and the SiLU activation function. The second layer consists of two FusedMBConv blocks in the Efficientnet V2 backbone network.
[0058] As Figure 2 shown, the FusedMBConv block structure includes a standard convolutional kernel of size 3×3 with a stride of 1, an SE layer, and a standard convolutional kernel of size 1×1. After the feature image is processed by a standard convolutional kernel of size 3×3 with a stride of 1, an SE layer, and a standard convolutional kernel of size 1×1, the processed image is added to the input feature image to obtain the output feature image.
[0059] The second stage consists of a first-level encoder, which contains four groups of FusedMBConv blocks. This encoder module passes the processed feature map to the second-level encoder and the dynamic road detail matcher of the first-level skip connection for learning road details;
[0060] The third stage consists of a second-level encoder, which contains four groups of FusedMBConv blocks. This encoder module passes the processed feature map to the third-level encoder and the second dynamic road detail matcher for learning road details.
[0061] The fourth stage consists of a third-level encoder, which includes two stacked convolutional blocks. One includes six groups of MBConv blocks, and the other consists of nine groups of MBConv blocks. The encoder at this stage passes the feature map to the fourth-level encoder and the first cross-context adaptive attention module.
[0062] As Figure 3 shown, the MBConv block structure includes a standard convolutional kernel of size 1×1, a depthwise separable convolutional kernel of size 3×3, an SE layer, and a standard convolutional kernel of size 1×1. After the feature image is processed by a standard convolutional kernel of size 1×1, a depthwise separable convolutional kernel of size 3×3, an SE layer, and a standard convolutional kernel of size 1×1, the processed image is added to the input feature image to obtain the output feature image.
[0063] The fifth stage consists of a fourth-level encoder, which includes two layers. The first layer contains 15 groups of MBConv blocks, and the second layer is configured with a standard convolutional kernel of size 3×3, a stride of 2, batch normalization, and the SiLU activation function. This encoder outputs the feature map to the decoder and the multi-scale global information integrator.
[0064] The encoders in the above five stages altogether output five batches of feature maps. Among them, the feature map output by the first stage is the outermost layer feature map, which is rich in the most road detail information. The feature maps output by the layers downward successively contain less detail information and more context information. And the feature map output by the final stage is the deepest layer feature map, which contains the most global information and deep semantic information.
[0065] The dynamic road detail matcher is set at the skip connections between the first layer and the second layer and belongs to the shallow part of the overall model. It performs road detail matching on the feature maps output by the encoders at the corresponding stage through strip-shaped convolutions with multiple scales and multiple directions, learns a large amount of detail information in the shallow layer feature maps, and then passes the feature maps output by the dynamic road detail matcher to the dynamic detail matcher in the next layer or the cross-context adaptive attention that can fully learn context information and perform cross-level feature interactions. This feature map is also passed to the decoder at the corresponding stage.
[0066] The dynamic road detail matcher includes strip-shaped convolutions, dynamic snake-shaped convolutions, and depthwise separable convolutions, which improves the recognition ability of dense road structures. Also, through the characteristic of dynamically adjusting deformable convolutions, it optimizes the accurate capture of road geometric forms. It realizes that the dynamic road detail matcher can make the most of the rich detail information in the original feature map to achieve more complete road feature extraction and can effectively solve the problem of missing road parts in dense scenes.
[0067] As Figure 4 shown, the detailed structure of the dynamic road detail matcher is as follows:
[0068] The dynamic road detail matcher combines the feature map passed from the skip connection of the previous layer and the feature map output by the encoder of this layer as the input of the module. Since their sizes and the number of channels do not match, first use depthwise separable convolutions and pointwise convolutions to process the feature map passed from the skip connection of the previous layer to make it match the feature map output by the encoder of this layer, and use the result obtained by adding the two as the original input result. Input the original input result into three parallel branches to learn the detail information of the road and match the geometric form of the road. The first branch successively includes depthwise separable convolutions of 1×15 and 1×13 horizontally and a dynamic snake-shaped convolution of 1×11. The second branch successively includes depthwise separable convolutions of 15×1 and 13×1 vertically and a dynamic snake-shaped convolution of 11×1. The third branch successively includes depthwise separable convolutions of 15×15, 13×13, and 11×11 at a large scale.
[0069] Add the detail matching results of the three branches to the original input result, then process them using 1×1 convolution, batch normalization, and ReLU activation function respectively, and then add the processed result to the original input result through skip connection. After enhancing the non-linear processing ability of the module through the GELU activation function, it is used as the final output of the dynamic road detail matcher.
[0070] The output feature maps of the two dynamic road detail matchers are respectively passed to the next level of the skip connection and the decoder module at the corresponding level.
[0071] The cross-context adaptive attention is set at the third and fourth level skip connections, which belongs to the middle part of the overall model. The cross-context adaptive attention receives the output of the previous layer's dynamic detail matcher or the output of the previous layer's cross-context adaptive attention, uses various types and multi-scale convolutions to learn the context information of the middle layer feature map, and uses cross multi-head attention to fully perform cross-level feature interaction between the shallow detail information and the context information. Then, the feature map processed by the cross-context adaptive attention is passed to the cross-context adaptive attention or multi-scale global information integrator of the next layer, and this feature map is also passed to the decoder at the corresponding stage.
[0072] The cross-context adaptive attention introduces a temperature parameter to adjust the sensitivity of the multi-head attention mechanism, enabling the model to automatically adjust its processing strategy according to the complexity of the input features. This method not only improves the recognition accuracy of road features but also enhances the model's ability to handle occlusion and complex backgrounds.
[0073] In addition, the cross-context adaptive attention integrates various types of convolutions, enhances the multi-scale context information and cross-level cross-attention mechanism, significantly improving the model's ability to accurately extract and restore road features in complex backgrounds. Especially in scenarios dealing with visual interference and road occlusion, it can effectively improve road continuity and optimize the segmentation result.
[0074] Such as Figure 5 shown, the detailed structure of the cross-context adaptive attention is as follows: The cross-context adaptive attention combines the feature map passed from the previous layer's skip connection and the feature map output by the encoder of this layer as the input of the module. Since their sizes and number of channels do not match, it is necessary to first use depthwise separable convolution and pointwise convolution to process the feature map passed from the previous layer's skip connection to make it match the feature map output by the encoder of this layer. After this processing, their sizes and number of channels are the same.
[0075] Two batches of feature maps are respectively input into the multi-scale context supplementation mixer in two branches for sufficient context information mixing. The multi-scale context supplementation mixer sequentially includes a standard convolution with a convolution kernel size of 5×5, a dilated convolution with a size of 7×7 and a dilation coefficient of 2, and a large-range depthwise separable convolution with a size of 9×9.
[0076] The feature maps with sufficient context information learned from the two branches are subjected to feature mapping using 1×1 convolution to generate the query vector, value vector, and key vector of the attention mechanism for the two batches of feature maps. The key vectors of the two batches of feature maps are cross-swapped, and the value vectors and key vectors of the other batch of feature maps are respectively used for attention mechanism weighting, so that the model can fully perform cross-level feature interaction learning on the shallow road details and the context information in the middle layer. After the attention-weighted feature maps are used to learn the dependencies between channels using 1×1 convolution, the feature maps of the two branches are added as the final output of the cross-context adaptive attention. The outputs of the two cross-context adaptive attentions are respectively passed to the next level of the skip connection and the decoder module at the corresponding level.
[0077] The multi-scale global information integrator integrates the global information in the outputs of the two cross-context adaptive attentions and the output of the deepest encoder through an efficient global average pooling operation, enhances the model's ability to understand global information, and helps the model learn more coherent roads. The multi-scale global information integrator multiplies the global vector after global average pooling by the output of the deepest encoder to highlight the global information of the feature map, and then simultaneously inputs the three batches of feature maps into the efficient RepViT Blocks for in-depth modeling of the global information. The three batches of feature maps are integrated as the output of the module and passed to the decoder part.
[0078] The advantages of the multi-scale global information integrator are mainly reflected in its cross-level feature integration and powerful global information processing capabilities, and it is particularly suitable for structured visual tasks such as road extraction. The multi-scale global information integrator emphasizes the global relevance of roads and improves the model's understanding of the overall image structure, and is particularly suitable for solving the problem of partial missing caused by road occlusion.
[0079] As Figure 6 shown, the detailed structure of the multi-scale global information integrator is as follows:
[0080] The input of the multi-scale global information integrator includes the outputs of two cross-context adaptive attentions and the output feature map of the deepest encoder. The module uses a simple and effective global average pooling operation to learn the global information of the three batches of feature maps, multiplies the learned global information vector by the output of the deepest encoder, enhances the feature map with global information using a dot product operation, and then inputs the enhanced feature map into efficient RepViT Blocks to perform global information modeling using the Transformer architecture. The three batches of feature maps and the output of the original deepest encoder are concatenated, and then processed sequentially using 1×1 convolution, batch normalization, and ReLU activation function to obtain the final output of the multi-scale global information integrator, and the output feature map is passed to the decoder for road reproduction and reconstruction.
[0081] The decoder mainly processes the feature map through transposed convolution, batch normalization, and ReLU activation function to reconstruct the extracted road.
[0082] The detailed structure of the encoder part is as follows:
[0083] The decoder part is divided into five stages, and the processing flows of the first to the fourth stages are the same.
[0084] The processing flows of the first to the fourth stages are as follows:
[0085] First, use 1×1 convolution to reduce the number of channels of the feature map to one-fourth of the original, and perform integration and non-linear enhancement through batch normalization and ReLU activation function; then, use a transposed convolution with a size of 3×3 and a stride of 2 to enlarge the feature map, and process it again using batch normalization and ReLU activation function; finally, adjust the number of channels through 1×1 convolution to match the feature map of the next stage, and perform final integration using batch normalization and ReLU activation function. The output of each stage will be passed to the next decoder stage.
[0086] The decoder in the fifth stage processes sequentially using a transposed convolution with a size of 4 and a stride of 2, ReLU activation function, a standard convolution with a size of 3, ReLU activation function, a standard convolution with a size of 3, and Sigmoid function, and obtains the final road segmentation result.
[0087] To further verify the feasibility and effectiveness of the present invention, the relevant configurations of the remote sensing road extraction network are built: Ubuntu 18.04, Python 3.9.6, Pytorch 1.13.1, CUDA 11.6. The test results of the model are quantified using four commonly used evaluation metrics in the semantic segmentation task, namely accuracy, recall, intersection over union, and harmonic mean.
[0088] Adopt Figure 7Experiments were conducted on the road images in the DeepGlobe dataset, Figure 8 is a remote sensing road image of the label in the DeepGlobe dataset, Figure 9 is a schematic diagram of the road segmentation result extracted from the DeepGlobe dataset. From Figure 7 , Figure 8 and Figure 9 it can be seen that the road segmentation result is very similar to the remote sensing road image of the label. The segmentation of the road dense area reaches a very high accuracy rate, and the recognition rate of the covered roads is also very high. Only the broken road at the bottom of the picture is not recognized.
[0089] The experimental results are shown in Table 1:
[0090]
[0091] Table 1 Specific indicators on the DeepGlobe road extraction dataset
[0092] Using Figure 10 the remote sensing images in the DMassachusetts dataset shown for experiments, Figure 11 is a remote sensing road image of the label in the Massachusetts dataset remote sensing image, Figure 12 is a schematic diagram of the road segmentation result extracted from the Massachusetts dataset. From Figure 10 , Figure 11 and Figure 12 it can be seen that the road segmentation result is very similar to the remote sensing road image of the label. The segmentation of the road dense area reaches a very high accuracy rate, and the recognition rate of the covered roads is also very high. Only some broken roads in the middle of the picture are not recognized.
[0093] The experimental results are shown in Table 2:
[0094]
[0095] Table 2 Specific indicators on the Massachusetts road extraction dataset
[0096] Compared with the existing technology, the present invention proposes a dynamic road detail matcher suitable for extracting a large amount of road detail information in the shallow feature map, a cross-context adaptive attention suitable for extracting sufficient context information and cross-level feature interaction, and a multi-scale global information integrator capable of efficiently enhancing the global information of the road and improving the road continuity. The present invention is applicable to the field of high-resolution remote sensing image road segmentation, and is particularly suitable for a model to solve the problem of missing road parts caused by road occlusion or dense distribution.
[0097] Figure 13The flowchart shows a method for constructing a road segmentation model for remote sensing images provided by an embodiment of this specification. It includes:
[0098] 110. Obtain a remote sensing image training sample set. Any group of remote sensing image training samples includes a sample image with a partially occluded road and its corresponding road label image.
[0099] 120. Input the sample images of the remote sensing image training sample set into an initial road segmentation model to obtain a road segmentation prediction image of the sample images.
[0100] 130. Based on the difference between the road label image and the road segmentation prediction image of the sample image, perform parameter iteration on the initial road segmentation model to obtain a road segmentation model.
[0101] The initial road segmentation model includes a dynamic road detail matcher, a cross-context adaptive attention mechanism, and a multi-scale global information integrator. Among them:
[0102] The dynamic road detail matcher is used to obtain a shallow feature map through the feature image of the sample image. The shallow feature map contains the geometric shape of the road.
[0103] The cross-context adaptive attention mechanism is used to obtain a middle feature map based on the feature image and the shallow feature map. The middle feature map contains the road surrounding environment information of the occluded road.
[0104] The multi-scale global information integrator is used to obtain a deep feature map based on the feature image and the middle feature map. The deep feature map contains the deep semantic information and global information of the road. The deep semantic information of the road includes the mutual relationship and attributes between the occluded road and surrounding objects, and the global information includes the overall layout and road structure information of the occluded road.
[0105] Based on a further embodiment, the initial road segmentation model further includes an encoder; the encoder is used to obtain the feature image of the sample image according to the sample image.
[0106] The encoder includes a preprocessor, a first-level encoder, a second-level encoder, a third-level encoder, and a fourth-level encoder.
[0107] The preprocessor is used to extract the image features of the sample image to obtain a first feature image, and output the first feature image to the first-level encoder and the dynamic road detail matcher. Among them, the size of the first feature image is smaller than that of the sample image.
[0108] The first-level encoder is used to extract the image features of the first feature image to obtain the second feature image, and output the second feature image to the second-level encoder and the dynamic road detail matcher. Among them, the size of the second feature image is smaller than that of the first feature image, the number of channels of the second feature image is larger than that of the first feature image, and the context information of the second feature image is more than that of the first feature image.
[0109] The second-level encoder is used to extract the image features of the second feature image to obtain the third feature image, and output the third feature image to the third-level encoder and the dynamic road detail matcher.
[0110] Among them, the size of the third feature image is smaller than that of the second feature image, the number of channels of the third feature image is larger than that of the second feature image, and the context information of the third feature image is more than that of the second feature image.
[0111] The third-level encoder is used to extract the image features of the third feature image to obtain the fourth feature image, and output the fourth feature image to the third-level encoder and the cross-context adaptive attention mechanism.
[0112] Among them, the size of the fourth feature image is smaller than that of the third feature image, the number of channels of the fourth feature image is larger than that of the third feature image, and the context information of the fourth feature image is more than that of the third feature image.
[0113] The fourth-level encoder is used to extract the image features of the fourth encoder feature image to obtain the fifth feature map of the sample image. Among them, the size of the fifth feature image is smaller than that of the fourth feature image, the number of channels of the fifth feature image is larger than that of the fourth feature image, and the context information of the fifth feature image is more than that of the fourth feature image.
[0114] Based on a further embodiment, the road segmentation model includes a first dynamic road detail matcher and a second dynamic road detail matcher.
[0115] The first dynamic road detail matcher includes at least two parallel branch processing modules, and at least one of the parallel branch processing modules includes a dynamic snake-shaped convolutional layer;
[0116] The first dynamic road detail matcher is used to obtain the first shallow feature map according to the first feature image and the second feature image;
[0117] The second dynamic road detail matcher includes at least two parallel branch processing modules, and at least one of the parallel branch processing modules includes a dynamic snake-shaped convolutional layer.
[0118] The second dynamic road detail matcher is used to obtain a second shallow feature map according to a second feature image and a first shallow feature map, where the size of the second shallow feature map is smaller than that of the first shallow feature map, the number of channels of the second shallow feature map is more than that of the first shallow feature map, and the second shallow feature map contains more context information than the first shallow feature map.
[0119] According to a further embodiment, the road segmentation model includes a first cross-context adaptive attention mechanism and a second cross-context adaptive attention mechanism. The first cross-context adaptive attention mechanism includes two multi-scale context complementary mixers, and the second cross-context adaptive attention mechanism includes two multi-scale context complementary mixers.
[0120] The first cross-context adaptive attention mechanism is used to input the matching result images obtained after matching the second shallow feature map and the fourth feature image into two multi-scale context complementary mixers respectively to obtain a first cross-attention feature image and a second cross-attention feature image.
[0121] Perform feature mapping on the first cross-attention feature image to obtain a query vector, a value vector, and a key vector of the first attention mechanism. Perform feature mapping on the second cross-attention feature image to obtain a query vector, a value vector, and a key vector of the second attention mechanism.
[0122] Exchange the query vector of the first attention mechanism and the query vector of the second attention mechanism;
[0123] After weighting the first cross-attention feature image by using the value vector and the key vector of the first attention mechanism and the query vector of the second attention mechanism, obtain a weighted first cross-attention feature image.
[0124] After weighting the second cross-attention feature image by using the value vector and the key vector of the second attention mechanism and the query vector of the first attention mechanism, obtain a weighted second cross-attention feature image.
[0125] Add the weighted first cross-attention feature image and the weighted second cross-attention feature image after learning the channel interdependence relationship to obtain a first middle-level feature map.
[0126] The second cross-context adaptive attention mechanism is used to input the matching result images obtained after matching the first middle-level feature map and the fifth feature image into two multi-scale context complementary mixers respectively to obtain a third cross-attention feature image and a fourth cross-attention feature image;
[0127] Perform feature mapping on the third cross-attention feature image to obtain the query vector, value vector, and key vector of the third attention mechanism; perform feature mapping on the fourth cross-attention feature image to obtain the query vector, value vector, and key vector of the fourth attention mechanism;
[0128] Exchange the query vector of the third attention mechanism and the query vector of the fourth attention mechanism;
[0129] After weighting the third cross-attention feature image using the value vector and key vector of the third attention mechanism, and the query vector of the fourth attention mechanism, obtain the weighted third cross-attention feature image;
[0130] After weighting the fourth cross-attention feature image using the value vector and key vector of the fourth attention mechanism, and the query vector of the third attention mechanism, obtain the weighted fourth cross-attention feature image;
[0131] Add the weighted third cross-attention feature image and the weighted fourth cross-attention feature image after learning the dependencies between channels to obtain the second middle-level feature map, where the size of the second middle-level feature map is smaller than that of the first middle-level feature map, the number of channels of the second middle-level feature map is more than that of the first middle-level feature map, and the second middle-level feature map contains more context information than the first middle-level feature map.
[0132] Based on a further embodiment, the multi-scale global information integrator includes three global average pooling layers;
[0133] The multi-scale global information integrator is used to parallelly input the fifth feature map, the first middle-level feature map, and the second middle-level feature map into the three global average pooling layers to obtain the global information vectors output by each global average pooling layer;
[0134] After multiplying each global information vector by the first feature map respectively, obtain the global information images corresponding to each global information vector;
[0135] Input each global information image into the RepViT Blocks module for global information modeling to obtain the global information feature images corresponding to each global information image;
[0136] Based on the global information feature images corresponding to each global information image and the first feature map, obtain the deep feature map.
[0137] Based on a further embodiment, the initial road segmentation model further includes a decoder; the decoder is used to obtain the road segmentation prediction image according to the feature image, the shallow feature map, the middle-level feature map, and the deep feature map;
[0138] The decoder includes a first decoder, a second decoder, a third decoder, a fourth decoder, and a fifth decoder;
[0139] A first decoder, configured to integrate a deep feature map and a second middle-level feature map to obtain a first decoded feature image;
[0140] A second decoder, configured to integrate the first decoded feature image and the first middle-level feature map to obtain a second decoded feature image, wherein the size of the second decoded feature image is larger than that of the first decoded feature image, and the number of channels of the second decoded feature image is smaller than that of the first decoded feature image;
[0141] A third decoder, configured to integrate the second decoded feature image and a second shallow-level feature map to obtain a third decoded feature image, wherein the size of the third decoded feature image is larger than that of the second decoded feature image, and the number of channels of the third decoded feature image is smaller than that of the second decoded feature image;
[0142] A fourth decoder, configured to integrate the third decoded feature image and a first shallow-level feature map to obtain a fourth decoded feature image, wherein the size of the fourth decoded feature image is larger than that of the third decoded feature image, and the number of channels of the fourth decoded feature image is smaller than that of the third decoded feature image;
[0143] A fifth decoder, configured to process the fourth decoded feature image to obtain a road segmentation prediction image of the sample image.
[0144] Corresponding to the above method provided by the present invention, Figure 14 The structure diagram of a device for constructing a road segmentation model of a remote sensing image provided by an embodiment of this specification is shown. The device includes:
[0145] An acquisition module, configured to obtain a remote sensing image training sample set, and any set of remote sensing image training samples includes a sample image with a partially covered road and its corresponding road label image;
[0146] An input module, configured to input the sample image of the remote sensing image training sample set into an initial road segmentation model to obtain a road segmentation prediction image of the sample image;
[0147] A training module, configured to perform parameter iteration on the initial road segmentation model based on the difference between the road label image and the road segmentation prediction image of the sample image to obtain a road segmentation model;
[0148] The initial road segmentation model includes a dynamic road detail matcher, a cross-context adaptive attention mechanism, and a multi-scale global information integrator, wherein:
[0149] The dynamic road detail matcher is configured to obtain a shallow-level feature map through the feature image of the sample image, and the geometric form of the road is highlighted in the shallow-level feature map;
[0150] A cross-context adaptive attention mechanism is used to obtain a middle-level feature map based on a feature image and a shallow feature map. The middle-level feature map contains information about the road surrounding environment for predicting the occluded road.
[0151] A multi-scale global information integrator is used to obtain a deep feature map based on the feature image and the middle-level feature map. The deep feature map contains deep semantic information and global information of the road. The deep semantic information of the road describes the mutual relationship and attributes between the occluded road and surrounding objects, and the global information describes information about the overall layout and road structure of the road for predicting the occluded road.
[0152] Based on a further embodiment, the initial road segmentation model further includes an encoder. The encoder is used to obtain a feature image of a sample image according to the sample image.
[0153] The encoder includes a pre-processor, a first-level encoder, a second-level encoder, a third-level encoder, and a fourth-level encoder.
[0154] The pre-processor is used to extract the image features of the sample image to obtain a first feature image, and output the first feature image to the first-level encoder and a dynamic road detail matcher. Among them, the size of the first feature image is smaller than that of the sample image.
[0155] The first-level encoder is used to extract the image features of the first feature image to obtain a second feature image, and output the second feature image to the second-level encoder and the dynamic road detail matcher. Among them, the size of the second feature image is smaller than that of the first feature image, the number of channels of the second feature image is greater than that of the first feature image, and the context information of the second feature image is more than that of the first feature image.
[0156] The second-level encoder is used to extract the image features of the second feature image to obtain a third feature image, and output the third feature image to the third-level encoder and the dynamic road detail matcher.
[0157] Among them, the size of the third feature image is smaller than that of the second feature image, the number of channels of the third feature image is greater than that of the second feature image, and the context information of the third feature image is more than that of the second feature image.
[0158] The third-level encoder is used to extract the image features of the third feature image to obtain a fourth feature image, and output the fourth feature image to the third-level encoder and the cross-context adaptive attention mechanism.
[0159] Among them, the size of the fourth feature image is smaller than that of the third feature image, the number of channels of the fourth feature image is greater than that of the third feature image, and the context information of the fourth feature image is more than that of the third feature image.
[0160] The fourth - level encoder is used to extract the image features of the fourth encoder feature image to obtain the fifth feature map of the sample image. Among them, the size of the fifth feature image is smaller than that of the fourth feature image, the number of channels of the fifth feature image is larger than that of the fourth feature image, and the context information of the fifth feature image is more than that of the fourth feature image.
[0161] Based on a further embodiment, the road segmentation model includes a first dynamic road detail matcher and a second dynamic road detail matcher.
[0162] The first dynamic road detail matcher includes at least two parallel branch processing modules, and at least one of the parallel branch processing modules includes a dynamic snake - shaped convolutional layer;
[0163] The first dynamic road detail matcher is used to obtain the first shallow feature map according to the first feature image and the second feature image;
[0164] The second dynamic road detail matcher includes at least two parallel branch processing modules, and at least one of the parallel branch processing modules includes a dynamic snake - shaped convolutional layer.
[0165] The second dynamic road detail matcher is used to obtain the second shallow feature map according to the second feature image and the first shallow feature map. Among them, the size of the second shallow feature map is smaller than that of the first shallow feature map, the number of channels of the second shallow feature map is larger than that of the first shallow feature map, and the context information contained in the second shallow feature map is more than that of the first shallow feature map.
[0166] Based on a further - more embodiment, the road segmentation model includes a first cross - context adaptive attention mechanism and a second cross - context adaptive attention mechanism. The first cross - context adaptive attention mechanism includes two multi - scale context - supplementing mixers, and the second cross - context adaptive attention mechanism includes two multi - scale context - supplementing mixers.
[0167] The first cross - context adaptive attention mechanism is used to input the matching result images obtained after matching the second shallow feature map and the fourth feature image into two multi - scale context - supplementing mixers respectively to obtain the first cross - attention feature image and the second cross - attention feature image.
[0168] Perform feature mapping on the first cross - attention feature image to obtain the query vector, value vector, and key vector of the first attention mechanism, and perform feature mapping on the second cross - attention feature image to obtain the query vector, value vector, and key vector of the second attention mechanism.
[0169] Exchange the query vector of the first attention mechanism and the query vector of the second attention mechanism;
[0170] After weighting the first cross-attention feature image using the value vector and key vector of the first attention mechanism, and the query vector of the second attention mechanism, the weighted first cross-attention feature image is obtained.
[0171] After weighting the second cross-attention feature image using the value vector and key vector of the second attention mechanism, and the query vector of the first attention mechanism, the weighted second cross-attention feature image is obtained.
[0172] After learning the channel interdependencies between the weighted first cross-attention feature image and the weighted second cross-attention feature image and adding them together, the first middle-level feature map is obtained.
[0173] The second cross-context adaptive attention unit is used to input the matching result images obtained by matching the first middle-level feature map and the fifth feature image into two multi-scale context supplementation mixers respectively, to obtain the third cross-attention feature image and the fourth cross-attention feature image;
[0174] Feature mapping is performed on the third cross-attention feature image to obtain the query vector, value vector, and key vector of the third attention mechanism; feature mapping is performed on the fourth cross-attention feature image to obtain the query vector, value vector, and key vector of the fourth attention mechanism;
[0175] Exchange the query vector of the third attention mechanism and the query vector of the fourth attention mechanism;
[0176] After weighting the third cross-attention feature image using the value vector and key vector of the third attention mechanism, and the query vector of the fourth attention mechanism, the weighted third cross-attention feature image is obtained;
[0177] After weighting the fourth cross-attention feature image using the value vector and key vector of the fourth attention mechanism, and the query vector of the third attention mechanism, the weighted fourth cross-attention feature image is obtained;
[0178] After learning the channel interdependencies between the weighted third cross-attention feature image and the weighted fourth cross-attention feature image and adding them together, the second middle-level feature map is obtained, where the size of the second middle-level feature map is smaller than that of the first middle-level feature map, the number of channels of the second middle-level feature map is more than that of the first middle-level feature map, and the context information contained in the second middle-level feature map is more than that of the first middle-level feature map.
[0179] Based on a further embodiment, the multi-scale global information integrator includes three global average pooling layers;
[0180] A multi-scale global information integrator that parallelly inputs a fifth feature map, a first middle-level feature map, and a second middle-level feature map into three global average pooling layers to obtain global information vectors output by each global average pooling layer;
[0181] After multiplying each global information vector by the first feature map respectively, global information images corresponding to each global information vector are obtained;
[0182] Input each global information image into the RepViT Blocks module for global information modeling to obtain global information feature images corresponding to each global information image;
[0183] Based on the global information feature images corresponding to each global information image and the first feature map, a deep feature map is obtained.
[0184] Based on a further embodiment, the initial road segmentation model further includes a decoder; the decoder is used to obtain a road segmentation prediction image according to the feature image, the shallow feature map, the middle-level feature map, and the deep feature map;
[0185] The decoder includes a first decoder, a second decoder, a third decoder, a fourth decoder, and a fifth decoder;
[0186] The first decoder is used to integrate and process the deep feature map and the second middle-level feature map to obtain a first decoded feature image;
[0187] The second decoder is used to integrate and process the first decoded feature image and the first middle-level feature map to obtain a second decoded feature image, wherein the size of the second decoded feature image is larger than that of the first decoded feature image, and the number of channels of the second decoded feature image is less than that of the first decoded feature image;
[0188] The third decoder is used to integrate and process the second decoded feature image and the second shallow feature map to obtain a third decoded feature image, wherein the size of the third decoded feature image is larger than that of the second decoded feature image, and the number of channels of the third decoded feature image is less than that of the second decoded feature image;
[0189] The fourth decoder is used to integrate and process the third decoded feature image and the first shallow feature map to obtain a fourth decoded feature image, wherein the size of the fourth decoded feature image is larger than that of the third decoded feature image, and the number of channels of the fourth decoded feature image is less than that of the third decoded feature image;
[0190] The fifth decoder is used to process the fourth decoded feature image to obtain a road segmentation prediction image of the sample image.
[0191] According to an embodiment of another aspect, a method for road segmentation of remote sensing images is further provided, including:
[0192] Obtain a target remote sensing image and a road segmentation model constructed according to the method for constructing a road segmentation model in the above embodiment. The target remote sensing image includes a target road partially covered.
[0193] Input the target remote sensing image into the road segmentation model to obtain an uncovered target road image in the target remote sensing image.
[0194] According to an embodiment of another aspect, there is provided an electronic device including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the method for constructing a road segmentation model of a remote sensing image in the above technical solution.
[0195] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having a computer program stored thereon. When the computer program is executed on a computer, the computer is made to execute the method for constructing a road segmentation model of a remote sensing image.
[0196] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in this application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0197] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of this application. It should be understood that the above is only the specific embodiment of this application and is not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of this application should be included in the protection scope of this application.
Claims
1. A method for constructing a road segmentation model of remote sensing images, characterized in that, The method includes: Obtaining a remote sensing image training sample set, where any set of remote sensing image training samples includes a sample image with a partially occluded road and its corresponding road label image; Inputting the sample images of the remote sensing image training sample set into an initial road segmentation model to obtain a road segmentation prediction image of the sample images; Based on the difference between the road label image of the sample image and the road segmentation prediction image, performing parameter iteration on the initial road segmentation model to obtain a road segmentation model; The initial road segmentation model includes a dynamic road detail matcher, a cross-context adaptive attention mechanism, and a multi-scale global information integrator, where: The dynamic road detail matcher is used to obtain a shallow feature map through the feature image of the sample image, and the shallow feature map contains the geometric form of the road; The cross-context adaptive attention mechanism is used to obtain a middle feature map according to the feature image and the shallow feature map, and the middle feature map contains the road surrounding environment information of the occluded road; The multi-scale global information integrator is used to obtain a deep feature map according to the feature image and the middle feature map, and the deep feature map contains the deep semantic information and global information of the road. The deep semantic information of the road includes the mutual relationship and attributes between the occluded road and surrounding objects, and the global information includes the overall layout of the road and the information of the road structure of the occluded road; Among them, the dynamic road detail matcher includes strip-shaped convolution, dynamic snake-shaped convolution, and depthwise separable convolution; the cross-context adaptive attention mechanism includes a multi-scale context supplement mixer, and the multi-scale context supplement mixer sequentially includes a standard convolution with a convolution kernel size of 5×5, a dilated convolution with a size of 7×7 and a dilation coefficient of 2, and a large-range depthwise separable convolution with a size of 9×9; the multi-scale global information integrator includes three global average pooling layers.
2. The method according to claim 1, wherein The initial road segmentation model further includes an encoder; the encoder is used to obtain the feature image of the sample image according to the sample image; The encoder includes a pre-processor, a first-level encoder, a second-level encoder, a third-level encoder, and a fourth-level encoder; The pre-processor is used to extract the image features of the sample image to obtain a first feature image, and output the first feature image to the first-level encoder and the dynamic road detail matcher. Among them, the size of the first feature image is smaller than that of the sample image; The first-level encoder is used to extract the image features of the first feature image to obtain a second feature image, and output the second feature image to the second-level encoder and the dynamic road detail matcher. Among them, the size of the second feature image is smaller than that of the first feature image, the number of channels of the second feature image is greater than that of the first feature image, and the context information of the second feature image is more than that of the first feature image; The second-level encoder is used to extract the image features of the second feature image, obtain a third feature image, and output the third feature image to the third-level encoder and the dynamic road detail matcher. Wherein, the size of the third feature image is smaller than that of the second feature image, the number of channels of the third feature image is larger than that of the second feature image, and the context information of the third feature image is more than that of the second feature image; The third-level encoder is used to extract the image features of the third feature image, obtain a fourth feature image, and output the fourth feature image to the fourth-level encoder and the cross-context adaptive attention mechanism. Wherein, the size of the fourth feature image is smaller than that of the third feature image, the number of channels of the fourth feature image is larger than that of the third feature image, and the context information of the fourth feature image is more than that of the third feature image; The fourth-level encoder is used to extract the image features of the fourth feature image to obtain a fifth feature image of the sample image. Wherein, the size of the fifth feature image is smaller than that of the fourth feature image, the number of channels of the fifth feature image is larger than that of the fourth feature image, and the context information of the fifth feature image is more than that of the fourth feature image.
3. The method according to claim 2, wherein, The road segmentation model includes a first dynamic road detail matcher and a second dynamic road detail matcher; The first dynamic road detail matcher includes at least two parallel branch processing modules, and at least one of the parallel branch processing modules includes a dynamic snake convolution layer; The first dynamic road detail matcher is used to obtain a first shallow feature map according to the first feature image and the second feature image; The second dynamic road detail matcher includes at least two parallel branch processing modules, and at least one of the parallel branch processing modules includes a dynamic snake convolution layer, The second dynamic road detail matcher is used to obtain a second shallow feature map according to the second feature image and the first shallow feature map. Wherein, the size of the second shallow feature map is smaller than that of the first shallow feature map, the number of channels of the second shallow feature map is larger than that of the first shallow feature map, and the context information contained in the second shallow feature map is more than that of the first shallow feature map.
4. The method according to claim 3, wherein, The road segmentation model includes a first cross-context adaptive attention mechanism and a second cross-context adaptive attention mechanism. The first cross-context adaptive attention mechanism includes two multi-scale context supplement mixing units, and the second cross-context adaptive attention mechanism includes two multi-scale context supplement mixing units; The first cross-context adaptive attention mechanism is used to match the second shallow feature map and the fourth feature image, and then input the matching results into the two multi-scale context supplement mixing units respectively to obtain a first cross-attention feature image and a second cross-attention feature image; Feature map the first cross-attention feature image to obtain the query vector, value vector, and key vector of the first attention mechanism, and feature map the second cross-attention feature image to obtain the query vector, value vector, and key vector of the second attention mechanism; Exchange the query vector of the first attention mechanism and the query vector of the second attention mechanism; After weighting the first cross-attention feature image using the value vector and key vector of the first attention mechanism, and the query vector of the second attention mechanism, obtain the weighted first cross-attention feature image; After weighting the second cross-attention feature image using the value vector and key vector of the second attention mechanism, and the query vector of the first attention mechanism, obtain the weighted second cross-attention feature image; Add the weighted first cross-attention feature image and the weighted second cross-attention feature image after learning the channel interdependencies to obtain the first middle-level feature map; The second cross-context adaptive attention unit is configured to, after matching the first middle-level feature map and the fifth feature image, input the matching results into two of the multi-scale context supplementation mixers respectively to obtain a third cross-attention feature image and a fourth cross-attention feature image; Feature map the third cross-attention feature image to obtain the query vector, value vector, and key vector of the third attention mechanism; feature map the fourth cross-attention feature image to obtain the query vector, value vector, and key vector of the fourth attention mechanism; Exchange the query vector of the third attention mechanism and the query vector of the fourth attention mechanism; After weighting the third cross-attention feature image using the value vector and key vector of the third attention mechanism, and the query vector of the fourth attention mechanism, obtain the weighted third cross-attention feature image; After weighting the fourth cross-attention feature image using the value vector and key vector of the fourth attention mechanism, and the query vector of the third attention mechanism, obtain the weighted fourth cross-attention feature image; Add the weighted third cross-attention feature image and the weighted fourth cross-attention feature image after learning the channel interdependencies to obtain a second middle-level feature map, where the size of the second middle-level feature map is smaller than that of the first middle-level feature map, the number of channels of the second middle-level feature map is more than that of the first middle-level feature map, and the second middle-level feature map contains more context information than the first middle-level feature map.
5. The method according to claim 4, wherein The multi-scale global information integrator includes three global average pooling layers; The multi-scale global information integrator is configured to parallelly input the fifth feature map, the first middle-level feature map, and the second middle-level feature map into the three global average pooling layers to obtain the global information vectors output by each of the global average pooling layers; After multiplying each of the global information vectors by the first feature map, global information images corresponding to the global information vectors are obtained; Each of the global information images is input into a RepViT Blocks module for global information modeling to obtain global information feature images corresponding to the global information images; Based on the global information feature images corresponding to the global information images and the first feature map, the deep feature map is obtained.
6. The method according to claim 5, wherein The initial road segmentation model further includes a decoder; the decoder is configured to obtain the road segmentation prediction image according to the feature image, the shallow feature map, the middle feature map, and the deep feature map; The decoder includes a first decoder, a second decoder, a third decoder, a fourth decoder, and a fifth decoder; The first decoder is configured to perform integration processing on the deep feature map and the second middle feature map to obtain a first decoded feature image; The second decoder is configured to perform integration processing on the first decoded feature image and the first middle feature map to obtain a second decoded feature image, wherein the size of the second decoded feature image is larger than that of the first decoded feature image, and the number of channels of the second decoded feature image is smaller than that of the first decoded feature image; The third decoder is configured to perform integration processing on the second decoded feature image and the second shallow feature map to obtain a third decoded feature image, wherein the size of the third decoded feature image is larger than that of the second decoded feature image, and the number of channels of the third decoded feature image is smaller than that of the second decoded feature image; The fourth decoder is configured to perform integration processing on the third decoded feature image and the first shallow feature map to obtain a fourth decoded feature image, wherein the size of the fourth decoded feature image is larger than that of the third decoded feature image, and the number of channels of the fourth decoded feature image is smaller than that of the third decoded feature image; The fifth decoder is configured to process the fourth decoded feature image to obtain the road segmentation prediction image of the sample image.
7. An apparatus for constructing a road segmentation model of a remote sensing image, characterized in that The apparatus includes: An acquisition module, configured to acquire a remote sensing image training sample set, and any group of remote sensing image training samples includes a sample image with a partially covered road and its corresponding road label image; An input module, configured to input the sample image of the remote sensing image training sample set into an initial road segmentation model to obtain the road segmentation prediction image of the sample image; A training module, configured to perform parameter iteration on the initial road segmentation model based on the difference between the road label image of the sample image and the road segmentation prediction image to obtain a road segmentation model; The initial road segmentation model includes a dynamic road detail matcher, a cross-context adaptive attention mechanism, and a multi-scale global information integrator, wherein: The dynamic road detail matcher is configured to obtain a shallow feature map through the feature image of the sample image, and the shallow feature map contains the geometric shape of the road; The cross-context adaptive attention mechanism is used to obtain a middle-level feature map based on the feature image and the shallow feature map, and the middle-level feature map contains the road surrounding environment information of the occluded road; The multi-scale global information integrator is used to obtain a deep feature map based on the feature image and the middle-level feature map. The deep feature map contains the deep semantic information and global information of the road. The deep semantic information of the road includes the mutual relationship and attributes between the occluded road and surrounding objects, and the global information includes the overall layout and road structure information of the occluded road; Among them, the dynamic road detail matcher includes strip-shaped convolution, dynamic snake-shaped convolution, and depthwise separable convolution; the cross-context adaptive attention mechanism includes a multi-scale context supplement mixer, and the multi-scale context supplement mixer sequentially includes a standard convolution with a convolution kernel size of 5×5, a dilated convolution with a size of 7×7 and a dilation coefficient of 2, and a large-range depthwise separable convolution with a size of 9×9; the multi-scale global information integrator includes three global average pooling layers.
8. A method for segmenting roads in a remote sensing image, characterized in that, The method includes: Obtaining a target remote sensing image and a road segmentation model constructed according to the construction method of the road segmentation model according to any one of claims 1 to 6, wherein the target remote sensing image includes a target road partially covered; Inputting the target remote sensing image into the road segmentation model to obtain an uncovered target road image in the target remote sensing image.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the program, the road segmentation method of the remote sensing image as claimed in claim 8 is implemented.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, When the processor executes the program, the construction method of the road segmentation model of the remote sensing image according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Remote sensing image road segmentation method fusing multi-scale features and double attention mechanism
CN117078943A
Remote sensing image dense road segmentation method using strip-shaped features
CN118429356A