High-resolution remote sensing image road adaptive extraction method, system and device based on multi-scale dynamic convolution enhancement, and medium
By designing a multi-scale adaptive serpentine convolution module and an encoder-decoder model, the adaptability and connectivity problems of road extraction in high-resolution remote sensing images are solved, and high-precision road extraction in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510657546.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, when extracting roads in high-resolution remote sensing images, there are problems such as insufficient adaptability of road feature extraction and geometric morphology, weak ability to integrate multi-scale features and long-distance dependency, resulting in poor road extraction accuracy and connectivity in complex scenarios.
A multi-scale adaptive serpentine convolution module is designed, combined with the encoder-decoder model, and adaptively adjust the rotation angle and offset of strip convolution to achieve the adaptation of complex road forms, and enhance the perception ability of road spatial continuity characteristics through multi-scale feature fusion strategy.
It significantly improves the completeness and accuracy of road extraction in complex scenarios, enhances the perception of road spatial continuity in occlusion environments, and improves the accuracy of road extraction.
Smart Images

Figure CN120451796A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of adaptive road extraction, and in particular relates to a method, system, equipment and medium for adaptive extraction of roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement. Background Art
[0002] With the continuous advancement of satellite and aerial photography technology, obtaining high-resolution remote sensing imagery has become increasingly convenient. Currently, the spatial resolution of some commercial remote sensing satellites has reached sub-meter levels, clearly capturing the detailed features of ground objects. This high resolution allows the texture, shape, and orientation of roads to be accurately displayed in the imagery, providing a rich data foundation for road extraction.
[0003] As an important geographic element, accurate extraction of roads is of great significance in urban planning, traffic management, disaster response, and other fields. However, road extraction from high-resolution remote sensing imagery faces many challenges. Roads vary in width, material, color, and shape, and there are a large number of irregular curved roads, dead-end roads, and temporary roads. Road construction standards and styles vary across regions. For example, winding mountain roads differ significantly from straight, regular roads in plains. In real-world scenarios, roads face many complex environmental factors. Trees and buildings on both sides of the road, as well as the shadows they cast, can obscure road information. Seasonal changes in vegetation cover can make the spectral characteristics of roads and vegetation similar, increasing the difficulty of extraction and leading to poor connectivity in road extraction results in complex scenarios.
[0004] From the perspective of existing technical solutions, traditional road extraction algorithms are difficult to train, have insufficient extraction accuracy, and suffer from low processing efficiency when faced with massive remote sensing data. Although deep learning-based intelligent interpretation technology has been widely used in road extraction tasks, its core technical solution still has significant flaws:
[0005] Classic convolutional neural network models represented by UNet and LinkNet use fixed-size square convolution kernels for sliding window feature extraction. This isotropic feature sampling method is naturally not compatible with the linear continuous structural characteristics of roads, and it is difficult to effectively capture the long strip spatial distribution characteristics of roads. Although the DeepLab series of models expands the receptive field through void convolution technology, its grid-based feature sampling method is not compatible with the linear continuity characteristics of roads, resulting in limited road segmentation accuracy in complex scenarios. Some technical solutions that use strip convolution have poor adaptability to the directional morphology of roads and difficulty in achieving effective fusion of multi-scale road features due to their technical characteristics of fixed convolution direction and single convolution kernel size.
[0006] Therefore, existing technologies have obvious technical shortcomings when dealing with the problems of extracting complex road features and maintaining connectivity in occluded scenes. It is urgent to study multi-scale adaptive road feature extraction methods to break through the existing technical bottlenecks and improve the accuracy and integrity of road extraction in complex scenes.
[0007] Patent application document with publication number CN116778318A discloses a convolutional neural network remote sensing image road extraction model and method. The method specifically discloses a convolutional neural network model combined with a multi-scale feature encoder module to improve road extraction accuracy. However, its multi-scale module uses three square convolution kernels of 1×1, 3×3, and 5×5, which cannot effectively extract strip road features.
[0008] The patent application document with publication number CN116543304A discloses a method and device for extracting roads from remote sensing images based on a convolutional network. The method specifically discloses an encoder-decoder model composed of a position strip-attention hole convolutional network to enhance the spatial feature extraction of roads. However, the strip-position attention module used in this method uses strip convolutions in four directions: horizontal, vertical, left diagonal, and right diagonal. On the one hand, the use of strip convolutions in four directions makes the model structure complex and redundant, and has poor flexibility; on the other hand, the strip convolutions used have a fixed size, and cannot better extract road features from the multi-scale strip convolution aspect.
[0009] In summary, the current remote sensing image road extraction technology has the following significant technical shortcomings:
[0010] 1) Insufficient adaptability between road feature extraction and geometric morphology
[0011] Some technical solutions use square convolution kernels, and their feature extraction methods are not sufficiently compatible with the linear continuous structure of roads; the strip convolution directions used in some solutions are limited to fixed angles such as horizontal and vertical, and cannot adaptively match any direction and angle of the road, resulting in poor road feature extraction capabilities.
[0012] 2) Weak multi-scale feature fusion and long-distance dependency extraction capabilities
[0013] Some solutions use a single, fixed-size convolution kernel, lacking the ability to extract multi-scale features. This leads to fragmented road extraction results due to their poor ability to extract long-range dependencies when dealing with shadows from trees and buildings. For example, a partially obscured road cannot be restored to connectivity through long-range dependencies during feature transfer. Summary of the Invention
[0014] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a method, system, equipment and medium for adaptive road extraction from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement, by designing a multi-scale direction-adaptive serpentine convolution module to enhance the extraction capability of linear continuous complex road features and the acquisition capability of multi-scale and long-distance information; and propose an encoder-decoder method by connecting a multi-scale direction-adaptive serpentine convolution module in series in the coding layer to improve the integrity and accuracy of road extraction in complex scenes.
[0015] In order to achieve the above object, the technical solution adopted by the present invention is:
[0016] A method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement includes the following steps:
[0017] Step 1: Construct a high-resolution remote sensing image road dataset, including original images and label data; and divide it into training set, validation set, and test set;
[0018] Step 2: Build an encoder-decoder model combined with a multi-scale direction-adaptive snake convolution module;
[0019] Step 3: Use the training set and validation set divided in step 1 to train the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module constructed in step 2, and save the trained encoder-decoder model combined with the multi-scale direction adaptive snake convolution module;
[0020] Step 4: Use the encoder-decoder model combined with the multi-scale direction-adaptive snake convolution module trained in step 3 to extract roads from the test set divided in step 1.
[0021] The specific process of step 1 is:
[0022] Step 1.1: Acquire high-resolution remote sensing images;
[0023] Step 1.2: According to the ground feature classification standard, mark the roads in the remote sensing image obtained in step 1.1 to generate a label image;
[0024] Step 1.3: Crop the remote sensing images obtained in step 1.1 and the labeled images generated in step 1.2 to a size suitable for the network model, perform data cleaning, delete images with poor labeling quality, and construct a remote sensing road dataset;
[0025] Step 1.4: Divide the remote sensing road dataset constructed in step 1.3 into training set, validation set and test set.
[0026] The specific process of step 2 is:
[0027] Step 2.1: Build a direction-adaptive snake convolution module;
[0028] The constructed direction-adaptive snake convolution module consists of three stages; among them:
[0029] In the first stage, the offset coordinates are obtained through three branches according to the input features:
[0030] Branch 1: The input features are sequentially processed through depthwise separable convolution, layer normalization, and multiplication by a scaling factor to obtain the rotation offset, i.e., the radian value θ of the rotation offset angle. The calculation formula for θ is as follows:
[0031] θ = α × LN(Conv2d(f))
[0032] Among them, α is the scaling factor, f is the input feature, Conv2d means using convolution operation, and LN means using layer normalization;
[0033] Branch 2: Calculate the basic coordinates according to the size of the input feature and the size of the strip convolution kernel used; first initialize the center coordinates of the convolution kernel according to the input feature size, and then expand the dimension according to the number of channels of the input feature and the size of the strip convolution kernel to obtain the basic coordinates (x base ,y base );
[0034] Branch 3: The input features are sequentially processed through depthwise separable convolution, layer normalization, and activation functions to obtain the vertical offset Δy, which ranges from (-1, 1). The calculation formula for Δy is as follows:
[0035] Δy=Tanh(LN(Conv2d(f)))
[0036] Among them, Tanh means using the hyperbolic tangent activation function;
[0037] The basic coordinates (x base ,y base ) Combined with the radian value θ of the rotation offset angle obtained in branch 1, calculate the rotated coordinates (x, y):
[0038]
[0039] Where l is the coordinate initialized according to the strip convolution kernel size K, and the value of l is (-K / / 2, K / / 2), where K / / 2 means dividing the strip convolution kernel size K by 2 and rounding up.
[0040] Add the rotated ordinate y to the vertical offset Δy obtained in branch 3 to calculate the offset ordinate value y':
[0041] y′=y+Δy
[0042] Then the offset coordinates (x, y') are obtained from stage 1;
[0043] In the second stage, the input features are interpolated according to the offset coordinates obtained in the first stage to obtain the offset features, which are expressed as:
[0044] f'=bilinear interpolation(f,x,y')
[0045] Among them, bilinear interpolation is bilinear feature interpolation;
[0046] In the third stage, the offset features obtained in the second stage are convolved using depth-wise separable strip convolution to obtain the output features.
[0047] Step 2.2: Construct a multi-scale direction-adaptive snake convolution module
[0048] The input features are passed through three branches to extract road features from different scales. The first branch is a global average pooling module to extract overall information. The second branch uses the direction-adaptive snake convolution module constructed in step 2.1. The third branch cascades the two direction-adaptive snake convolution modules constructed in step 2.1 to increase the range of the receptive field. Finally, the features extracted by the three branches are added together to obtain the output feature f that integrates information at different scales. o , the calculation formula is as follows:
[0049]
[0050] f o =f o1 +f o2 +f o3
[0051] Among them, f i Represents the input features, AvgPool represents the global average pooling module, DAS (Direction-Adaptive Snake Convolution) represents the direction-adaptive snake convolution module, f oj represents the output feature of the j-th branch, f o Represents the output features after the input features pass through the multi-scale direction adaptive snake convolution module;
[0052] Step 2.3: Build the encoder module;
[0053] The constructed encoder module uses the ConvNeXt-Tiny version model as the basic skeleton of the encoder module, and connects four encoder stages in series. In encoder Stage 1, the convolution (conv2d), layer normalization (LayerNorm) layer and ConvNeXt Block layer are connected in series, and the multi-scale direction adaptive snake convolution module constructed in step 2.2 is connected in series after the ConvNeXt Block layer; in encoder Stage 2, encoder Stage 3 and encoder Stage 4, the downsampling layer (Downsample) and ConvNeXt Block layer are connected in series, and the multi-scale direction adaptive snake convolution module constructed in step 2.2 is connected in series after the ConvNeXt Block layer;
[0054] The input image is input into the encoder module and passes through 4 encoder stages in sequence. Each encoder stage takes the output features of the previous encoder stage as input and outputs the processed output features to the next encoder stage. Finally, the output features of 4 encoder stages are obtained.
[0055] Step 2.4: Build the decoder module;
[0056] Using the LinkNet decoder, the decoder module includes four stages of decoder layer modules. The output features of each stage of the decoder layer module are obtained through three layers in sequence. The first layer uses convolution, batch normalization and Relu activation function; the second layer uses transposed convolution, batch normalization and Relu activation function for upsampling; the third layer uses convolution, batch normalization and Relu activation function; the features output by encoder Stage 1, encoder Stage 2 and encoder Stage 3 in step 2.3 are input into the corresponding stages of decoder Stage 4, decoder Stage 3 and decoder Stage 2 constructed in step 2.4 using skip connections;
[0057] The input features of decoder Stage 1 are the output features of encoder Stage 4 in step 2.3; the input features of decoder Stage 2 are the features obtained by adding the output features of decoder Stage 1 and the output features of encoder Stage 3 in step 2.3; the input features of decoder Stage 3 are the features obtained by adding the output features of decoder Stage 2 and the output features of encoder Stage 2 in step 2.3; the input features of decoder Stage 4 are the features obtained by adding the output features of decoder Stage 3 and the output features of encoder Stage 1 in step 2.3; the output features of decoder Stage 4 are the final output features of the decoder module;
[0058] Step 2.5: Build the segmentation head module;
[0059] The segmentation head module consists of three layers. The first layer uses transposed convolution, batch normalization, and Relu activation function for upsampling; the second layer uses convolution, batch normalization, and Relu activation function; and the third layer uses convolution plus threshold prediction. The final output features of the decoder module constructed in step 2.4 are used as input to the segmentation head module, and the road segmentation result image is obtained through the segmentation head module.
[0060] Finally, the construction of the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module is completed.
[0061] The present invention also provides a high-resolution remote sensing image road adaptive extraction system based on multi-scale dynamic convolution enhancement, comprising:
[0062] The dataset construction and division module is used to construct a high-resolution remote sensing image road dataset, including original images and label data; and divide it into training set, validation set and test set;
[0063] A model building module for building an encoder-decoder model combined with a multi-scale direction-adaptive snake convolution module;
[0064] A model training module is used to train the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module using the training set and the validation set, and save the trained encoder-decoder model combined with the multi-scale direction adaptive snake convolution module;
[0065] The road extraction module is used to extract roads from the test set using the trained encoder-decoder model combined with the multi-scale direction-adaptive snake convolution module.
[0066] The present invention also provides a high-resolution remote sensing image road adaptive extraction device based on multi-scale dynamic convolution enhancement, comprising:
[0067] Memory: a computer-readable device storing a computer program for the above-mentioned method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement;
[0068] Processor: used to implement the method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement when executing the computer program.
[0069] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the high-resolution remote sensing image road adaptive extraction method based on multi-scale dynamic convolution enhancement.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] 1. The present invention designs a multi-scale direction-adaptive serpentine convolution module, which can not only adaptively adjust the rotation angle and offset of the strip convolution according to the actual morphological characteristics of the road in the remote sensing image, but also get rid of the application limitations of the traditional square convolution kernel that is not suitable for road morphology and the fixed direction of strip convolution, and significantly improve the model's adaptability to complex road morphology in high-resolution remote sensing images; and through the multi-scale feature fusion strategy, it realizes the multi-scale capture of road contextual semantic information in complex scenes, effectively enhancing the model's perception of the spatial continuity characteristics of roads in occluded environments.
[0072] 2. This paper designs an encoder-decoder model that incorporates a multi-scale, directionally adaptive snake convolution module. By cascading these modules in each encoder stage, the model achieves hierarchical extraction and fusion of road features at different scales. This enhances the ability to model the association between local features and global structure and to extract road features.
[0073] In summary, the present invention not only achieves accurate modeling of the topological structure level of complex road morphology through the designed multi-scale direction adaptive serpentine convolution module and the encoder-decoder model architecture combined with the multi-scale direction adaptive serpentine convolution module, but also deeply mines the road space semantic information through the multi-scale feature fusion method, effectively bridging the road break gaps caused by occlusion and shadows in complex scenes. In practical applications, this method not only significantly improves the integrity of road extraction from high-resolution remote sensing images and ensures the stable output of road connectivity, but also greatly improves the extraction accuracy by optimizing feature expression capabilities. Experimental verification shows that its road extraction performance in complex scenes is greatly improved compared to existing technologies, providing a new solution for the field of remote sensing image road extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is a flow chart of the method of the present invention.
[0075] Figure 2 This is a diagram of the direction-adaptive serpentine convolution module of the method of the present invention.
[0076] Figure 3 This is a diagram of the multi-scale directional adaptive snake convolution module of the method of the present invention.
[0077] Figure 4 It is the encoder module of the method of the present invention.
[0078] Figure 5 It is the decoder module of the method of the present invention.
[0079] Figure 6 It is the segmentation head module of the method of the present invention.
[0080] Figure 7 It is the overall framework diagram of the model of the method of the present invention.
[0081] Figure 8 2 is a test comparison result diagram of the method of the present invention. DETAILED DESCRIPTION
[0082] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0083] A method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement includes the following steps:
[0084] Step 1: Construct a high-resolution remote sensing image road dataset, including original images and manually annotated label data; and divide it into training set, validation set, and test set. The specific process is as follows:
[0085] Step 1.1: Obtain 0.3m high-resolution remote sensing images of northern Shaanxi, Guanzhong, southern Shaanxi, and parts of Hangzhou from satellite platforms or mapping software platforms such as Bigemap software;
[0086] Step 1.2: According to the ground feature classification standard, the roads in the high-resolution remote sensing image obtained in step 1.1 are annotated to generate a labeled image;
[0087] Step 1.3: Crop the high-resolution remote sensing images obtained in step 1.1 and the labeled images generated in step 1.2 into 512 × 512 pixel images. Perform data cleaning and delete images with poor labeling quality to construct a high-resolution remote sensing road dataset with a total of 5120 images.
[0088] Step 1.4: Divide the high-resolution remote sensing road dataset constructed in step 1.3 into training set, validation set, and test set in a ratio of 8:1:1.
[0089] Step 2: Build an encoder-decoder model that combines a multi-scale direction-adaptive snake convolution module. The specific process is as follows:
[0090] Step 2.1: Build a direction-adaptive snake convolution module;
[0091] like Figure 2As shown in the figure, building a direction-adaptive snake convolution module includes three stages. In the figure, (B, C, N, N) represent the batch size of the input features, Channels (number of channels), (N, N) are the pixel sizes of the input features, K represents the kernel size of the strip convolution kernel used; LayerNorm represents the use of layer normalization, and Tanh represents the use of hyperbolic tangent activation function.
[0092] In the first stage, the offset coordinates are obtained through three branches according to the input features:
[0093] Branch 1: The input features are sequentially processed through a 7×7 depthwise separable convolution, layer normalization, and multiplication by a scaling factor to obtain the rotation offset, i.e., the radian value θ of the rotation offset angle. The calculation formula is as follows:
[0094] θ = α × LN(Conv2d(f))
[0095] Among them, α is the scaling factor, f is the input feature, Conv2d means using convolution operation, and LN means using layer normalization;
[0096] The scaling factor α of the present invention is set to When calculating the radian value θ of the rotation offset angle based on the input feature f, after the input feature f undergoes depthwise separable convolution and LayerNorm (layer normalization), the output value obeys a normal distribution with a mean of 0 and a standard deviation (σ) of 1. According to the "3σ" principle in the normal distribution, there is a 95% probability that its value will be in the range of (-2σ, 2σ) or (-2, 2). Therefore, multiplying by the scaling factor After that, the radian value of the rotation offset angle θ has a 95% probability of being within the range Its value range can cover 360° and can flexibly adapt to the angle of the road in the remote sensing image.
[0097] Branch 2: Calculate the basic coordinates according to the size of the input feature and the size of the K=7 strip convolution kernel. First, initialize the center coordinates of the convolution kernel according to the input feature size, and then expand the dimension according to the number of channels of the input feature and the size of the strip convolution kernel to obtain the basic coordinates (x base ,y base );
[0098] Branch 3: The input features are sequentially passed through a 3×3 depthwise separable convolution, layer normalization, and a hyperbolic tangent activation function to obtain a vertical offset Δy, which ranges from (-1, 1). The calculation formula for Δy is as follows:
[0099] Δy=Tanh(LN(Conv2d(f)))
[0100] Among them, Tanh means using the hyperbolic tangent activation function;
[0101] The basic coordinates (x base ,y base ) Combined with the radian value θ of the rotation offset angle obtained in branch 1, calculate the rotated coordinates (x, y):
[0102]
[0103] Among them, l is the coordinate initialized according to the strip convolution kernel size K, and the value of l is (-K / / 2, K / / 2), where K / / 2 means dividing the strip convolution kernel size K by 2 and rounding it up; since the radian value of the rotation offset angle θ has a 95% probability of ranging from It can cover a 360° angle, so the strip convolution with the center point as the rotation center can adapt to roads at different angles in the image.
[0104] Add the rotated ordinate y to the vertical offset Δy obtained in branch 3 to calculate the offset ordinate value y':
[0105] y′=y+Δy
[0106] Then the offset coordinates (x, y') are obtained from stage 1;
[0107] In the second stage, the input features are interpolated according to the offset coordinates obtained in the first stage, and the offset features are expressed as:
[0108] f'=bilinear interpolation(f,x,y')
[0109] Among them, bilinear interpolation is bilinear feature interpolation;
[0110] In the third stage, a 7×1 depthwise separable strip convolution is used to convolve the offset features obtained in the second stage to obtain the output features. The output features are the features extracted after enhancement based on the road shape characteristics.
[0111] The constructed direction-adaptive snake convolution module uses depthwise separable convolution to reduce the number of parameters in both the convolution operation used in stage 1 to calculate the offset coordinates and the strip convolution operation in stage 3. This module can more simply and flexibly adapt to complex road angles and shapes in images, thereby better extracting road features.
[0112] Step 2.2: Construct a multi-scale direction-adaptive snake convolution module
[0113] like Figure 3As shown in the figure, the input features are extracted from different scales through three branches. The first branch is a global average pooling module to extract overall information; the second branch uses the direction-adaptive snake convolution module constructed in step 2.1; the third branch cascades the two direction-adaptive snake convolution modules constructed in step 2.1 to increase the range of the receptive field. Finally, the features extracted by the three branches are added together to obtain the output feature f that integrates information of different scales. o , the calculation formula is as follows:
[0114]
[0115] f o =f o1 +f o2 +f o3
[0116] Among them, f i Represents the input features, AvgPool represents the global average pooling module, DAS (Direction-Adaptive Snake Convolution) represents the direction-adaptive snake convolution module, f oj represents the output feature of the j-th branch, f o Represents the output features after the input features pass through the multi-scale direction adaptive snake convolution module;
[0117] The obtained output features model road context information from multiple scales, enhancing the ability of road feature extraction.
[0118] Step 2.3: Build the encoder module;
[0119] like Figure 4 As shown in the figure, the constructed encoder module uses the ConvNeXt-Tiny version model as the basic skeleton of the encoder module, and connects 4 encoder stages (Stage) in series. In encoder Stage1, the convolution (conv2d), layer normalization (LayerNorm) layer and ConvNeXt Block layer are connected in series, and the multi-scale direction adaptive snake convolution module constructed in step 2.2 is connected in series after the ConvNeXt Block layer; in encoder Stage2, encoder Stage3 and encoder Stage 4, the downsampling layer (Downsample) and ConvNeXt Block layer are connected in series in series, and the multi-scale direction adaptive snake convolution module constructed in step 2.2 is connected in series after the ConvNeXt Block layer; the number of ConvNeXt Block layers used in the 4 decoder stages is [3,3,9,3].
[0120] The input image is input into the encoder module and passes through 4 encoder stages (Stage) in sequence. The input image first passes through the convolution (conv2d), layer normalization (LayerNorm) layer and ConvNeXt Block layer in encoder Stage1 to obtain the output feature. The output feature is used as the input feature of the multi-scale direction adaptive snake convolution module constructed in step 2.2 in series. The output feature obtained by passing through the multi-scale direction adaptive snake convolution module is used as the output feature of encoder Stage1; in encoder Stage2, encoder Stage3 and encoder Stage4, the output feature of the previous encoder stage (Stage) is used as the input feature, and passes through the downsample layer (Downsample) and ConvNeXt Block layer in series. The output feature obtained is used as the input feature of the multi-scale direction adaptive snake convolution module constructed in step 2.2 in series, and then the output feature obtained by passing through the multi-scale direction adaptive snake convolution module is used as the output feature of encoder Stage2, encoder Stage3 and encoder Stage4.
[0121] Step 2.4: Build the decoder module;
[0122] like Figure 5 As shown in the figure, the decoder of LinkNet is used. The decoder module includes four stages of decoder layer modules. The output features are obtained through three layers in sequence in the decoder layer module of each stage. The first layer uses 1×1 convolution, batch normalization and Relu activation function to fuse features and reduce the number of channels of input features; the second layer uses 3×3 transposed convolution, batch normalization and Relu activation function for upsampling; the third layer uses 1×1 convolution, batch normalization and Relu activation function to change the number of channels of output features to half of the number of input channels.
[0123] The output features of encoder Stage 1, encoder Stage 2, and encoder Stage 3 in step 2.3 are input into the corresponding stages of decoder Stage 4, decoder Stage 3, and decoder Stage 2 constructed in step 2.4 using jump connections; the input features of decoder Stage 1 are the output features of encoder Stage 4 in step 2.3; the input features of decoder Stage 2 are the features obtained by adding the output features of decoder Stage 1 and the output features of encoder Stage 3 in step 2.3; the input features of decoder Stage 3 are the features obtained by adding the output features of decoder Stage 2 and the output features of encoder Stage 2 in step 2.3; the input features of decoder Stage 4 are the features obtained by adding the output features of decoder Stage 3 and the output features of encoder Stage 1 in step 2.3; the output features of decoder Stage 4 are the final output features of the decoder module.
[0124] Step 2.5: Build the segmentation head module;
[0125] like Figure 6 As shown in , it consists of three layers. The first layer uses 3×3 transposed convolution, batch normalization, and Relu activation function for upsampling; the second layer uses 3×3 convolution, batch normalization, and Relu activation function; the third layer uses 1×1 convolution to obtain the output value and perform threshold prediction. The prediction value greater than 0.5 is road, and the rest is background. The final output feature of the decoder module constructed in step 2.4 is used as the input of the segmentation head module, and the road segmentation result image is obtained after the segmentation head module. The construction of the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module is completed, as shown in Figure 7 shown.
[0126] Step 3: Input the training set data divided in step 1 into the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module constructed in step 2 for training. The validation set divided in step 1 is used to assist in detecting the training effect of the model.
[0127] During training, the pre-trained weights of the ConvNeXtTiny basic skeleton on ImageNet-1K are loaded, the batch size is set to 16, the initial learning rate is set to 3e-4, the learning rate strategy of WarmUp plus Poly is adopted, the optimizer is Adam, the training rounds are set to 200, and the Loss is used. BCE As the loss function, its calculation formula is as follows:
[0128]
[0129] Where N is the number of image pixels, g irepresents the value of the i-th pixel label, p i Represents the predicted probability value of the corresponding pixel.
[0130] When the loss drops to a stable convergence and the IoU on the validation set no longer increases, the model training is completed and the trained model is saved. IoU represents the ratio of the intersection and union of the prediction result and the label, and its calculation formula is as follows:
[0131]
[0132] Among them, TP (True Positives) represents the number of road pixels correctly predicted as road categories, FP (False Positives) represents the number of background pixels incorrectly predicted as road categories, and FN (False Negatives) represents the number of road pixels incorrectly predicted as background categories.
[0133] Step 4: Use the encoder-decoder model combined with the multi-scale direction-adaptive snake convolution module trained in step 3 to extract roads from the test set divided in step 1 and obtain the road extraction result map.
[0134] Experimental analysis
[0135] 1) Simulation experiment environment:
[0136] The hardware platform of the experimental environment of the method of the present invention is Huawei's ModelArts platform with 8 cards Ascend910, and the software platform is Huawei EulerOS 2.0SP8 and Pytorch 1.11.0.
[0137] 2) Experimental content
[0138] The simulation experiments were conducted using 0.3m high-resolution, large-scale optical remote sensing images from northern Shaanxi, Guanzhong, southern Shaanxi, and parts of Hangzhou. The dataset was cropped using a 512×512 non-overlapping format, resulting in 5120 image blocks. 512 of these blocks, 10%, were randomly selected as a test set to verify model performance. 4096 of these blocks were randomly selected as a training set, and 512 were used as a validation set for model training.
[0139] The network models were trained using the method of the present invention and a baseline model in which the multi-scale direction-adaptive snake convolution module was deleted from the encoder-decoder model combined with the multi-scale direction-adaptive snake convolution module of the present invention, respectively, under the ImageNet-1K pre-training weights, and the road extraction performance of the models on the test set was compared.
[0140] 3) Experimental results:
[0141] The present invention uses four evaluation metrics to measure the accuracy of the road extraction results of each model, including Intersection over Union (IoU), Precision, Recall, and F1 value. Among them, IoU represents the ratio of the intersection and union of the prediction result and the label, Precision represents the proportion of correctly predicted pixels among the pixels predicted as roads, Recall represents the proportion of pixels correctly predicted as roads among all road pixels, and F1 value is the harmonic mean of Precision and Recall. The calculation formulas of the above four metrics are as follows:
[0142]
[0143] Among them, TP (True Positives) represents the number of road pixels correctly predicted as road categories, FP (False Positives) represents the number of background pixels incorrectly predicted as road categories, and FN (False Negatives) represents the number of road pixels incorrectly predicted as background categories.
[0144] The trained encoder-decoder model of the present invention combined with the multi-scale direction adaptive snake convolution module and the baseline model are used to extract roads from the test set divided in step 1 and the extraction results are compared. The obtained evaluation indicators are shown in Table 1.
[0145] Table 1 Comparison of various evaluation indicators between the proposed method and the benchmark model in the same data set
[0146] IoU (%) Precision (%) Recall (%) F1(%) Benchmark 64.6 78.7 78.2 78.5 The present invention 66.4 79.6 80.0 79.8
[0147] It can be seen from Table 1 that the method of the present invention effectively improves the accuracy of road extraction.
[0148] Some of the comparison results obtained from the test are as follows Figure 8 As shown in the figure, the comparison results show that in complex scenes with obvious buildings, trees, or shadows, the proposed method achieves better road connectivity compared to the baseline model. Experimental comparison results demonstrate that the proposed method has superior feature extraction and road context modeling capabilities for roads in high-resolution remote sensing imagery. The proposed method enhances road connectivity in complex scenes and improves road extraction accuracy.
[0149] The key points and protection points of the present invention are:
[0150] 1) This paper designs a multi-scale direction-adaptive serpentine convolution module, which can dynamically adjust the rotation angle and offset of the strip convolution according to the actual morphological characteristics of the roads in remote sensing images, accurately adapt to the multi-scale geometric morphology of different roads, and realize multi-level capture of road contextual semantic information in complex scenes, enhancing the ability to capture local details of road features and understand the global structure.
[0151] 2) This paper designs an encoder-decoder model that incorporates a multi-scale, directionally adaptive snake convolution module. By cascading these modules at each encoder stage, the model achieves layered extraction and deep fusion of road features at different scales, strengthening the modeling of the association between local features and global structure. This model effectively enhances the spatial connectivity of road extraction in complex scenarios, significantly improving the accuracy and completeness of road extraction, and provides a new technical approach and model architecture for road extraction from high-resolution remote sensing imagery.
[0152] The present invention also provides a high-resolution remote sensing image road adaptive extraction system based on multi-scale dynamic convolution enhancement, comprising:
[0153] The dataset construction and division module is used to implement the construction of a high-resolution remote sensing image road dataset in step 1, including original images and label data; and divide it into a training set, a validation set, and a test set;
[0154] A model construction module is used to implement the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module in step 2;
[0155] The model training module is used to implement the training set and validation set divided in step 1 in step 3 to train the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module constructed in step 2, and save the trained encoder-decoder model combined with the multi-scale direction adaptive snake convolution module;
[0156] The road extraction module is used to implement the road extraction in step 4 using the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module trained in step 3 on the test set divided in step 1.
[0157] The present invention also provides a high-resolution remote sensing image road adaptive extraction device based on multi-scale dynamic convolution enhancement, comprising:
[0158] Memory: a computer-readable device storing a computer program for the above-mentioned method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement;
[0159] Processor: used to implement the method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement when executing the computer program.
[0160] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the high-resolution remote sensing image road adaptive extraction method based on multi-scale dynamic convolution enhancement.
Claims
1. A method for adaptive road extraction from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement, characterized in that: The following steps are involved: Step 1: Construct a high-resolution remote sensing image road dataset, including original images and label data; and divide it into training set, validation set, and test set; Step 2: Build an encoder-decoder model combined with a multi-scale direction-adaptive snake convolution module; Step 3: Use the training set and validation set divided in step 1 to train the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module constructed in step 2, and save the trained encoder-decoder model combined with the multi-scale direction adaptive snake convolution module; Step 4: Use the encoder-decoder model combined with the multi-scale direction-adaptive snake convolution module trained in step 3 to extract roads from the test set divided in step 1.
2. The method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement according to claim 1, characterized in that: The specific process of step 1 is: Step 1.1: Acquire high-resolution remote sensing images; Step 1.2: According to the ground feature classification standard, mark the roads in the remote sensing image obtained in step 1.1 to generate a label image; Step 1.3: Crop the remote sensing images obtained in step 1.1 and the labeled images generated in step 1.2 to a size suitable for the network model, perform data cleaning, delete images with poor labeling quality, and construct a remote sensing road dataset; Step 1.4: Divide the remote sensing road dataset constructed in step 1.3 into training set, validation set and test set.
3. The method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement according to claim 1, characterized in that: The specific process of step 2 is: Step 2.1: Build a direction-adaptive snake convolution module; The constructed direction-adaptive snake convolution module consists of three stages; among them: In the first stage, the offset coordinates are obtained through three branches according to the input features: Branch 1: The input features are sequentially processed through depthwise separable convolution, layer normalization, and multiplication by a scaling factor to obtain the rotation offset, i.e., the radian value θ of the rotation offset angle. The calculation formula for θ is as follows: θ = α × LN(Conv2d(f)) Among them, α is the scaling factor, f is the input feature, Conv2d means using convolution operation, and LN means using layer normalization; Branch 2: Calculate the basic coordinates according to the size of the input feature and the size of the strip convolution kernel used; first initialize the center coordinates of the convolution kernel according to the input feature size, and then expand the dimension according to the number of channels of the input feature and the size of the strip convolution kernel to obtain the basic coordinates (x base ,y base ); Branch 3: The input features are sequentially processed through depthwise separable convolution, layer normalization, and activation functions to obtain the vertical offset Δy, which ranges from (-1, 1). The calculation formula for Δy is as follows: Δy=Tanh(LN(Conv2d(f))) Among them, Tanh means using the hyperbolic tangent activation function; The basic coordinates (x base ,y base ) Combined with the radian value θ of the rotation offset angle obtained in branch 1, calculate the rotated coordinates (x, y): Where l is the coordinate initialized according to the strip convolution kernel size K, and the value of l is (-K / / 2, K / / 2), where K / / 2 means dividing the strip convolution kernel size K by 2 and rounding up. Add the rotated ordinate y to the vertical offset Δy obtained in branch 3 to calculate the offset ordinate value y': y′=y+Δy Then the offset coordinates (x, y') are obtained from stage 1; In the second stage, the input features are interpolated according to the offset coordinates obtained in the first stage to obtain the offset features, which are expressed as: f'=bilinear interpolation(f,x,y') Among them, bilinear interpolation is bilinear feature interpolation; In the third stage, the offset features obtained in the second stage are convolved using depth-wise separable strip convolution to obtain the output features. Step 2.2: Construct a multi-scale direction-adaptive snake convolution module The input features are passed through three branches to extract road features from different scales. The first branch is a global average pooling module to extract overall information. The second branch uses the direction-adaptive snake convolution module constructed in step 2.
1. The third branch cascades the two direction-adaptive snake convolution modules constructed in step 2.1 to increase the range of the receptive field. Finally, the features extracted by the three branches are added together to obtain the output feature f that integrates information at different scales. o , the calculation formula is as follows: f o =f o1 +f o2 +f o3 Among them, f i Represents the input features, AvgPool represents the global average pooling module, DAS (Direction-AdaptiveSnake Convolution) represents the direction-adaptive snake convolution module, f oj represents the output feature of the j-th branch, f o Represents the output features after the input features pass through the multi-scale direction adaptive snake convolution module; Step 2.3: Build the encoder module; The constructed encoder module uses the ConvNeXt-Tiny version model as the basic skeleton of the encoder module, and connects four encoder stages in series. In encoder Stage 1, the convolution (conv2d), layer normalization (LayerNorm) layer and ConvNeXtBlock layer are connected in series, and the multi-scale direction adaptive snake convolution module constructed in step 2.2 is connected in series after the ConvNeXtBlock layer; in encoder Stage 2, encoder Stage 3 and encoder Stage 4, the downsampling layer (Downsample) and ConvNeXtBlock layer are connected in series, and the multi-scale direction adaptive snake convolution module constructed in step 2.2 is connected in series after the ConvNeXtBlock layer; The input image is input into the encoder module and passes through 4 encoder stages in sequence. Each encoder stage takes the output features of the previous encoder stage as input and outputs the processed output features to the next encoder stage. Finally, the output features of 4 encoder stages are obtained. Step 2.4: Build the decoder module; Using the LinkNet decoder, the decoder module includes four stages of decoder layer modules. The output features of each stage of the decoder layer module are obtained through three layers in sequence. The first layer uses convolution, batch normalization and Relu activation function; the second layer uses transposed convolution, batch normalization and Relu activation function for upsampling; the third layer uses convolution, batch normalization and Relu activation function; the features output by encoder Stage 1, encoder Stage 2 and encoder Stage 3 in step 2.3 are input into the corresponding stages of decoder Stage 4, decoder Stage 3 and decoder Stage 2 constructed in step 2.4 using skip connections; The input features of decoder Stage 1 are the output features of encoder Stage 4 in step 2.3; the input features of decoder Stage 2 are the features obtained by adding the output features of decoder Stage 1 and the output features of encoder Stage 3 in step 2.3; the input features of decoder Stage 3 are the features obtained by adding the output features of decoder Stage 2 and the output features of encoder Stage 2 in step 2.3; the input features of decoder Stage 4 are the features obtained by adding the output features of decoder Stage 3 and the output features of encoder Stage 1 in step 2.3; the output features of decoder Stage 4 are the final output features of the decoder module; Step 2.5: Build the segmentation head module; The segmentation head module consists of three layers. The first layer uses transposed convolution, batch normalization, and Relu activation function for upsampling; the second layer uses convolution, batch normalization, and Relu activation function; and the third layer uses convolution plus threshold prediction. The final output features of the decoder module constructed in step 2.4 are used as input to the segmentation head module, and the road segmentation result image is obtained through the segmentation head module. Finally, the construction of the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module is completed.
4. A high-resolution remote sensing image road adaptive extraction system based on multi-scale dynamic convolution enhancement based on the method according to any one of claims 1 to 3, characterized in that: include: The dataset construction and partitioning module is used to construct a high-resolution remote sensing image road dataset, including original images and label data; And divided into training set, validation set and test set; A model building module for building an encoder-decoder model combined with a multi-scale direction-adaptive snake convolution module; A model training module is used to train the encoder-decoder model combined with the multi-scale direction adaptive snake convolution module using the training set and the validation set, and save the trained encoder-decoder model combined with the multi-scale direction adaptive snake convolution module; The road extraction module is used to extract roads from the test set using the trained encoder-decoder model combined with the multi-scale direction-adaptive snake convolution module.
5. A high-resolution remote sensing image road adaptive extraction device based on multi-scale dynamic convolution enhancement, characterized in that: include: Memory: a computer-readable device storing a computer program for a method for adaptively extracting roads from high-resolution remote sensing images based on multi-scale dynamic convolution enhancement as described in any one of claims 1 to 3; Processor: used to implement the high-resolution remote sensing image road adaptive extraction method based on multi-scale dynamic convolution enhancement as described in any one of claims 1-3 when executing the computer program.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, can implement a high-resolution remote sensing image road adaptive extraction method based on multi-scale dynamic convolution enhancement as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Remote sensing image road extraction method and device based on convolutional network
CN116543304A
Convolutional neural network remote sensing image road extraction model and method
CN116778318A