Remote sensing image road extraction method based on CNN (Convolutional Neural Network) and Swin Transform double encoders
By adopting the DPMSRE-Net model based on CNN and Swin Transformer dual encoder in remote sensing image processing, the problem of low road extraction accuracy in complex backgrounds is solved, and more efficient and accurate road feature extraction is achieved.
Patent Information
- Application Number
- CN202510061509.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
AI Technical Summary
In the complex context, it is difficult for the prior art to achieve high-precision and efficient road extraction, especially when local features and global information are inconsistent.
The remote sensing image road extraction model DPMSRE-Net based on CNN and Swin Transformer dual encoder is adopted. Through the dual-path encoder, the fusion module MSAF of multi-scale bar convolution attention and the multi-scale cross-direction attention module MSCA, the deep fusion of local features and global information and the enhancement of context information are achieved.
It improves the accuracy and accuracy of road feature extraction, and can more effectively deal with road extraction tasks in complex contexts.
Smart Images

Figure CN119963853A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of remote sensing image processing and relates to a remote sensing image road extraction method based on CNN and Swin Transformer dual encoders. Background Art
[0002] Roads are an important part of geographic information systems. Road information extraction is crucial for many fields such as urban planning, autonomous driving, emergency rescue, residential area planning and road monitoring. It can provide accurate and practical road information for decision makers, researchers and the public, and help improve urban efficiency, safety and sustainable development. The early road feature recognition and manual road contour extraction methods that relied on manual visual interpretation required a lot of manpower and time, and were affected by subjective factors, resulting in low efficiency.
[0003] In recent years, with the development of remote sensing technology and the popularization of high-resolution satellite images, traditional road extraction methods can no longer meet the needs of high efficiency and accuracy. Deep learning has gradually become the mainstream method for road extraction research due to its excellent performance in image processing. Deep learning originated from artificial neural networks. By increasing the depth and number of layers of the network, the model can automatically extract and learn complex features in the data. Since AlexNet proposed by Krizhevsky et al. emerged in the ImageNet competition in 2012, deep learning has made breakthrough progress in the field of computer vision. The remote sensing extraction of linear features based on deep learning mainly involves the improvement of the three basic models of convolutional networks, UNet and Transformer to improve the performance of linear feature extraction. The improvement of the convolutional network model mainly involves multi-source data fusion, loss function design and other aspects. The road extraction method based on the fully convolutional network (FCN) is specifically used to realize the automatic extraction of road features from images or remote sensing images. However, rough extraction effects still appear in complex backgrounds.
[0004] The improvement of the UNet model mainly focuses on improving the model accuracy from the aspects of reducing parameters, expanding the receptive field, capturing contextual information, improving multi-scale features and detail capture capabilities. The network based on the UNet model encoder-decoder can learn features of different scales through skip connections and multi-layer structures, thereby achieving better extraction effects. Although many studies have adopted the UNet-based road extraction model, fragmentation is still the main challenge faced by multi-scale road extraction in complex backgrounds.
[0005] Transformer was originally applied in the field of natural language processing. With its self-attention mechanism, it performs well in processing long-distance dependencies and extracting global information. In recent years, Transformer has been introduced into the field of computer vision and has shown great potential in tasks such as image classification and object detection. In the road extraction task, Transformer also shows its unique advantages. Although the road extraction problem is solved by introducing Transformer, the inconsistency between local features and global information is not taken into account. A deeper combination needs to be considered in road feature extraction. Summary of the invention
[0006] In view of this, the purpose of the present invention is to provide a remote sensing image road extraction model DPMSRE-Net (Dual-path and Muti-scale road extractionnetwork based on CNN and Swin Transformer, DPMSRE-Net) based on CNN and Swin Transformer dual encoders to extract roads from remote sensing images.
[0007] In order to achieve the above object, the present invention provides the following technical solutions:
[0008] A remote sensing image road extraction method based on CNN and Swin Transformer dual encoders includes the following steps:
[0009] S1: Obtain relevant road image datasets, divide them into training set, validation set and test set after preprocessing;
[0010] S2: Construct a remote sensing image road extraction model DPMSRE-Net based on CNN and Swin Transformer dual encoders. The model includes a dual-path encoder, a fusion module based on multi-scale strip convolution attention MSAF (Multi-scale Strip Attention Fusion) and a multi-scale cross-direction attention module MSCA (Multi-scale Cross-direction Attention);
[0011] S3: Train the model, calculate the loss of real images and predicted images by joint cross entropy loss function and DICE loss function, reduce the loss by back propagation, and adjust the learning strategy using AdamW optimizer; use validation set and test set for verification and testing;
[0012] S4: Use the trained DPMSRE-Net for road extraction.
[0013] Furthermore, the preprocessing in step S1 includes cropping the image, assigning the corresponding label background pixel value to 0, and the road pixel value to 1; and performing data enhancement operations on the training set and the validation set data.
[0014] Furthermore, the dual-path encoder of the DPMSRE-Net is a four-layer encoder consisting of a CNN local feature extraction path and a SwinTransformer global feature extraction path.
[0015] Furthermore, the fusion module MSAF based on multi-scale strip convolution attention of the DPMSRE-Net realizes the deep fusion of CNN and Transformer features through multi-scale strip convolution kernel and channel fusion, specifically including:
[0016] Adjust the shape and channel of the features from the Transformer path and concatenate them with the features from the CNN path in the channel dimension;
[0017] The strip convolution kernel is used to adapt the directional characteristics of the road. After 1x1 convolution, the feature information is sent to the convolution layer with a convolution kernel size of 5x5 to learn preliminary features. Then, multi-scale features are extracted through three different strip convolution kernel paths, and then channel fusion is performed on these features to mix information of different scales:
[0018] F0=Conv 1×1 (F concat ) (1)
[0019] F1=Conv 5×5 (F0) (2)
[0020] F2=Conv 7×1 (Conv 1×7 (F1)) (3)
[0021] F3=Conv 11×1 (Conv 1×11 (F1)) (4)
[0022] F4=Conv 21×1 (Conv 1×21 (F1)) (5)
[0023]
[0024] F out =Conv 1×1 (F5)+F0 (7)
[0025] in, is element-wise multiplication, is element-wise addition.
[0026] After residual connection, the fused features are segmented on the channel;
[0027] Then, through the residual connection, it is fused with the feature information of the corresponding path and sent to the Swin Transformer and CNN encoder of the next layer respectively to continue feature extraction and processing.
[0028] Furthermore, the multi-scale cross-directional attention module MSCA of the DPMSRE-Net is in the central bridge part of the network, which enhances the context information transmission by cross-fusing multi-scale features in vertical and horizontal directions;
[0029] The features are then further processed through the cross-attention mechanism: for the features Fx and Fy from the horizontal and vertical directions, the features of Fx are used as the query matrix Q, and the features of Fy are used as the key matrix K and the value matrix V, and Wq, Wk and Wv are generated respectively through linear transformation:
[0030] Q=F x W q ,K=F y W k ,V=F y W v (8)
[0031] Among them, Wq, Wk, Wv are learnable weight matrices;
[0032] Finally, based on the Squeeze-and-Excitation Attention (SE) mechanism, the weight of each channel is adaptively adjusted.
[0033] Furthermore, in step S3, the loss function is calculated jointly using the cross entropy function and the Dice loss:
[0034]
[0035] L=L dice +L ce (11)
[0036] Among them, Y represents the true label, represents the predicted label, N represents the total number of pixels in the image, and y i ∈{0,1} represents the true label of pixel i, represents the probability that pixel i belongs to category 1.
[0037] Furthermore, during the training process, the AdamW optimizer is used to adjust the learning strategy, and the batch size is selected according to the GPU.
[0038] Further, in step S4, the test set is sent to the trained model, and in the generated image obtained, 0 represents the background and 1 represents the road, which is converted into RGB values, the 0 value is converted into 3-channel 255, representing black, and the 1 value is converted into 3-channel 0, representing white, so as to obtain a road extraction image.
[0039] The beneficial effect of the present invention is that the present invention takes into account the inconsistency between local features and global information, thereby improving the precision and accuracy of road feature extraction.
[0040] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0042] Figure 1 This is a flow chart of the road extraction method from remote sensing images based on CNN and Swin Transformer dual encoders;
[0043] Figure 2 This is the structure diagram of DPMSRE-Net;
[0044] Figure 3 It is the MSAF module diagram;
[0045] Figure 4 This is the MSCA module diagram. DETAILED DESCRIPTION
[0046] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0047] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0048] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0049] like Figure 1 As shown, the present invention provides a remote sensing image road extraction method based on CNN and Swin Transformer dual encoder, comprising the following steps:
[0050] First, download the optical remote sensing road image and its corresponding label image from the Internet, crop it into a 256*256 image by overlapping cropping, and assign the corresponding label background pixel value to 0 and the road pixel value to 1.
[0051] The preprocessed dataset is divided into training set, validation set and test set with an 8:1:1 ratio. During the training process, data enhancement operations such as horizontal flipping, vertical flipping, 10° rotation and random cropping are performed on the training set and validation set.
[0052] The DPMSRE-Net structure diagram is as follows Figure 2 As shown, the processed data is fed into DPMSRE-Net. DPMSRE-Net includes the model including dual-path encoder, MSAF, MSCA and decoder.
[0053] The dual-path encoder is a four-layer encoder consisting of a CNN local feature extraction path and a Transformer global feature extraction path. The CNN path encoder uses ResNet34 pre-trained on the ImageNet dataset. First, a convolution layer with a convolution kernel size of 7x7, a stride of 2, and a padding of 3 and a maximum pooling layer are used to perform primary extraction of image features. The shape of the features changes from 512x512x3 to 128x128x64. Then, four residual convolution blocks are followed, and the number of them is [3, 4, 6, 3]. This part uses the weights pre-trained on ImageNet, which can help the network speed up the training convergence without additional training costs. The core of the Swin Transformer path encoder is the Window-based Multi-Head Self-Attention (W-MSA) and Sliding Window Multi-Head Self-Attention (SW-MSA). The sliding window makes up for the limitations of the VIT (Vision Transformer) window self-attention mechanism and realizes the exchange of feature information across windows. The adopted Swin Transformer generates feature maps of different resolutions through 4 stages after Patch Partition and Linear Embedding, namely (128x128), (64x64), (32x32), and (16x16). Each stage contains two base layers, and each base layer is composed of W-MSA, MLP (multi-layer perceptron), SW-MSA, and MLP in sequence, interspersed with four LayerNorm layers.
[0054] The feature information initially extracted by the dual-path encoder of each layer is put into MSAF for deep fusion of local features and global features. Figure 3 As shown in the figure, the two features are first concatenated on the channel, and then the linear features of the adaptive road are extracted through 1*7, 1*11, and 1*21 strip convolutions. The feature information extracted at different scales is then added and fused, and the model's sensitivity to the extracted features is enhanced by multiplying the feature information before fusion.
[0055] F0=Conv 1×1 (F concat )
[0056] F1=Conv 5×5 (F0)
[0057] F2=Conv 7×1 (Conv 1×7(F1)
[0058] F3=Conv 11×1 (Conv 1×11 (F1)
[0059] F4=Conv 21×1 (Conv 1×21 (F1)
[0060]
[0061] F out =Conv 1×1 (F5)+F0
[0062] in, is element-wise multiplication, is element-wise addition.
[0063] After that, the feature is cut from the channel, and the cut features are combined with the original path features to add new features and prevent the degradation of the model. Finally, they are sent to the encoder of the next layer as input, and sent to the decoder of the corresponding layer through skip connection.
[0064] like Figure 4 As shown, after the encoder extraction, the core bridge module in the encoder-decoder model normalizes the feature information and sends it to the strip convolution groups in the vertical and horizontal directions respectively. Then, the multi-scale feature information in the two directions is added, and then normalized respectively, and the fusion features in the two directions are cross-fused. For the features Fx and Fy from the horizontal and vertical directions. The features of Fx are used as the query matrix Q, and the features of Fy are used as the key matrix K and the value matrix V. The features of Fx are integrated into Fy. Similarly, the features of Fy are used as the query matrix Q, and the features of Fx are used as the key matrix K and the value matrix V, and the features of Fy are integrated into Fx.
[0065] Q=F x W q ,K=F y W k ,V=F y W v
[0066] Among them, Wq, Wk, Wv are learnable weight matrices.
[0067] After average pooling, the learnable parameter w allows the model to spontaneously adjust the weights of features from different directions, thereby improving the effect of feature fusion. Finally, the squeeze-and-excitation attention mechanism (SE) is used to adaptively adjust the weight of each channel to enhance the model's attention to important features. This mechanism can effectively improve the model's feature expression ability, allowing it to focus more on features that are helpful for road extraction.
[0068] The fused features are restored to the original image size layer by layer through the decoder to obtain the predicted image. During the training process, since the effective road pixels always account for a small proportion of the total image, the loss function is calculated jointly by the cross entropy function and the Dice loss.
[0069]
[0070] L=L dice +L ce
[0071] Among them, Y represents the true label, represents the predicted label, N represents the total number of pixels in the image, and y i ∈{0,1} represents the true label of pixel i, represents the probability that pixel i belongs to category 1.
[0072] And by using the AdamW optimizer to adjust the learning strategy, the epoch is set to 100, and the batch size is selected according to the GPU.
[0073] Finally, the test set is sent to the trained model. In the generated image, 0 represents the background and 1 represents the road. It is converted into RGB values, and the 0 value is converted into 3-channel 255, representing black, and the 1 value is converted into 3-channel 0, representing white, so as to obtain the road extraction image.
[0074] In the above embodiments, the description's reference to "this embodiment" indicates that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.
[0075] In the above-described embodiments, although the invention has been described in conjunction with specific embodiments of the invention, many substitutions, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed. Embodiments of the invention are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of the appended claims.
[0076] This embodiment further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, any one of the methods in this embodiment is implemented.
[0077] This embodiment also provides an electronic terminal, including: a processor and a memory;
[0078] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes any one of the methods in this embodiment.
[0079] The computer-readable storage medium in this embodiment can be understood by ordinary technicians in this field: all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes.
[0080] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used to communicate, and the processor and the transceiver are used to run computer programs so that the electronic terminal executes each step of the above method.
[0081] In this embodiment, the memory may include a random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0082] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0083] The present invention can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.
[0084] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A method for extracting roads from remote sensing images based on CNN and Swin Transformer dual encoders, characterized in that: The following steps are involved: S1: Obtain relevant road image datasets, divide them into training set, validation set and test set after preprocessing; S2: Construct a remote sensing image road extraction model DPMSRE-Net based on CNN and Swin Transformer dual encoders. The model includes a dual-path encoder, a fusion module MSAF based on multi-scale strip convolution attention, and a multi-scale cross-direction attention module MSCA; S3: Train the model, calculate the loss of real images and predicted images by joint cross entropy loss function and DICE loss function, reduce the loss by back propagation, and adjust the learning strategy using AdamW optimizer; use validation set and test set for verification and testing; S4: Use the trained DPMSRE-Net for road extraction.
2. The method for extracting roads from remote sensing images based on CNN and Swin Transformer dual encoders according to claim 1, characterized in that: The preprocessing in step S1 includes cropping the image, assigning the corresponding label background pixel value to 0, and the road pixel value to 1; and performing data enhancement operations on the training set and the validation set data.
3. The method for extracting roads from remote sensing images based on CNN and Swin Transformer dual encoders according to claim 1, characterized in that: The dual-path encoder of the DPMSRE-Net is a four-layer encoder composed of a CNN local feature extraction path and a SwinTransformer global feature extraction path.
4. The method for extracting roads from remote sensing images based on CNN and Swin Transformer dual encoders according to claim 1, characterized in that: The fusion module MSAF based on multi-scale strip convolution attention of the DPMSRE-Net realizes the deep fusion of CNN and Transformer features through multi-scale strip convolution kernel and channel fusion, specifically including: Adjust the shape and channel of the features from the Transformer path and concatenate them with the features from the CNN path in the channel dimension; The strip convolution kernel is used to adapt the directional characteristics of the road. After 1x1 convolution, the feature information is sent to the convolution layer with a convolution kernel size of 5x5 to learn preliminary features. Then, multi-scale features are extracted through three different strip convolution kernel paths, and then channel fusion is performed on these features to mix information of different scales: F0=Conv 1×1 (F concat ) (1) F1=Conv 5×5 (F0) (2) F2=Conv 7×1 (Conv 1×7 (F1)) (3) F3=Conv 11×1 (Conv 1×11 (F1)) (4) F4=Conv 21×1 (Conv 1×21 (F1)) (5) F out =Conv 1×1 (F5)+F0 (7) in, is element-wise multiplication, is element-wise addition; After residual connection, the fused features are segmented on the channel; Then, through the residual connection, it is fused with the feature information of the corresponding path and sent to the SwinTransformer and CNN encoder of the next layer respectively to continue feature extraction and processing.
5. The method for extracting roads from remote sensing images based on CNN and Swin Transformer dual encoders according to claim 1, characterized in that: The multi-scale cross-directional attention module MSCA of the DPMSRE-Net is in the central bridge part of the network, which enhances the context information transmission by cross-fusion of multi-scale features in vertical and horizontal directions; The features are then further processed through the cross-attention mechanism: for the features Fx and Fy from the horizontal and vertical directions, the features of Fx are used as the query matrix Q, and the features of Fy are used as the key matrix K and the value matrix V, and Wq, Wk and Wv are generated respectively through linear transformation: Q=F x W q ,K=F y W k ,V=F y W v , (8) Among them, Wq, Wk, Wv are learnable weight matrices; Finally, based on the squeeze and excitation attention mechanism SE, the weight of each channel is adaptively adjusted.
6. The method for extracting roads from remote sensing images based on CNN and Swin Transformer dual encoders according to claim 1, characterized in that: In step S3, the loss function is calculated jointly by the cross entropy function and the Dice loss: L=L dice +L ce (11) Among them, Y represents the true label, represents the predicted label, N represents the total number of pixels in the image, and y i ∈{0,1} represents the true label of pixel i, represents the probability that pixel i belongs to category 1.
7. The method for extracting roads from remote sensing images based on CNN and Swin Transformer dual encoders according to claim 1, characterized in that: During the training process, the AdamW optimizer is used to adjust the learning strategy, and the batch size is selected according to the GPU.
8. The method for extracting roads from remote sensing images based on CNN and Swin Transformer dual encoders according to claim 1, characterized in that: In step S4, the test set is sent to the trained model, and in the generated image obtained, 0 represents the background and 1 represents the road. It is converted into RGB values, and the 0 value is converted into 3-channel 255, representing black, and the 1 value is converted into 3-channel 0, representing white, so as to obtain a road extraction image.
Citation Information
Cited By
CNN-RWKV fusion-based remote sensing image road extraction method and system
CN121053554A