A remote sensing image segmentation method and system based on improved SwinTransformer
By combining SwinTransformer backbone network and global information enhancement module, the problem of high global information acquisition and calculation complexity in remote sensing image segmentation is solved, and a more efficient remote sensing image segmentation effect is achieved.
Patent Information
- Application Number
- CN202311068314.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-23
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-08-23
AI Technical Summary
In the existing remote sensing image segmentation methods, CNN-based methods are difficult to obtain global context information, while Vision Transformers-based methods have high computational complexity, which limits their performance in remote sensing image segmentation tasks.
The SwinTransformer backbone network is combined with the global information enhancement module. Through global average pooling operations in horizontal and vertical directions, the global information mining capabilities of the SwinTransformer backbone network are made up for, and feature fusion is carried out through the decoder's attention mechanism to retain more local information and obtain broader context information.
It improves the accuracy of remote sensing image segmentation, reduces the computational complexity and parameter quantity, and can better capture the long-distance dependence between semantic features.
Smart Images

Figure CN117095168B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a remote sensing image segmentation method and system based on an improved SwinTransformer. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Remote sensing image segmentation can provide accurate classification and distribution information of land objects, and has broad practical significance in land use planning, environmental monitoring, urban planning, agriculture and forestry management, resource exploration, and infrastructure management. The task of remote sensing image segmentation is to classify each pixel and separate different land objects in the image.
[0004] Pure convolutional neural networks (CNNs) are widely used in remote sensing image segmentation tasks due to their outstanding ability to represent data at multiple scales and capture local semantic information. CNNs' powerful learning capabilities enable them to obtain more accurate and rich features than traditional methods for studying high-dimensional information, given sufficient training samples. They rely on an encoder-decoder architecture; the entire input sequence is first read and encoded into a fixed-length internal representation, which is then used by a decoder network to generate words until the end of the sequence is reached. Most deep learning-based semantic segmentation techniques employ an encoder-decoder architecture. Two prominent CNN-based encoder-decoder networks are SegNet (Badrinarayanan et al., 2017) and U-Net (Badrinarayanan et al., 2017; Ronneberger et al., 2015).
[0005] Transformers were first introduced in 2017 and were initially used for sequence-to-sequence learning. Because they easily model long-range dependencies, they have been widely used in natural language processing (NLP) and are gradually showing promise in computer vision. The Vision Transformers proposed by Dosovitskiy et al. first introduced Transformers to the image domain. They divide images into fixed-size blocks and then convert these blocks into vector representations. These vectors are processed through positional encoding and Transformer modules, enabling global context modeling and feature extraction of the image. This overcomes, to a certain extent, the limitations of CNNs in capturing long-range dependencies. Therefore, using Transformers as encoders or decoders is a new research direction.
[0006] Current problems with remote sensing image segmentation include: 1. For remote sensing image segmentation methods based on pure convolutional neural networks: CNN uses hierarchical feature representation and demonstrates strong local information extraction capabilities. However, due to the local nature of convolution operations, it is difficult to directly obtain global context information. To overcome this limitation, some works have added dilated convolutions or attention modules to their architectures to handle long-range dependencies. However, these methods still find it difficult to free the network from the local nature of convolution, thus limiting their performance on complex remote sensing images. 2. For remote sensing image segmentation methods based on Vision Transformers: The standard Transformer uses multi-head self-attention to calculate the relationships between all tokens, but this method results in a quadratic increase in computational complexity. Such complexity may lead to significant computational resource requirements when processing dense input tasks such as remote sensing images, limiting the scalability of Transformers in practical applications. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a remote sensing image segmentation method and system based on an improved SwinTransformer. The Swin transformer backbone network is combined with a global information enhancement module that performs global average pooling in the horizontal and vertical directions, thereby making up for the global information mining capability of the SwinTransformer backbone network that is limited by the window mechanism. The decoder fuses the encoder output through the attention mechanism, which can retain more local information while obtaining more extensive contextual information, thereby improving the accuracy of remote sensing image segmentation.
[0008] To achieve the above objectives, the first aspect of the present invention provides a remote sensing image segmentation method based on an improved SwinTransformer, comprising:
[0009] Obtaining a remote sensing image to be segmented;
[0010] Input the remote sensing image to be segmented into the trained segmentation model to obtain the segmentation result;
[0011] The segmentation model includes an encoder and a decoder. The encoder includes a Swin Transformer backbone network and a global information enhancement module. The different extraction stages of the Swin Transformer backbone network are used to extract the global context information of the remote sensing image to be segmented, and obtain global feature maps of different scales and levels.
[0012] The global information enhancement module is used to perform global average pooling operations on the input features of different extraction stages in the horizontal and vertical directions to obtain enhanced feature maps; the enhanced feature maps are fused with the global feature maps obtained at different extraction stages of the corresponding SwinTransformer backbone network to obtain global enhanced feature maps of different scales and levels;
[0013] The decoder fuses global enhanced feature maps of different scales and levels based on the attention mechanism to obtain the segmentation result.
[0014] A second aspect of the present invention provides a remote sensing image segmentation system based on an improved SwinTransformer, comprising:
[0015] An acquisition module, used for acquiring remote sensing images to be segmented;
[0016] The segmentation module is used to: input the remote sensing image to be segmented into the trained segmentation model to obtain the segmentation result;
[0017] The segmentation model includes an encoder and a decoder. The encoder includes a Swin Transformer backbone network and a global information enhancement module. The different extraction stages of the Swin Transformer backbone network are used to extract the global context information of the remote sensing image to be segmented, and obtain global feature maps of different scales and levels.
[0018] The global information enhancement module is used to perform global average pooling operations on the input features of different extraction stages in the horizontal and vertical directions to obtain enhanced feature maps; the enhanced feature maps are fused with the global feature maps obtained at different extraction stages of the corresponding SwinTransformer backbone network to obtain global enhanced feature maps of different scales and levels;
[0019] The decoder fuses global enhanced feature maps of different scales and levels based on the attention mechanism to obtain the segmentation result.
[0020] The third aspect of the present invention provides a computer device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, a remote sensing image method based on an improved SwinTransformer is performed.
[0021] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, a remote sensing image method based on an improved SwinTransformer is executed.
[0022] One or more of the above technical solutions have the following beneficial effects:
[0023] In the present invention, an encoder is constructed using a Swin transformer backbone network and a global information enhancement module. The Swin transformer backbone network fully mines the global contextual information of remote sensing images to obtain global feature maps of different scales and levels. The global information enhancement module is used to perform global average pooling operations on the feature maps in the horizontal and vertical directions. The global information enhancement module is fused with the global feature maps output by the corresponding Swin transformer backbone network at different stages, thereby making up for the global information mining capability of the Swin Transformer backbone network limited by the window mechanism. The decoder fuses the global enhanced feature maps of different scales and levels output by the encoder through an attention mechanism, thereby retaining more local information while obtaining more extensive contextual information. In addition, the segmentation model of the present invention can better capture the long-distance dependencies between semantic features than the CNN-based backbone network; and has fewer parameters and lower computational complexity than the Transformer-based backbone network.
[0024] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0026] Figure 1 This is a flow chart of the remote sensing image segmentation method in Example 1 of the present invention;
[0027] Figure 2 This is the overall network structure diagram of the segmentation model in Example 1 of the present invention;
[0028] Figure 3 This is a network structure diagram of the GIEM module in Example 1 of the present invention;
[0029] Figure 4 A visual comparison diagram of the method of this embodiment and other models in Example 1 of the present invention;
[0030] Figure 5Schematic diagram of the channel attention module structure in embodiment 1 of the present invention;
[0031] Figure 6 Schematic diagram of the spatial attention structure in Example 1 of the present invention. DETAILED DESCRIPTION
[0032] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0033] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.
[0034] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0035] Example 1
[0036] This embodiment discloses a remote sensing image segmentation method based on an improved SwinTransformer, comprising:
[0037] Obtaining a remote sensing image to be segmented;
[0038] Input the remote sensing image to be segmented into the trained segmentation model to obtain the segmentation result;
[0039] The segmentation model includes an encoder and a decoder. The encoder includes a Swin Transformer backbone network and a global information enhancement module. The different extraction stages of the Swin Transformer backbone network are used to extract the global context information of the remote sensing image to be segmented, obtaining global feature maps of different scales and levels.
[0040] The global information enhancement module is used to perform global pooling operations in the horizontal and vertical directions on the input features of different extraction stages to obtain enhanced feature maps; the enhanced feature maps are fused with the global feature maps obtained at different extraction stages of the corresponding Swin Transformer backbone network to obtain global enhanced feature maps of different scales and levels;
[0041] The decoder fuses the global enhanced feature maps of different scales and levels to obtain the segmentation result.
[0042] In this example, two publicly available remote sensing image segmentation datasets, Potsdam and Vaihingen, were selected as training data. The Potsdam dataset contains 38 image patches; 24 of them are used for training, and the remaining 14 are used for testing. Each patch is 6000 × 6000 pixels in size. The Vaihingen dataset contains 33 image patches; 17 of them are used for training, and the remaining 16 are used for testing. Each patch has a different size, approximately 2000 × 2500 pixels.
[0043] In this embodiment, the acquired remote sensing images are preprocessed. The preprocessing includes: first, randomly scaling and cropping the original remote sensing images and the ground truth segmentation images in the dataset to a size of 1024×1024. Second, the cropped training images and the corresponding ground truth segmentation images are horizontally and vertically flipped, randomly rotated, and normalized. This not only effectively compensates for the limited number of training images in the remote sensing dataset, improves model robustness, but also enhances the model's ability to resist overfitting.
[0044] like Figure 2 The figure shows the segmentation model of the high-resolution remote sensing image segmentation method based on SwinTransformer feature enhancement in this embodiment. The segmentation model consists of two parts: the first part is a SwinTransformer encoder with a GIEM global information enhancement module, and the second part is a feature fusion decoder composed of enhanced attention. The specific implementation of these two parts is as follows:
[0045] 1. The SwinTransformer encoder with the GIEM global information enhancement module mainly consists of four stages, and the specific operations include:
[0046] The input original remote sensing image is 3×1024×1024 in size, where 3 is the number of channels in the feature image and 1024×1024 represents the height and width of the feature image. Patch Partitioning is used to partition the image into non-overlapping blocks of 256×256, and each block is projected into a 48-dimensional feature space. The first stage involves mapping the 48-dimensional features into a 96-dimensional space using a linear embedding layer. The feature map is then fed into the SwinTransformer Block module and the GIEM global information enhancement module designed in this example. The outputs of these two modules are summed to produce the output of the first stage.
[0047] The 96×256×256 feature map output from the first stage is fed into the second stage. In the second stage, the feature map is first merged in the spatial dimension through the PatchMerging layer to obtain a feature map of size 192×128×128. The obtained 192×128×128 feature map is fed into the SwinTransformer Block module and the GIEM module respectively, and the outputs of these two modules are added together to obtain the output of the second stage.
[0048] The 192×128×128 feature map output from the second stage is sent to the third stage. In the third stage, the feature map is first merged in the spatial dimension through the PatchMerging layer to obtain a 384×64×64 feature map. The obtained 384×64×64 feature map is respectively sent to the SwinTransformerBlock module and the GIEM module, and then the outputs of these two modules are added together to obtain the output of the third stage.
[0049] The 384×64×64 feature map output from the third stage is sent to the fourth stage. In the fourth stage, the feature map is first merged in the spatial dimension through the PatchMerging layer to obtain a feature map of 768×32×32 size; the obtained 768×32×32 feature map is sent to the SwinTransformerBlock and GIEM modules respectively, and then the outputs of these two modules are added together to obtain the output of the fourth stage.
[0050] Among them, SwinTransformerBlock includes window multi-head self-attention W-MSA and shift window multi-head self-attention SW-MAS.
[0051] like Figure 3 As described above, in the GIEM module of this embodiment, the specific operation is as follows: the output of the linear embedding of the first stage or the output feature s(h,w,c) of the second to fourth stages of PatchMerging is fed into a 3×3 dilated convolution layer with an expansion rate of 2 to reconstruct the structural information of the feature map by expanding the receptive field; at the same time, in order to reduce memory overhead, the number of channels is reduced to c / 2. Next, a global average pooling operation is performed on the data in the vertical and horizontal directions of the input feature map. Specifically, the calculation formulas for the elements in the vertical and horizontal directions are as follows:
[0052]
[0053]
[0054] where i, j, and k are the indices in the vertical direction, horizontal direction, and channels respectively, where 0 ≤ i < h, 0 ≤ j < w, and 0 ≤ k < c / 2. Feature f(·) is an dilated convolutional layer with batch normalization and GELU activation function. Denote the aggregated tensors in the horizontal and vertical directions as v h and v w . v w ∈R h×1×(c / 2) , v h ∈R 1×w×(c / 2) ; then multiply v h and v w to obtain the attention map W with enhanced spatial relationship, W ∈ R h×w×(c / 2) ; finally, add the output result to the output feature t l+1 of the Swin Transformer Block through convolution, normalization, and gelu activation function, and pass the output result to the next layer.
[0055] 2. Decoder composed of enhanced attention modules:
[0056] Fuse the feature map obtained by 1×1 convolution of the feature map with a size of 768×32×32 obtained in the fourth stage of the encoder, and the feature obtained after sequentially passing through 1×1 convolution operation, downsampling operation, spatial attention enhancement module, and downsampling operation of the feature map with a size of 192×128×128 obtained in the second stage of the encoder, to generate the fourth feature map FF4 with a size of 768×32×32.
[0057] Fuse the feature map obtained by sending the feature map with a size of 384×64×64 obtained in the third stage of the encoder through 1×1 convolution into the spatial attention enhancement module, and the feature map obtained after sequentially passing through 1×1 convolution operation, downsampling operation, channel attention enhancement module, and downsampling operation of the feature map with a size of 96×256×256 obtained in the first stage of the encoder, as well as the feature map obtained after upsampling the fourth feature map FF4, to generate the third feature map FF3 with a size of 384×64×64.
[0058] Fuse the feature map obtained by sending the feature map with a size of 192×128×128 obtained in the second stage of the encoder through 1×1 convolution into the channel attention enhancement module, and the feature map obtained after upsampling the third feature map FF3, to generate the second feature map FF2 with a size of 192×128×128.
[0059] Fuse the feature map obtained after 1×1 convolution of the feature map with a size of 96×256×256 obtained in the first stage of the encoder, and the feature map obtained after upsampling the second feature map FF2 to generate the first feature map FF1 with a size of 96×256×256.
[0060] Finally, four feature fusions FF1, FF2, FF3, and FF4 are generated. The final feature map is passed to the segmentation head, which consists of a series of convolution and bilinear interpolation upsampling operations to obtain the final prediction map.
[0061] (1) Downsample: Downsample can connect features at different levels in the encoding and decoding process and further fuse features at different scales and levels. It is defined as follows:
[0062]
[0063] Here, Z represents the input vector, i.e., the output of the first or second stage of the encoder after a 1×1 convolution; σ1 represents the Reluctant Unified Unit (ReLU) activation function. δ1 and μ1 represent 3×3 convolutional layers with a stride of 2, and θ1 represents a 3×3 convolutional layer with a stride of 1. Both convolutional layers include normalization. The number of input and output channels is determined by m and n, respectively.
[0064] (2) Spatial Attention Enhancement Module: Based on the linear attention mechanism, the SAEM spatial attention enhancement module is designed. The output of the third stage of the encoder is subjected to 1×1 convolution and the output of the second stage is downsampled as the input of SAEM. The SAEM module is used to enhance the weight of more important spatial dimension features. Figure 5 As shown. The formula is defined as:
[0065]
[0066] Among them, Q(X1), K(X1) and V(X1) represent the corresponding convolution operations on the input feature map X1 to generate the query matrix Bond Matrix Value Matrix P represents the number of pixels in the input feature map. c represents the channel dimension, i.e., the number of channels in the input feature map of the spatial attention enhancement module. d represents the spatial dimension, i.e., the spatial position of each pixel in the feature map.
[0067] (3) Channel Attention Enhancement Module: The output of the second stage of the encoder is subjected to 1×1 convolution and the output of the first stage is downsampled and used as the input of CAEM. The channel attention enhancement module CAEM is used to enhance the feature weights between channel dimensions. Figure 6 As shown. The formula is defined as:
[0068]
[0069] Among them, W(X2) represents the reintegration of the spatial dimension of the input feature map X2 of the channel attention enhancement module, and P represents the number of pixels in the input feature map.
[0070] (4) Feature fusion: The final four feature fusion modules of the decoder are calculated as follows:
[0071] FF4=T4+D(SAEM(D(T2))) (6)
[0072] FF3 = SAEM(T3)+D(CAEM(D(T1)))+S(FF4) (7)
[0073] FF2=CAEM(T2)+S(FF3) (8)
[0074] FF1 = T1 + S (FF2) (9)
[0075] Among them, S represents bilinear interpolation upsampling and D represents downsampling operation.
[0076] T1, T2, T3, and T4 represent the outputs of the first, second, third, and fourth stages of the encoder, respectively, after 1×1 convolution.
[0077] The loss function is used to calculate the error between the predicted value of the segmentation model and the actual segmented image. For the feature map of each stage, the weighted sum of the cross entropy loss function and the Dice loss function is used as the loss function of the model:
[0078] L=L CE +L Dice (10)
[0079] This example uses the AdamW optimizer, with an initial learning rate of 1e-3 and a weight decay of 2.5e-4. Setting the weight decay coefficient can prevent overfitting of the model, and adaptive learning rate adjustment can accelerate the convergence of the model.
[0080] In this implementation, for model training and testing, the training images were preprocessed and augmented as described above. The resulting training images were fed into a SwinTransformer encoder with a GIEM global information enhancement module, and then into a decoder with an enhanced attention module to produce the model's final predictions. The loss between the predicted and true segmentation images was calculated using the loss function designed above. Finally, the AdamW optimizer was used for gradient updates. Each training session consisted of 8 samples, with 50 epochs for the Potsdam dataset and 100 epochs for the Vaihingen dataset. MIoU, F1, and OA were used as evaluation metrics.
[0081] The comparison models used in the experiment are several popular remote sensing image segmentation methods. The experimental data comparison with other models on the Vaihingen dataset is shown in Table 1, and the visual comparison with other models is shown in Figure 4 .
[0082] Table 1: Experimental comparison results of the method in this embodiment and other models on the Vaihingen dataset:
[0083]
[0084] This embodiment performs semantic segmentation on remote sensing images based on SwinTransformer, and the model uses an encoding-decoding structure. Compared with the CNN-based backbone network, it can better capture the long-distance dependencies between semantic features. Compared with the Transformer-based backbone network, it has fewer parameters and lower computational complexity. Specifically, the GIEM module can effectively make up for the global modeling capability of SwinTransformer limited by the window mechanism, and EFFM can effectively aggregate semantic features of different levels and scales. The combination of cross entropy loss and Dice loss enables the model to have faster convergence speed and better performance. Experimental results on the Potsdam and Vaihingen datasets show that the model has good accuracy.
[0085] Example 2
[0086] The purpose of this embodiment is to provide a remote sensing image segmentation system based on an improved SwinTransformer, including:
[0087] An acquisition module, used for acquiring remote sensing images to be segmented;
[0088] The segmentation module is used to: input the remote sensing image to be segmented into the trained segmentation model to obtain the segmentation result;
[0089] The segmentation model includes an encoder and a decoder. The encoder includes a Swin Transformer backbone network and a global information enhancement module. The different extraction stages of the Swin Transformer backbone network are used to extract the global context information of the remote sensing image to be segmented, and obtain global feature maps of different scales and levels.
[0090] The global information enhancement module is used to perform global pooling operations on the input features of different extraction stages in the horizontal and vertical directions to obtain enhanced feature maps; the enhanced feature maps are fused with the global feature maps obtained at different extraction stages of the corresponding Swin Transformer backbone network to obtain global enhanced feature maps of different scales and levels;
[0091] The decoder fuses global enhanced feature maps of different scales and levels based on the attention mechanism to obtain the segmentation result.
[0092] Example 3
[0093] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the program.
[0094] Example 4
[0095] The purpose of this embodiment is to provide a computer-readable storage medium.
[0096] A computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the above method.
[0097] The steps involved in the apparatuses of Examples 2, 3, and 4 above correspond to those of Method Example 1. For detailed implementations, please refer to the relevant description of Example 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and causing the processor to perform any method of the present invention.
[0098] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0099] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A remote sensing image segmentation method based on improved SwinTransformer, characterized in that: include: Obtaining a remote sensing image to be segmented; Input the remote sensing image to be segmented into the trained segmentation model to obtain the segmentation result; The segmentation model includes an encoder and a decoder. The encoder includes a Swin Transformer backbone network and a global information enhancement module. The different extraction stages of the Swin Transformer backbone network are used to extract the global context information of the remote sensing image to be segmented, and obtain global feature maps of different scales and levels. The global information enhancement module is used to perform global average pooling operations on the input features of different extraction stages in the horizontal and vertical directions to obtain enhanced feature maps; the enhanced feature maps are fused with the global feature maps obtained at different extraction stages of the corresponding Swin Transformer backbone network to obtain global enhanced feature maps of different scales and levels; The decoder fuses global enhanced feature maps of different scales and levels based on the attention mechanism to obtain the segmentation result.
2. The remote sensing image segmentation method based on the improved SwinTransformer according to claim 1, characterized in that: In the encoder, the input remote sensing image to be segmented is divided into blocks, the blocked images are spatially mapped, and the feature maps after spatial dimension mapping are sequentially subjected to different extraction stages of the Swin Transformer backbone network for feature extraction. In each stage, specifically: for the input feature map of each extraction stage, an enhanced feature map and a global feature map are obtained based on the global information enhancement module and the Swin TransformerBlock module, and the enhanced feature map and the global feature map are added to obtain a global enhanced feature map.
3. The remote sensing image segmentation method based on the improved SwinTransformer according to claim 1, characterized in that: In the global information enhancement module, the specific operations are: The input features are reconstructed through the dilated convolution layer to obtain the reconstructed feature map; Perform global pooling operations on the reconstructed feature map in the horizontal and vertical directions to obtain horizontal pooling features and vertical pooling features; Multiply the horizontal pooling features and the vertical pooling features, and pass the multiplication results through the convolution layer, normalization layer and activation function respectively to obtain the enhanced feature map.
4. The remote sensing image segmentation method based on the improved SwinTransformer according to claim 1, characterized in that: In the decoder, the specific operations are: The fourth feature map is obtained by fusing the convolutional feature map of the output of the fourth stage of the encoder with the convolutional, downsampling, spatial attention module and downsampling feature map of the output of the second stage of the encoder. The feature map generated by the output of the third stage of the encoder after convolution and spatial attention module is fused with the feature map after the output of the first stage of the encoder after convolution, downsampling, channel attention module and downsampling, and the features of the fourth feature map after upsampling to obtain the third feature map; The feature map generated by the output of the second stage of the encoder after the convolution operation and the channel attention module is fused with the feature map after the upsampling of the third feature map to obtain a second feature map; The feature map after the convolution operation on the output of the first stage of the encoder is fused with the feature map after the upsampling of the second feature map to obtain a first feature map; The first feature map is sent to the segmentation head, and the segmentation result is output.
5. The remote sensing image segmentation method based on the improved SwinTransformer according to claim 4, characterized in that: The specific operation of downsampling on the input vector is: The input vector is sequentially passed through a 3×3 convolutional layer with a stride of 2 and an activation function to obtain the first output result; The input vector is sequentially passed through a 3×3 convolutional layer with a stride of 1 and a 3×3 convolutional layer with a stride of 2 to obtain the second output result; The first output result and the second output result are added to obtain the down-sampled output result.
6. The remote sensing image segmentation method based on the improved SwinTransformer according to claim 1, characterized in that: The cross entropy loss function and the dice loss function are added together as the loss function of the segmentation model, and the segmentation model is trained using the obtained loss function.
7. A remote sensing image segmentation system based on improved SwinTransformer, characterized in that: include: An acquisition module, used for acquiring remote sensing images to be segmented; The segmentation module is used to: input the remote sensing image to be segmented into the trained segmentation model to obtain the segmentation result; The segmentation model includes an encoder and a decoder. The encoder includes a Swin Transformer backbone network and a global information enhancement module. The different extraction stages of the Swin Transformer backbone network are used to extract the global context information of the remote sensing image to be segmented, and obtain global feature maps of different scales and levels. The global information enhancement module is used to perform global pooling operations on the input features of different extraction stages in the horizontal and vertical directions to obtain enhanced feature maps; the enhanced feature maps are fused with the global feature maps obtained at different extraction stages of the corresponding Swin Transformer backbone network to obtain global enhanced feature maps of different scales and levels; The decoder fuses global enhanced feature maps of different scales and levels based on the attention mechanism to obtain the segmentation result.
8. The remote sensing image segmentation system based on the improved SwinTransformer according to claim 7, characterized in that: In the segmentation module, the specific operations of the decoder are: The fourth feature map is obtained by fusing the convolutional feature map of the output of the fourth stage of the encoder with the convolutional, downsampling, spatial attention module and downsampling feature map of the output of the second stage of the encoder. The feature map generated by the output of the third stage of the encoder after convolution and spatial attention module is fused with the feature map after the output of the first stage of the encoder after convolution, downsampling, channel attention module and downsampling, and the features of the fourth feature map after upsampling to obtain the third feature map; The feature map generated by the output of the second stage of the encoder after the convolution operation and the channel attention module is fused with the feature map after the upsampling of the third feature map to obtain a second feature map; The feature map after the convolution operation on the output of the first stage of the encoder is fused with the feature map after the upsampling of the second feature map to obtain a first feature map; The first feature map is sent to the segmentation head, and the segmentation result is output.
9. A computer device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, a remote sensing image segmentation method based on an improved SwinTransformer as described in any one of claims 1 to 6 is performed.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the remote sensing image segmentation method based on the improved SwinTransformer according to any one of claims 1 to 6 is executed.