Remote sensing image earth surface type segmentation method based on global-local transfomer block
Through the remote sensing image segmentation method of the global-local Transformer block, combined with the ResNet18 encoder and the global-local Transformer block decoder, the problem of inefficient global context modeling of CNN in remote sensing images is solved, and efficient and accurate surface type segmentation of remote sensing images is achieved.
Patent Information
- Application Number
- CN202510673223.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
AI Technical Summary
Existing CNN-based semantic segmentation methods for remote sensing images have difficulty in effectively modeling global contextual information, resulting in high computational complexity and limited real-time applications in high-resolution remote sensing images.
A global-local Transformer block-based method is adopted, combining a lightweight ResNet18 encoder and a global-local Transformer block decoder. The global and local context features are captured through a dual-branch structure, and feature fusion is performed through a weighted sum operation to achieve efficient semantic segmentation.
It significantly improves the accuracy and speed of remote sensing image segmentation, reduces computational complexity, and is suitable for complex object segmentation in real-time scenarios.
Smart Images

Figure CN120599255A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing, and in particular relates to a remote sensing image surface type segmentation method based on global-local transformer blocks. Background Art
[0002] Semantic segmentation of remote sensing images is a key technology that aims to classify each pixel in a remote sensing image into predefined semantic categories. With the advancement of remote sensing technology and the development of sensors, the acquired remote sensing images have higher resolution and wider coverage, which poses challenges and opportunities for semantic segmentation of remote sensing images. Semantic segmentation has broad application prospects in agriculture, urban planning, environmental monitoring and other fields. For example, in agriculture, semantic segmentation of remote sensing images can be used to monitor the types and growth status of farmland crops, which helps in the management and decision-making of agricultural production; in urban planning, semantic segmentation can be used to identify different types of land features such as buildings, roads, and green spaces, providing important information support for urban planning and management; in environmental monitoring, semantic segmentation technology can be used to monitor and analyze natural environments such as forests, lakes, and grasslands, thereby realizing timely monitoring and early warning of environmental changes.
[0003] However, due to the complexity and diversity of remote sensing images, accurately extracting semantic information about objects remains a challenging problem. Traditional pixel-based classification methods are often limited by the limitations of feature extraction and classifier design, making them difficult to handle complex objects and scenes in remote sensing images. Therefore, in recent years, semantic segmentation methods based on deep learning have gradually become a research hotspot. Deep learning models can learn high-level feature representations of images, thus achieving good results in semantic segmentation tasks. Commonly used deep learning models include FCN (Fully Convolutional Network) and U-Net, which have achieved remarkable results in the field of semantic segmentation.
[0004] Driven by continuous advances in sensor technology, the ability to capture high-resolution remote sensing imagery of urban scenes has significantly increased worldwide. These images, with their rich spatial details (such as building textures and road orientation) and latent semantic content (such as land use types and functional zoning), have become a core data foundation for refined urban management. Semantic segmentation (i.e., pixel-level object classification and segmentation) of urban scene images has spawned key applications such as land cover mapping, surface change detection, and road and building extraction, supporting the development of important areas such as smart city planning and environmental monitoring. In recent years, deep learning techniques, represented by convolutional neural networks (CNNs), have dominated semantic segmentation. Through hierarchical feature extraction using multiple layers of convolutional kernels, CNNs can automatically capture multi-scale local contextual information (such as edges and textures) in images. Compared to traditional methods such as support vector machines and random forests, end-to-end learning frameworks significantly improve the robustness of feature representation and classification accuracy. Typical architectures, such as U-Net and DeepLab, further optimize the aggregation of local features through designs such as dilated convolutions and feature pyramids.
[0005] However, existing CNN-based methods suffer from a fundamental technical bottleneck: while the fixed-receptive-field convolutional operations they rely on excel at extracting local patterns, they struggle to model the global context and long-range dependencies prevalent in urban scenes. For example, accurate classification of a "road" requires considering the spatial distribution of surrounding "buildings" and "vegetation," while discrimination based solely on local textures can easily lead to semantic confusion among similar features. While the self-attention mechanism addresses this shortcoming by calculating global dependencies between pixels, its high temporal complexity leads to hundreds of GB of memory consumption and seconds-long computational delays in high-resolution remote sensing imagery (often reaching thousands of pixels), severely limiting the model's application in real-time scenarios.
[0006] Therefore, the existing deep learning technology represented by CNN has the problem of "low efficiency of global context modeling" in semantic segmentation tasks. Summary of the Invention
[0007] In response to the core pain point of "inefficient global context modeling" mentioned above, the present invention proposes a surface type segmentation method for remote sensing images based on global-local transformer blocks. While retaining the advantages of CNN local feature extraction, it achieves efficient modeling of global semantic dependencies with low computational complexity. It is expected to significantly reduce the global relationship calculation cost, improve the segmentation accuracy and inference speed in complex scenarios, and provide a solution for surface type segmentation of remote sensing images that is both accurate and efficient.
[0008] To fully utilize the global context extraction capabilities of the Transformer architecture without incurring excessive complexity, this paper proposes a method for remote sensing image surface type segmentation based on a global-local transformer block. In this method, the Transformer has a CNN-based encoder and a Transformer-based decoder, which is used to effectively semantically segment remote sensing urban scene images. Specifically, this method uses a lightweight backbone network, namely ResNet18, as the encoder and develops an efficient global-local attention mechanism to construct the Transformer block in the decoder. The efficient global-local attention mechanism proposed in this paper adopts a dual-branch structure, namely a global branch and a local branch. This structure allows the attention block to capture both global and local context, thereby surpassing the single-branch efficient attention mechanism in the Transformer that only captures global context. Finally, the conversion from the original image to the longitude and latitude coordinates of the contour of a fixed surface type is completed through multiple mapping functions.
[0009] The technical solution adopted by the present invention to solve the technical problem includes the following steps:
[0010] Step 1: For RGB image Here, h and w represent the height and width of the image, respectively. The pre-trained ResNet18 is selected as the CNN-Based encoder to extract multi-scale semantic features with a lower computational cost.
[0011] Input the image into the encoding layer and pass the input image through the CNN-Based encoder
[0012] e=CNN Encoder(img)
[0013] Here, CNN Encoder (·) represents the CNN-Based encoder, and e represents the features encoded by the encoder. The CNN Encoder here uses the pre-trained ResNet18 as the encoder. ResNet18 consists of four layers of Resblocks, each of which is downsampled by a scaling factor of 2. The feature maps generated by each layer are fused with the corresponding feature maps of the decoder through a 1×1 convolution with a channel dimension of 64, i.e., a skip connection. Specifically, a weighted sum operation is used to aggregate the semantic features generated by Resblocks with the features generated by the global-local Transformer block of the decoder as follows.
[0014] FF=α·RF+(1-α)·GLF
[0015] Where FF represents fused features, RF represents features generated by Resblocks, and GLF represents features generated by global-local Transformer blocks.
[0016] Step 2: Decode the features encoded by the Encoder through the decoder of the transformer architecture.
[0017] S=Transformer Decoder(e)
[0018] Transformer Decoder (·) represents the global-local Transformer block decoder, which includes two parallel branches constructed with global-local attention to extract global and local context respectively, where the local branch uses two parallel convolutional layers with kernel sizes of 3 and 1 to extract local context. Then, two batch normalization operations are appended before the final summation operation. The global branch deploys window-based multi-head self-attention to capture global context. The input 2D feature map ∈ R is first convolved using a standard 1×1 convolution. B×C×H×W The channel dimension is expanded to three times. Then, the window splitting operation is used to convert the one-dimensional sequence It is split into query vector (Q), keyword vector (K) and value vector (V), and the channel size C is set to 64. The window size w and the number of heads h are both set to 8.
[0019] Step 3: Use the fused features obtained in step 1 as the input to the feature refinement head. Two paths are constructed to strengthen the channel-wise and spatial feature representations.
[0020] R=FRH(FF)
[0021] Where R represents the segmentation result after feature refinement head processing, FRH(·) represents the feature refinement head, and FF represents the fused features obtained by step 1. Specifically, the channel path uses a global average pooling layer to generate the channel attention map C∈R 1×1×c , where c represents the channel dimension. The reduce & expand operation consists of two 1×1 convolutional layers, which first reduce the channel dimension c to 4 times of the original and then expand it to the original. The spatial path uses depthwise convolution to generate a spatial attention map S∈R h×w×1 , where h and w represent the spatial resolution of the feature map. The attention features generated by the two paths are further fused through the summation operation. Finally, the final segmentation map is obtained through post-processing 1×1 convolutional layers and upsampling operations. It is worth noting that the residual connection is introduced to prevent network degradation. The specific structure is as follows Figure 2 shown.
[0022] Step 4: Map the binary segmentation result obtained in step 3 into an RGB three-channel image through the RGB mapping relationship.
[0023] S=RGBmap(R)
[0024] RGBmap(·) represents the binary to RGB mapping function. The findContours library function provided by OpenCV is then used to extract the contour coordinates of the fixed RGB value. The left side of the pixel is then converted to longitude and latitude coordinates through the mapping relationship between pixel coordinates and longitude and latitude coordinates and saved in the geojson file.
[0025] The beneficial effects of the present invention are as follows:
[0026] Semantic segmentation of terrain and landforms is a key component of the terrain and landform analysis module. Traditional detection and segmentation algorithms are designed to solve specific problems. Traditional CNN-based methods have low computational overhead but lack versatility. Deep learning methods are highly scalable, but they require large amounts of data and have high computational complexity. To balance interpretation accuracy and resource consumption, a combination of two approaches is needed to maximize system efficiency. This paper uses a UNet-like transformer (UNetFormer) for real-time terrain and geological scene segmentation. To achieve efficient segmentation, this method uses a lightweight ResNet-18 encoder and develops an efficient global-local attention mechanism to model global and local information in the decoder. This method is not only faster but also more accurate. It utilizes three global-local transformer blocks (GLTBs) and a feature refinement head to construct a lightweight transformer-based decoder. Through this layered and lightweight design, the decoder can capture global and local contextual features at multiple scales while maintaining high efficiency, addressing the technical difficulties of pixel-level classification. Features generated by the residual block are aggregated with those generated by the decoder's GLTBs through a weighted summation. The weighted sum operation selectively weights two features according to their contribution to segmentation accuracy, thereby learning more general fusion features and solving the difficulty of complex ground background. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Schematic diagram of the model structure of the remote sensing image surface type segmentation method based on global-local transformer blocks.
[0028] Figure 2 Detailed model structure of the feature refinement head. DETAILED DESCRIPTION
[0029] The present invention will be further described below with reference to the accompanying drawings and examples.
[0030] like Figure 1 As shown, to fully utilize the global context extraction capabilities of the Transformer architecture without incurring excessive complexity, the present invention proposes a remote sensing image surface type segmentation method based on a global-local transformer block. In this method, the Transformer has a CNN-based encoder and a Transformer-based decoder, which is used to effectively semantically segment remote sensing urban scene images. Specifically, to overcome the shortcomings of the existing technology, this method uses a lightweight backbone network, namely ResNet18, as the encoder and develops an efficient global-local attention mechanism to construct the Transformer block in the decoder. The efficient global-local attention mechanism proposed in this method adopts a dual-branch structure, namely a global branch and a local branch. This structure allows the attention block to capture both global and local context, thereby surpassing the single-branch efficient attention mechanism in the Transformer that only captures global context. Finally, the conversion from the original image to the longitude and latitude coordinates of the contour of a fixed surface type is completed through multiple mapping functions.
[0031] A method for surface type segmentation of remote sensing images based on global-local transformer blocks includes the following aspects and steps:
[0032] Remote sensing surface type segmentation model structure:
[0033] Part 1: For RGB images Here, h and w represent the height and width of the image, respectively. The pre-trained ResNet18 is selected as the CNN-Based encoder to extract multi-scale semantic features with a lower computational cost.
[0034] Input the image into the encoding layer and pass the input image through the CNN-Based encoder
[0035] e=CNN Encoder(img)
[0036] Here, CNN Encoder (·) represents the CNN-Based encoder, and e represents the features encoded by the encoder. The CNN Encoder here uses the pre-trained ResNet18 as the encoder. ResNet18 consists of four layers of Resblocks, each of which is downsampled by a scaling factor of 2. The feature maps generated by each layer are fused with the corresponding feature maps of the decoder through a 1×1 convolution with a channel dimension of 64, i.e., a skip connection. Specifically, a weighted sum operation is used to aggregate the semantic features generated by Resblocks with the features generated by the global-local Transformer block of the decoder as follows.
[0037] FF=α·RF+(1-α)·GLF
[0038] Where FF represents fused features, RF represents features generated by Resblocks, and GLF represents features generated by global-local Transformer blocks.
[0039] Part 2: The features encoded by the Encoder are decoded by the decoder of the transformer architecture.
[0040] s=Transformer Decoder(e)
[0041] Transformer Decoder (·) represents the global-local Transformer block decoder, which includes two parallel branches constructed with global-local attention to extract global and local context respectively, where the local branch uses two parallel convolutional layers with kernel sizes of 3 and 1 to extract local context. Then, two batch normalization operations are appended before the final summation operation. The global branch deploys window-based multi-head self-attention to capture global context. The input 2D feature map ∈ R is first convolved using a standard 1×1 convolution. B×C×H×W The channel dimension is expanded to three times. Then, the window splitting operation is used to convert the one-dimensional sequence Split into query vector (Q), keyword vector (K) and value vector (V), the channel size C is set to 64. The window size w and the number of heads h are both set to 8.
[0042] Part 3: The fused features obtained in Part 1 are used as the input to the feature refinement head. Two paths are constructed to strengthen the channel-wise and spatial-wise feature representations.
[0043] R=FRH(FF)
[0044] Where R represents the binary segmentation result after feature refinement head processing, FRH(·) represents the feature refinement head, and FF represents the fused features obtained from part 1. Specifically, the channel path uses a global average pooling layer to generate the channel attention map C∈R 1×1×c , where c represents the channel dimension. The reduce & expand operation consists of two 1×1 convolutional layers, which first reduce the channel dimension c to 4 times of the original and then expand it to the original. The spatial path uses depthwise convolution to generate a spatial attention map S∈R h×w×1 , where h and w represent the spatial resolution of the feature map. The attention features generated by the two paths are further fused through the summation operation. Finally, the final segmentation map is obtained through post-processing 1×1 convolutional layers and upsampling operations. It is worth noting that the residual connection is introduced to prevent network degradation. The specific structure is as follows Figure 2 shown.
[0045] Part 4: Map the binary segmentation result obtained in Part 3 into an RGB three-channel image through the RGB mapping relationship.
[0046] S=RGBmap(R)
[0047] RGBmap(·) represents the binary to RGB mapping function. The findContours library function provided by OpenCV is used to extract the contour coordinates of the fixed RGB value. The pixel coordinates are then converted into longitude and latitude coordinates through the mapping relationship between pixel coordinates and longitude and latitude coordinates and saved in the geojson file.
[0048] Based on the above analysis, the present invention proposes a remote sensing image surface type segmentation method based on global-local transformer blocks, which includes the following steps:
[0049] Step 1: For RGB image Where h and w represent the height and width of the image respectively. The pre-trained ResNet18 is selected as the CNN-Based encoder to extract multi-scale semantic features.
[0050] Input the image into the encoding layer and pass the input image through the CNN-Based encoder
[0051] e=CNNEncoder(img)
[0052] Among them, CNNEncoder·) represents the CNN-Based encoder, and e represents the feature after being encoded by the encoder;
[0053] In step 1, ResNet18 consists of four layers of Resblock, each layer is downsampled with a scaling factor of 2, and the feature maps generated by each layer are fused with the corresponding feature maps of the decoder through 1×1 convolution with a channel dimension of 64, that is, skip connection.
[0054] Step 2: Decode the features encoded by the CNN-Based encoder through the decoder of the transformer architecture;
[0055] In step 2, the decoder of the transformer architecture consists of two parallel branches constructed by global-local attention to extract global and local context respectively.
[0056] The method for extracting local context in a local branch is as follows:
[0057] Two parallel convolutional layers with kernel sizes of 3 and 1 are employed to extract local context; then, two batch normalization operations are appended before the final summation operation.
[0058] The global branch deploys window-based multi-head self-attention to capture global context. The global branch extracts global context as follows:
[0059] First, the input 2D feature map ∈ R is transformed using a standard 1×1 convolution B×C×H×W The channel dimension is expanded to three times; then, the one-dimensional sequence is divided into It is split into query vector (Q), keyword vector (K) and value vector (V), and the channel size C is set to 64. The window size w and the number of heads h are both set to 8.
[0060] Step 3: Aggregate the semantic features produced by Resblocks with the features generated by the global-local Transformer block of the decoder using a weighted sum operation as follows;
[0061] FF=α·RF+(1-α)·GLF
[0062] Where FF represents fused features, RF represents features generated by Resblocks, and GLF represents features generated by global-local Transformer blocks;
[0063] Step 4: The fused features are used as the input of the feature refinement head to construct the channel path and the spatial path. The attention features generated by the two paths are further fused through the summation operation. Finally, the final segmentation map is obtained through post-processing 1×1 convolution layer and upsampling operation.
[0064] R=FRH(FF)
[0065] Where R represents the segmentation result after being processed by the feature refinement head, FRH(·) represents the feature refinement head, and FF represents the fused feature obtained in part 1;
[0066] In step 4, the channel path uses a global average pooling layer to generate the channel attention map C∈R 1×1×c , where c represents the channel dimension; the reduce & expand operation contains two 1×1 convolutional layers, which first reduce the channel dimension c to 4 times of the original one and then expand it to the original one.
[0067] In step 4, the spatial path uses depth-wise convolution to generate a spatial attention map S∈R h×w×1 , where h and w represent the spatial resolution of the feature map.
[0068] Step 5: Map the binary segmentation result into an RGB three-channel image through the RGB mapping relationship; then use the findContours library function provided by OpenCV to extract the contour coordinates of the fixed RGB value, and convert the left side of the pixel into longitude and latitude coordinates through the mapping relationship between pixel coordinates and longitude and latitude coordinates and save it in the geojson file;
[0069] S=RGBmap(R)
[0070] Where RGBmap(·) represents the mapping function from binary to RGB. Specific embodiment:
[0072] The present invention provides a method for surface type segmentation of remote sensing images based on global-local transformer blocks. The specific process is as follows:
[0073] 1. Data preprocessing
[0074] For a given training set There is an RGB image and ground truth label pair {img i ,lab i}, for a given test set No label exists i When training the model, the RGB data is normalized to the range of [0, 1], and the label data is a grayscale image, where the grayscale value represents the surface type category corresponding to the pixel.
[0075] In addition, for the input image img i and its corresponding label i Random cropping, random horizontal flipping, and random vertical flipping are used to enhance the generalization ability of the model.
[0076] 2. Image segmentation based on Unetfomer
[0077] A lightweight transformer-based decoder is constructed using three GLTBs (global-local transformer blocks) and a feature refinement head. Through this layered and lightweight design, the decoder is able to capture global and local contextual features at multiple scales while maintaining high efficiency to address the technical difficulties of pixel-level classification. The features generated by the residual block are aggregated with the features generated by the decoder's GLTB through a weighted summation. The weighted sum operation selectively weights the two features according to their contribution to segmentation accuracy, thereby learning a more general fused feature and ultimately obtaining the segmentation result.
[0078] 3. Contour extraction method based on pixel RGB
[0079] The image segmentation category result mask obtained by Unetfomer is stored in the form of a binary image. The mask image is converted into an RGB three-channel image through a predefined binary to RGB conversion relationship, and then the contour information of the target area is extracted through a contour extraction method based on pixel RGB. The specific operation is: traverse each pixel point of the segmentation result, extract the contour coordinates of the fixed RGB value through the findContours library function provided by OpenCV, and then convert the pixel coordinates into longitude and latitude coordinates through the mapping relationship between pixel coordinates and longitude and latitude coordinates and save them in the geojson file.
[0080] After the image converts the coordinate information into longitude and latitude information, it forms a geojson file and outputs it.
Claims
1. A method for surface type segmentation of remote sensing images based on global-local transformer blocks, characterized in that: The following steps are involved: Step 1: For RGB image Where h and w represent the height and width of the image respectively. The pre-trained ResNet18 is selected as the CNN-Based encoder to extract multi-scale semantic features. Input the image into the encoding layer and pass the input image through the CNN-Based encoder e=CNN Encoder(img) Where CNN Encoder(·) represents the CNN-Based encoder, and e represents the feature after being encoded by the encoder; Step 2: Decode the features encoded by the CNN-Based encoder through the decoder of the transformer architecture; Step 3: Aggregate the semantic features produced by Resblocks with the features generated by the global-local Transformer block of the decoder using a weighted sum operation as follows; FF=α·RF+(1-α)·GLF Where FF represents fused features, RF represents features generated by Resblocks, and GLF represents features generated by global-local Transformer blocks; Step 4: The fused features are used as the input of the feature refinement head to construct the channel path and the spatial path. The attention features generated by the two paths are further fused through the summation operation. Finally, the final segmentation map is obtained through post-processing 1×1 convolution layer and upsampling operation. R=FRH(FF) Where R represents the segmentation result after being processed by the feature refinement head, FRH(·) represents the feature refinement head, and FF represents the fused feature obtained in part 1; Step 5: Map the binary segmentation result into an RGB three-channel image through the RGB mapping relationship; then use the findContours library function provided by OpenCV to extract the contour coordinates of the fixed RGB value, and convert the left side of the pixel into longitude and latitude coordinates through the mapping relationship between pixel coordinates and longitude and latitude coordinates and save it in the geojson file; S=RGBmap(R) Where RGBmap(·) represents the mapping function from binary to RGB.
2. The method for surface type segmentation of remote sensing images based on global-local transformer blocks according to claim 1, characterized in that: In step 1, ResNet18 consists of four layers of Resblock, each layer is downsampled with a scaling factor of 2, and the feature maps generated by each layer are fused with the corresponding feature maps of the decoder through 1×1 convolution with a channel dimension of 64, that is, skip connection.
3. The method for surface type segmentation of remote sensing images based on global-local transformer blocks according to claim 2, characterized in that: In step 2, the decoder of the transformer architecture consists of two parallel branches constructed by global-local attention to extract global and local context respectively.
4. The method for surface type segmentation of remote sensing images based on global-local transformer blocks according to claim 3, wherein: In step 2, the local branch extracts the local context as follows: Two parallel convolutional layers with kernel sizes of 3 and 1 are employed to extract local context; then, two batch normalization operations are appended before the final summation operation.
5. The method for surface type segmentation of remote sensing images based on global-local transformer blocks according to claim 3, wherein: In step 2, the global branch deploys window-based multi-head self-attention to capture the global context.
6. The method for surface segmentation of remote sensing images based on global-local transformer blocks according to claim 5, characterized in that: In step 2, the global branch extracts the global context as follows: First, a standard 1×1 convolution is used to transform the input 2D feature map ∈ R B×C×H×W The channel dimension of is tripled; then, It is split into query vector (Q), keyword vector (K) and value vector (V), and the channel size C is set to 64. The window size w and the number of headers h are both set to 8.
7. The method for surface segmentation of remote sensing images based on global-local transformer blocks according to claim 1, wherein: In step 3, the channel path uses a global average pooling layer to generate the channel attention map C∈R 1×1×c , where c represents the channel dimension; the reduce & expand operation contains two 1×1 convolutional layers, which first reduce the channel dimension c to 4 times of the original one and then expand it to the original one.
8. The method for surface segmentation of remote sensing images based on global-local transformer blocks according to claim 1, wherein: In step 4, the spatial path uses depth-wise convolution to generate a spatial attention map S∈R h×w×1 , where h and w represent the spatial resolution of the feature map.
9. A method for surface segmentation of remote sensing images based on global-local transformer blocks according to any one of claims 1 to 8, characterized in that: Transformer has a CNN-based encoder and a Transformer-based decoder for effective semantic segmentation of remote sensing urban scene images; The CNN Encoder uses the pre-trained ResNet18 as the encoder. ResNet18 consists of four layers of Resblock. Each layer is downsampled by a scaling factor of 2, and the feature maps generated by each layer are fused with the corresponding feature maps of the decoder through a 1×1 convolution with a channel dimension of 64, i.e., a skip connection. The Transformer-based decoder is a global-local Transformer block decoder, which constructs two parallel branches by global-local attention to extract global and local context respectively; the local branch uses two parallel convolutional layers with kernel sizes of 3 and 1 to extract local context; the global branch deploys window-based multi-head self-attention to capture global context.
10. The method for surface segmentation of remote sensing images based on global-local transformer blocks according to claim 9, characterized in that: The local branch uses two parallel convolutional layers with kernel sizes of 3 and 1 to extract local context; then, two batch normalization operations are appended before the final summation operation; the global branch extracts global context as follows: First, a standard 1×1 convolution is used to transform the input 2D feature map ∈ R B×C×H×W The channel dimension of is tripled; then, It is split into query vector (Q), keyword vector (K) and value vector (V), and the channel size C is set to 64. The window size w and the number of headers h are both set to 8.