Remote sensing image geographic information interpretation method based on DNN and AM

By introducing deep neural networks and attention mechanisms into remote sensing image processing, the problem of noise distinction and feature extraction in semantic segmentation of hyperspectral and PolSAR images is solved, and more accurate semantic segmentation and spatial position information provision are achieved.

CN119992331APending Publication Date: 2025-05-13CHINA TELECOM SICHUAN BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510100322.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively segment hyperspectral images and synthetic aperture radar images, especially when processing spatial and spectral channel features, it is difficult to distinguish noise from effective physical areas, affecting the semantic accuracy of the features.

Method used

The geographic information interpretation method of remote sensing image based on deep neural network and attention mechanism is adopted, and the spectral spatial information of high-spectral images is extracted through data preprocessing, encoder-decoder transformer framework and multi-source fusion attention mechanism, and the spectral spatial information of high-spectral images is extracted, and the multi-axis sequence attention segmentation network is used in PolSAR images for semantic segmentation.

Benefits of technology

More accurate semantic segmentation of hyperspectral images and PolSAR images is achieved, which can effectively distinguish the importance of different spectral channels, reduce noise interference, distinguish nuances of ground features, and provide more accurate spatial position information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005254019790000061
    Figure BDA0005254019790000061
  • Figure BDA0005254019790000071
    Figure BDA0005254019790000071
  • Figure BDA0005254019790000072
    Figure BDA0005254019790000072
Patent Text Reader

Abstract

The invention discloses a remote sensing image geographic information interpretation method based on DNN and AM. The method comprises the following steps: preprocessing data; inputting the data into a corresponding deep neural network, extracting high-dimensional features of an image at an encoder part, and restoring semantic information of a target in the image at a decoder part; an attention mechanism is utilized in a feature transmission process, so that a neural network automatically pays attention to key features of a corresponding ground target, high-value spectrum channels are screened, and different types of radar image scattering noise are distinguished; in the decoding stage of the deep neural network, segmentation of a ground target is realized, and key information of a remote sensing image is interpreted; training a neural network, and enabling the recognition of the network on the ground target to gradually fit a real ground category; and deploying an algorithm to the reasoning server, opening an API interface, and providing a data receiving and sending disconnection function. The method has remarkable performance advantages in the aspects of hyperspectral image classification and synthetic aperture radar image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of remote sensing image interpretation, and specifically to a method and system for interpreting geographic information of remote sensing images based on deep neural networks and attention mechanisms. Background Art

[0002] Hyperspectral imaging (HSI) is a technique that analyzes a wide spectrum of light and can generate a variety of spectral information. It is widely used in various ground monitoring fields such as geology, agriculture, forestry, and environment. The classification of hyperspectral images (HSI) assigns a unique semantic label to each pixel and is an important challenge in hyperspectral remote sensing (HRS).

[0003] Synthetic aperture radar (PolSAR) imaging is a satellite remote sensing microwave imaging technology that has been widely used in disaster monitoring, wetland monitoring, ocean monitoring, land cover classification and other fields. Therefore, the semantic segmentation of PolSAR images has important research value. Compared with visible light remote sensing, PolSAR has higher penetration ability, can penetrate clouds, and is not affected by day-night changes and severe weather conditions.

[0004] Hyperspectral image classification can use deep learning algorithms such as convolutional neural networks (CNN), 3D convolutional neural networks (3D-CNN), recurrent neural networks (RNN) and multi-scale convolutional neural networks. RNN can also be used to model the dependencies between different spectral bands, and then convolutional recurrent neural networks (CRNN) can be used to learn more discriminative features for hyperspectral image classification. Recently, a deeper neural network has been proposed to explore local contextual information for hyperspectral image classification.

[0005] With the application of deep neural networks in image processing, they have also shown impressive capabilities in PolSAR images. Convolutional neural networks (CNNs) are the most commonly used model structure in deep learning for visual processing, extracting image features through convolution to handle downstream tasks. However, for semantic segmentation of PolSAR images, the feed-forward network randomly initializes weights, making it impossible to distinguish between noise, background, or valid physical areas during pixel-level feature extraction, thus affecting the semantic accuracy of the features.

[0006] In order to direct the network's attention to valuable areas, the Transformer framework based on the self-attention mechanism in natural language processing has been introduced into visual tasks, such as ViT and Swin-Transformer. Transformer can establish long-range dependencies, provide global relationships of features, and provide position and channel information in the target space. It pays more attention to specific targets rather than noise information during feature extraction, and has certain advantages in various PolSAR applications.

[0007] However, the aforementioned CNN-based methods are not sufficient to simultaneously encode the features of spatial and spectral channels in hyperspectral images. On the one hand, the deep spatial spectral information of hyperspectral images is difficult to simply describe and extract; on the other hand, CNN extracts features through local convolution operations, large convolution kernels will lose local details, while small convolution kernels lack a global view. Therefore, existing methods cannot model spatial context information and spectral channel information simultaneously in a targeted manner. At the same time, in the semantic segmentation of PolSAR images, due to the inevitable presence of speckle noise, pixel-level feature classification relies more on the contextual relationship between local adjacent pixels. In addition, the segmentation accuracy is also affected by the spatial position information. Summary of the invention

[0008] The present invention aims to provide a remote sensing image geographic information interpretation method based on deep neural network and attention mechanism to solve the shortcomings of the existing technology.

[0009] In order to achieve the above object, the present invention adopts the following technical solutions:

[0010] A remote sensing image geographic information interpretation method based on deep neural network (DNN) and attention mechanism (AM) is used to identify ground targets described in remote sensing images of two remote sensing data types, including the following steps:

[0011] The first step is data preprocessing. For hyperspectral images, the data is sliced ​​and divided into the smallest processing unit. For synthetic aperture radar images, the data is denoised, pseudo-color images are synthesized, and the smallest processing units are divided.

[0012] The second step is to input the data into the corresponding deep neural network, extract the high-dimensional features of the image in the encoder part, and restore the semantic information of the target in the image in the decoder part;

[0013] The third step is to use the attention mechanism in the feature transfer process to allow the neural network to automatically focus on the key features of the corresponding ground targets, screen high-value spectral channels, and distinguish different types of radar image scattering noise;

[0014] The fourth step is to segment ground targets and interpret key information of remote sensing images in the decoding stage of the deep neural network.

[0015] The fifth step is to train the neural network so that the network's recognition of ground targets gradually fits the real ground categories;

[0016] The sixth step is to deploy the algorithm to the inference server, open the API interface, and provide data receiving and sending disconnection functions.

[0017] In the first step, the pseudo-color synthesis method of synthetic aperture radar images includes single-polarization and multi-polarization SAR image preprocessing, which is used to reduce noise interference as much as possible and increase image readability before performing the segmentation task.

[0018] In the second step, for the hyperspectral image, the encoder includes a Transformer block and a hierarchical Transformer framework for encoding long-range context information. Further, each Transformer block includes a LayerNorm (LN) layer, a multi-source fusion attention mechanism, a residual connection, and an MLP. The hyperspectral image classification is implemented using an encoder-decoder transformer framework, and the spectral spatial information of the hyperspectral image is effectively extracted and utilized by combining the multi-source fusion attention mechanism with the Transformer framework. The encoder-decoder transformer is used for the encoder part, including two consecutive transformer blocks, and the output of the previous transformer block is used as the input of the next transformer block; the hierarchical Transformer framework based on the Transformer block is used as an encoder to generate hierarchical feature representations and obtain features of different scales; in the decoder, bilinear interpolation is used to perform upsampling operations to restore the spatial resolution of the feature map; and skip connections and convolution blocks are used to establish global dependencies between features of different scales.

[0019] In the second step, for synthetic aperture radar images, the PolSAR image feature extractor is used to serialize the feature map along two axes, and global and local attention is calculated to extract important spatial feature information and reduce noise interference. Specifically, for synthetic aperture radar images, an encoder-decoder structure is deployed using a multi-axis gated attention network to achieve effective semantic segmentation; the encoder consists of multi-level feature extraction, and the feature extractor is designed and implemented by stacking multi-axis sequence attention blocks to effectively extract PolSAR features at multiple scales, while mitigating the inter-class similarity and intra-class differences of speckle noise; at the same time, the serialized residual connection design process enables spatial information to propagate throughout the network, thereby improving the overall spatial perception ability of MASA-SegNet; the decoder consists of convolution and linear interpolation upsampling to complete the semantic segmentation task.

[0020] In the third step, for hyperspectral images, a multi-source fusion attention mechanism is used to comprehensively consider spectral channels and spatial attention to help the network achieve better classification; each pixel in the hyperspectral image is associated with pixels in adjacent local areas and remote context information in the entire image; the multi-source fusion attention mechanism consists of a spatial attention module and a spectral channel attention module, the spatial attention module is used to encode the spatial context information of the hyperspectral image; the spectral channel attention module is used to model the channel information and find important channels for the feature detector.

[0021] In the third step, the fusion method of the multi-source fusion attention mechanism is: directly adding the obtained spatial attention and spectral channel attention scores; or connecting the obtained spatial attention feature map and the spectral attention feature map in the channel dimension, and then processing them using a multi-layer perceptron (MLP); or applying the Hadamard product evaluation between the spatial attention matrix and the spectral channel attention matrix.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] (1) In the data preprocessing stage, the present invention proposes two PolSAR image pseudo-color synthesis methods, which are used as preprocessing for single-polarization and multi-polarization SAR images, and can minimize noise interference and increase image readability before performing the segmentation task.

[0024] (2) In the encoding stage, in order to encode long-range context information, the present invention combines the advantages of hierarchical transformers and utilizes the powerful feature extraction capabilities of Transformer blocks and hierarchical Transformer frameworks to obtain both shallow local features and high-level global rich semantic information. At the same time, a feature extractor is designed for PolSAR images, which serializes the feature maps along two axes and calculates global and local attention, thereby extracting important spatial feature information and reducing noise interference.

[0025] (3) In order to better encode the rich spectral spatial information of hyperspectral images, the present invention proposes a multi-source fusion attention mechanism that takes into account spectral channels and spatial attention, which can help the network achieve better classification.

[0026] (4) The present invention can distinguish the importance of different spectral channels to the interpretation of hyperspectral images, reduce the impact of noise on the extraction of high-dimensional feature information from PolSAR images, identify subtle differences in ground features with similar scattering characteristics, and provide sufficient and accurate spatial location information for classification. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solutions and advantages of the implementation methods of the present application clearer, the technical solutions in the implementation methods of the present application will be clearly and completely described below.

[0028] The method for interpreting geographic information of remote sensing images based on DNN and AM of the present invention comprises the following steps:

[0029] The first step is data preprocessing. For hyperspectral images, the data is sliced ​​and divided into the smallest processing units. For synthetic aperture radar images, the data is denoised, pseudo-color images are synthesized, and the data is divided into the smallest processing units.

[0030] In the second step, the data is input into the corresponding deep neural network, the encoder part extracts the high-dimensional features of the image, and the decoder part restores the semantic information of the target in the image.

[0031] The third step is to use the attention mechanism in the feature transfer process to allow the neural network to automatically focus on the key features of the corresponding ground targets, screen high-value spectral channels, and distinguish different types of radar image scattering noise.

[0032] The fourth step is to segment ground targets and interpret the key information of remote sensing images in the decoding stage of the deep neural network.

[0033] The fifth step is to train the neural network so that the network's recognition of ground targets gradually fits the real ground categories.

[0034] The sixth step is to deploy the algorithm to the inference server, open the API interface, and provide data receiving and sending disconnection functions.

[0035] 1. Hyperspectral imaging

[0036] For hyperspectral images, the present application provides an encoder-decoder Transformer framework based on a multi-source fusion attention mechanism, including an encoder part and a decoder part. The multi-source fusion attention mechanism consists of a spatial attention module and a spectral channel attention module.

[0037] 1. Encoder

[0038] In the encoder part, two consecutive transformer blocks are used as encoder transformer blocks, namely the window-based multi-source fusion attention module (W-FA) and the shifted window-based multi-source fusion attention module (SW-FA).

[0039] In the encoder transformer block, the output feature Z of the previous transformer block L-1 Used as input to the current transformer block; Z L-1 Linear normalization (LN) is applied and input to W-FA for spatial attention and spectral channel attention operations.

[0040] The operating equation of spatial attention is as follows:

[0041]

[0042] The operating equation of spectral channel attention is as follows:

[0043]

[0044] Among them, V represents the extracted eigenvalue matrix, Q represents the learned relationship query matrix, K represents the key matrix between pixels, T represents the matrix transpose, C represents the number of spectral channel dimensions, and B represents the relative position deviation.

[0045] The output of W-FA is obtained through a multi-source fusion attention mechanism, and the shifted window attention mechanism is added to the spectral channel attention mechanism in SW-FA to achieve the interaction of spatial and spectral channel information across windows.

[0046] The three fusion methods of the two attention features are feature addition, feature concatenation, and feature multiplication. The formula is as follows:

[0047]

[0048] W-FA output and Z L-1 Perform residual connection to obtain Z^ L , and then after a linear normalization (LN), MLP and residual connection are used to obtain Z L . Similar to W-FA, Z L is the input of SW-FA after linear normalization, and the output of SW-FA is related to Z L Perform residual connection to obtain Z^ L+1 , and finally after another linear normalization (LN), MLP and residual connection are used to obtain the Z output of the encoder transformer block L+1 .

[0049] The entire encoder process can be described as follows:

[0050] Z^Z L =W-FA(LN(Z L-1 ))+Z L-1

[0051] Z L =MLP(LN(Z^ L ))+Z L

[0052] Z^ L+1 =SW-FA(LN(Z L ))+Z L

[0053] Z L+1 =MLP(LN(Z^ L+1 ))+Z L+1

[0054] Specific implementation of the encoder: First, the input HSI is divided into non-overlapping windows, each with a patch size of 4×4, generating H / 4×W / 4 patches, with the feature size of the patch being 4×4×the number of channels, and the feature is considered as a concatenation of the original pixel HSI values. Then, these patches are mapped to a predefined dimension C (set to 96) through a linear layer. In the first stage, these patches are used as the input of the encoder transformer block. The number of patches (H / 4×W / 4) remains unchanged. In the second stage, a patch merging operation is performed to merge the 2×2 patches into one, the number of patches is changed to H / 8×W / 8, and the feature dimension is changed from C to 4C. These merged features are the input of the second encoder transformer block, and the output dimension is set to 2C. 2x downsampling is achieved in this process. The number of patches is reduced by patch merging, thereby generating hierarchical representations at different network depths. In the third and fourth stages, as the number of patches decreases and the channel dimension increases, the number of transformer blocks for feature encoding is increased accordingly.

[0055] In order to balance the depth of the network and the complexity of the model, the number of transformer blocks in each of the four stages is set to = (1, 1, 3, 1). The encoder can obtain deep features and hierarchical representations through four stages of downsampling. The output resolution of each encoder transformer block is H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, H / 32×W / 32 respectively. In the first and second stages of encoding, it is used to learn shallow local features, and in the subsequent encoding stages, it is used to learn deep global features.

[0056] 2. Decoder

[0057] In the decoder, bilinear interpolation is used to perform upsampling operations to restore the spatial resolution of feature maps; skip connections and convolutional blocks are used to establish global dependencies between features of different scales.

[0058] The decoder consists of four stages, each of which includes upsampling (bilinear interpolation), skip connections, and convolution blocks. The convolution blocks consist of 3×3 convolutions followed by a group normalization layer and a ReLU activation function. Specifically, the output of the 4th stage in the encoder is used as the initial input to the decoder. The input features are upsampled by ×2, and the shape changes from H / 32×W / 32×8C to H / 16×W / 16×4C. In order to preserve the shallow feature representation and prevent feature loss, skip connections are used, that is, the output of the encoder stage is added to the upsampled features, and the features of different scales from the encoder are fused during the decoding process. In order to keep the channel size of the multi-scale feature map in the skip connection consistent with the upsampled feature map, a linear layer is used to keep the size consistent. At the end of the decoder, 1×1 convolutions and softmax are used to map the final feature map to each class. In the encoder, the low-level output feature map contains rich global information, while the high-level output feature map containing more information contains rich semantic details and local information. Therefore, the decoder structure can effectively fuse the output feature maps of different stages in the encoder.

[0059] 2. Synthetic Aperture Radar Imagery

[0060] For synthetic aperture radar images, this application provides an encoder-decoder framework of a multi-axis sequence attention segmentation network (MASA-SegNet) for semantic segmentation of PolSAR data. Specifically, in the encoder, a feature extractor is designed and implemented by stacking multi-axis sequence attention blocks to effectively extract PolSAR features at multiple scales while mitigating inter-class similarities and intra-class differences in speckle noise. In addition, the serialized residual connection design process enables spatial information to propagate throughout the network, thereby improving the overall spatial perception capability of MASA-SegNet. In the decoder, the semantic segmentation task is completed.

[0061] Before interpreting the synthetic aperture radar image, the system preprocesses the single-polarization image and the multi-level image. The formula is as follows:

[0062]

[0063] and

[0064]

[0065] Gray value Grey=0.299×C2+0.587×C1+0.114×C3.

[0066] Among them, R, G, and B represent the three color channels of red, green, and blue respectively; P S Represents the pixel value of the single-polarization SAR image; I GRepresents the intensity value of the green channel. C1, C2, C3 represent the values ​​of the red, green and blue channels in the RGB color space respectively; p HH Represents the polarization state of horizontal transmission and horizontal reception; p VV Represents the polarization state of vertical transmission and vertical reception; p HV represents the polarization state of horizontal transmission and vertical reception; p VH Represents the polarization state of vertical transmission and horizontal reception.

[0067] In the encoder-decoder framework for semantic segmentation of PolSAR images, the encoder consists of stacked blocks of feature extractors, which receive low-level and high-level features from the encoder to complete pixel-level semantic segmentation. The PolSAR image is input into the feature extractor, and low-level and high-level features are extracted by the MASA block, which is input into the decoder for the semantic segmentation task of PolSAR images. The MASA block consists of three parts: multi-axis sequence attention (MASA), downsampling, and residual connection.

[0068] (1) Downsampling: Downsampling consists of depthwise separable convolution (DWConv), batch normalization, Gaussian error linear unit (GELU) activation function, and max pooling. The pseudo RGB image x from the PolSAR image with shape (H, W, 3) is used as the input for the downsampling part. Deep separable convolution is used to extract deep features, which are then normalized, activated (GLUE), and Maxpooling is performed to obtain the feature f with shape (h, w, c), which is denoted as: f (h,w,c) = Maxpooling(GND(x)), where GND represents GELU, BatchNormalization, and depthwise separable convolution, h and w are the height and width of the feature map, and c is the number of channels. DWConv with 3×3 convolution not only extracts features such as texture and color from the image, but also implicitly encodes spatial position information, preserving the relative position information of each pixel.

[0069] (2) Multi-axis sequence attention block: Convolutional filters often lack attention capabilities and cannot effectively suppress the impact of noise on feature extraction, thereby weakening the decoding performance of the model. To solve this problem, multi-axis sequence attention (MASA) is introduced into the PolSAR feature extraction task. MASA is used to perform pixel-level attention and region-level attention on the features, and convolution with a kernel size of 1×1 is used to further deepen the features to (h,w,2c). The features are then split into two axes along the channel dimension, and the input feature size of each axis is (h,w,c).

[0070] First, serialize the features along both axes:

[0071] Pixel-level Sequence Attention: On the pixel-level attention axis, features of size (h, w, c) are serialized into a sequence of tensors of shape (h / b×w / b,b×b,c), with each block of pixels being (b×b).

[0072] Region-level sequential attention: On the region-level attention axis, the entire feature map is divided into (q×q) sequential windows of shape (q×q,h / q×w / q,c), and the size of each window is (h / q×w / q).

[0073] b and q represent the predetermined token window sizes, which remain fixed during serialization. Next, the gMLP framework is deployed across multiple axes to compute the internal attention of each token. On the pixel-level axis, the spatial attention between local pixels is computed. Meanwhile, on the region-level axis, the global region spatial attention is computed. Then, the output features of both axes are reshaped into (h, w, c). The information of the upper and lower axes is merged and concatenated in the channel dimension, resulting in a shape of (h, w, 2c). Finally, the features are restored to the original size (h, w, c) through a convolution with a kernel size of 1×1.

[0074] (3) Serial residual connection: In order to prevent network degradation and share shallow information, a residual structure is introduced in the MASA block. In order to map to the channel dimension to match the attention sequence, a serialized residual structure is introduced to convert features into sequences and transfer features to deeper layers without degradation problems. During the serialization process, 1×1 convolution is used to maintain absolute spatial positions, which can better understand the global distribution of land features and objects. This simple design makes the backbone network both stackable and scalable, while also being able to stably preserve spatial location information.

[0075] Specifically, the encoding part of the network is a feature extractor consisting of four layers of MASA blocks. The preprocessed pseudo-color image, with size (H, W, 3), is input to the network. Two layers of MASA blocks are stacked in sequence to produce low-level features of size (h / 4, w / 4, 2c), and c is set to 64. Each MASA block performs 2x downsampling. The next two layers of MASA blocks are stacked to obtain high-level features of size (h / 16, w / 16, 8c), and the encoder pipeline is constructed in this way. It is worth noting that the number of MASA blocks used to extract high-level and low-level features can be set to be variable. In order to facilitate the concatenation of low-level and high-level features, the dimension of the low-level features is reduced from 2c to 48, i.e., (h / 4, w / 4, c), using 1×1 convolution. The high-level features are upsampled by 4 times to match the size of the low-level features, i.e., (h / 4, w / 4, 8c). Then, the high-level features and low-level features are concatenated along the channel dimension, i.e., (h / 4, w / 4, 8c+c). Finally, the (h / 4,w / 4,4c) feature is obtained through 3×3 convolution, and after 4 times upsampling and 1×1 convolution, the classification of each pixel is realized, that is, (H,W,class).

[0076] The above description is only an example of a specific implementation of the present application, and the protection scope of the present application is not limited thereto. Any technical solution that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed in the present application should be included in the protection scope of the present application.

Claims

1. A remote sensing image geographic information interpretation method based on DNN and AM, which is used to identify ground targets described in remote sensing images of two remote sensing data types, characterized in that: The following steps are involved: The first step is data preprocessing. For hyperspectral images, the data is sliced ​​and divided into the smallest processing unit. For synthetic aperture radar images, the data is denoised, pseudo-color images are synthesized, and the smallest processing units are divided. The second step is to input the data into the corresponding deep neural network, extract the high-dimensional features of the image in the encoder part, and restore the semantic information of the target in the image in the decoder part; The third step is to use the attention mechanism in the feature transfer process to allow the neural network to automatically focus on the key features of the corresponding ground targets, screen high-value spectral channels, and distinguish different types of radar image scattering noise; The fourth step is to segment ground targets and interpret key information of remote sensing images in the decoding stage of the deep neural network. The fifth step is to train the neural network so that the network's recognition of ground targets gradually fits the real ground categories; The sixth step is to deploy the algorithm to the inference server, open the API interface, and provide data receiving and sending disconnection functions.

2. According to claim 1, a remote sensing image geographic information interpretation method based on DNN and AM is characterized in that: In the first step, the method of synthesizing a pseudo-color image from an aperture radar image includes single-polarization and multi-polarization SAR image preprocessing, which is used to reduce noise interference as much as possible and increase image readability before performing the segmentation task.

3. The method for interpreting geographic information of remote sensing images based on DNN and AM according to claim 1, characterized in that: In the second step, an encoder-decoder transformer framework is used to implement hyperspectral image classification, wherein the encoder includes a Transformer block and a hierarchical Transformer architecture for encoding long-range context information; the Transformer block includes a LayerNorm layer, a multi-source fusion attention mechanism, a residual connection, and an MLP; by combining the multi-source fusion attention mechanism with the Transformer framework, the spectral spatial information of the hyperspectral image is effectively extracted and utilized.

4. The method for interpreting geographic information of remote sensing images based on DNN and AM according to claim 3 is characterized in that: In the encoder-decoder transformer, the encoder includes two consecutive transformer blocks, and the output of the previous transformer block is used as the input of the next transformer block; a hierarchical Transformer framework based on Transformer blocks is used as an encoder to generate hierarchical feature representations and obtain features of different scales; in the decoder part, bilinear interpolation is used to perform upsampling operations to restore the spatial resolution of feature maps; skip connections and convolution blocks are used to establish global dependencies between features of different scales.

5. The method for interpreting geographic information of remote sensing images based on DNN and AM according to claim 1, characterized in that: In the second step, for synthetic aperture radar images, the PolSAR image feature extractor is used to serialize the feature map along two axes, and calculate global and local attention to extract important spatial feature information and reduce noise interference.

6. The method for interpreting geographic information of remote sensing images based on DNN and AM according to claim 1, characterized in that: For synthetic aperture radar images, an encoder-decoder structure is deployed using a multi-axis gated attention network to achieve effective semantic segmentation; the encoder consists of multi-level feature extraction, and a feature extractor is designed and implemented by stacking multi-axis sequential attention blocks to effectively extract PolSAR features at multiple scales while mitigating inter-class similarities and intra-class differences in speckle noise; at the same time, the serialized residual connection design process enables spatial information to propagate throughout the network, thereby improving the overall spatial perception ability of MASA-SegNet; the decoder consists of convolution and linear interpolation upsampling to complete the semantic segmentation task.

7. The method for interpreting geographic information of remote sensing images based on DNN and AM according to claim 1, characterized in that: For hyperspectral images, a multi-source fusion attention mechanism is used to comprehensively consider spectral channels and spatial attention to help the network achieve better classification; each pixel in the hyperspectral image is associated with pixels in adjacent local areas and remote context information in the entire image; the multi-source fusion attention mechanism consists of a spatial attention module and a spectral channel attention module, the spatial attention module is used to encode the spatial context information of the hyperspectral image; the spectral channel attention module is used to model the channel information and find important channels for the feature detector.

8. The method for interpreting geographic information of remote sensing images based on DNN and AM according to claim 7, characterized in that: The fusion method of the multi-source fusion attention mechanism is: directly adding the obtained spatial attention and spectral channel attention scores; or connecting the obtained spatial attention feature map and spectral attention feature map in the channel dimension, and then processing them using a multi-layer perceptron; or applying the Hadamard product evaluation between the spatial attention matrix and the spectral channel attention matrix.