A Remote Sensing Image Semantic Segmentation Method Based on Scene-Aware Class Attention

By introducing scene-aware attention submodule and local-global class attention in remote sensing image semantic segmentation, the problem of difficulty in utilizing space correlation and processing complex background noise in the prior art is solved, and more efficient semantic segmentation performance is achieved.

CN115965789BActive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310061100.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-21
Publication Date
2025-06-27
Estimated Expiration
2043-01-21

AI Technical Summary

Technical Problem

The prior art is difficult to fully utilize the spatial correlation of land objects in semantic segmentation of remote sensing images, and the traditional attention mechanism has noise problems when dealing with complex backgrounds and large intra-class variance.

Method used

The scene-aware attention submodule is introduced, and the scene perception of pixels is embedded through context information embedding and position prior embedding, and combined with local-global class attention, improve feature expression capabilities.

Benefits of technology

Effectively utilize the spatial correlation of land objects in remote sensing images to improve the performance of semantic segmentation, reduce background noise interference, and deal with problems with large intra-class variance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965789B_ABST
    Figure CN115965789B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for semantic segmentation of remote sensing images based on scene-aware class attention. Aiming at the characteristics of the inherent spatial correlation of ground objects in high-resolution remote sensing images and the problems such as complex background and large intra-class variance, the present invention generates local class centers and global class centers respectively through a class center generation sub-module, and further uses a scene-aware attention sub-module to embed context information and position prior information into the feature representation of pixels. At the same time, the global class center is indirectly associated by introducing the local class center as an intermediate perception element. The present invention not only utilizes the spatial correlation of ground objects in remote sensing images to strengthen context modeling, but also solves the problems of more background noise and large intra-class variance. By combining scene perception and class-level context aggregation, the present invention provides a new solution for the high-resolution remote sensing image segmentation task and improves the accuracy of semantic segmentation of remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention applies technologies related to the fields of deep learning and computer vision, and specifically invents and applies a high-resolution remote sensing image semantic segmentation method based on scene-aware class-level context aggregation. Background Art

[0002] Semantic segmentation aims to predict the semantic category of each pixel in an image and is one of the fundamental and extremely challenging tasks in remote sensing image analysis. Semantic segmentation plays an important role in fields such as road extraction, urban planning, and environmental detection by providing semantic and localization information for ground objects of interest. Compared with natural images, the ground objects in remote sensing images have inherent spatial correlations, and these correlations are often observable. For example, vehicles usually stay on roads, and buildings are densely distributed on both sides of roads, etc.

[0003] In recent years, due to its powerful feature extraction ability, the Convolutional Neural Network (CNN) has become an important method to promote the development of semantic segmentation effects. However, due to its fixed geometric structure, CNN has natural limitations such as being able to effectively capture only local receptive fields and short-range context information. Therefore, context modeling, including spatial context modeling and relational context modeling, has become an important option for capturing long-range dependencies.

[0004] Spatial context modeling methods, such as PSPNet and DeepLabv3+, respectively use spatial pyramid pooling and multi-scale dilated convolutions to aggregate context information. These methods focus on capturing homogeneous context dependencies but often ignore category differences, which may lead to the introduction of unreliable context if there are confused categories in the image scene.

[0005] Relational context modeling methods adopt an attention mechanism, that is, calculating the similarity at the pixel level in the image to weight and aggregate heterogeneous context information, and have achieved remarkable results in semantic segmentation tasks. However, these methods mainly focus on the relationships between pixels and ignore the perception of pixels for the scene (i.e., global context information and location priors), resulting in insufficient exploration of the spatial correlations of ground objects in remote sensing images.

[0006] Based on this, the present invention first improves the spatial attention mechanism and proposes a scene-aware attention sub-module to utilize the spatial correlation of ground objects in remote sensing images by embedding the scene awareness of pixels. Scene awareness is divided into two parts. One is context information embedding, which is to identify different pairwise relationships between ground objects in different scenes. For example, in urban areas, roads usually coexist with buildings, but in rural areas, it may be surrounded by farmland. The other is position prior embedding, which is to identify the intrinsic pattern distribution followed by ground objects in space. For example, pixels that are close to each other usually show a high correlation, and pixels of the same object usually follow a certain positional relationship.

[0007] In addition, remote sensing images have the characteristics of complex backgrounds and large intra-class variances. The traditional attention mechanism introduces a large amount of background noise due to dense affinity operations, and it is difficult to handle the problem of large intra-class variances in remote sensing images by simply using global class representations. Summary of the Invention

[0008] The technical problem to be solved by the present invention is how to fuse local-global class attention on the basis of constructing position priors and context priors of the scenes where pixels are located, and improve the feature expression ability of each pixel through scene awareness and class-level context aggregation, and provide a semantic segmentation method for remote sensing images based on scene-aware class attention. The present invention associates pixels with global class representations by introducing local-global class attention, and uses local class representations as intermediate perception elements, which greatly reduces the required attention operations while improving the model accuracy.

[0009] The specific technical solution adopted by the present invention is as follows:

[0010] A semantic segmentation method for remote sensing images based on scene-aware class attention, and the specific approach is as follows: Input the remote sensing image to be semantically segmented into a semantic segmentation model composed of an encoder module and a decoder module to obtain a semantic segmentation result;

[0011] In the encoder module, first perform feature extraction through a backbone network, and use the features output by the backbone network as rough feature representations;

[0012] The decoder module includes a class center generation sub-module (CCG) and a scene-aware attention sub-module (SSA). The decoder module takes the rough feature representation output by the encoder as input. When the decoder works, first, it performs a pre-classification operation on the rough feature representation output by the encoder to obtain a global class probability distribution. Then, it takes the rough feature representation and the global class probability distribution as inputs to the class center generation sub-module to obtain global class centers. The global class centers are sliced along the spatial dimension to obtain multiple sliced global class center local blocks. At the same time, the decoder module slices the rough feature representation and the global class probability distribution along the spatial dimension respectively to obtain multiple pairs of rough feature representation local blocks and global class probability distribution local blocks of the same size, and inputs each pair of rough feature representation local blocks and global class probability distribution local blocks into the class center generation sub-module to obtain local class centers. Then, it inputs the sliced rough feature representation local blocks, the sliced global class center local blocks, and the local class centers into the scene-aware attention sub-module simultaneously to obtain an enhanced feature representation, and then reassembles the enhanced feature representations of all local blocks according to their positions before slicing to restore the same spatial dimension as the rough feature representation. Finally, it concatenates the rough feature representation and the reassembled enhanced feature representation along the channel direction to obtain an output feature representation, and performs upsampling on the output feature representation to obtain the semantic segmentation result of the input remote sensing image.

[0013] The input of the class center generation sub-module is the global or local feature representation and its corresponding class probability distribution, and the output is the global or local class center. This module first performs an affinity operation on the input class probability distribution and feature representation to obtain class representation information, then performs an Argmax operation on the class representation information to obtain a pre-classification mask, and finally puts the class representation information back to the corresponding pixel positions in the original rough feature representation according to the pre-classification mask, so as to obtain the class center.

[0014] The scene-aware attention sub-module introduces context information embedding and position prior embedding in the attention operation to embed the scene awareness of pixels. In this sub-module, first, it obtains position prior information from the rough feature representation local block through position prior embedding, and at the same time constructs a context diagonal matrix for the rough feature representation local block through context information embedding and contextualizes it. Then, the contextualized feature representation first aggregates local class centers, and then adds them to the position prior information element by element to obtain an affinity matrix. Finally, it aggregates global class centers according to the affinity matrix to obtain an enhanced feature representation after embedding the scene awareness of pixels.

[0015] Preferably, the context information embedding is used to construct a context diagonal matrix so that the attention can be adjusted according to the given context. The specific method is as follows: First, perform parallel global average pooling and max pooling on the input rough feature representation local block through two branches, and use feature mapping twice for the pooling results of the two branches respectively to obtain context vectors. The feature mappings used by the two branches share the same weights. Finally, add the context vectors obtained by the two branches element-wise, output through the Sigmoid function, and convert it into a context diagonal matrix, thereby contextualizing the rough features.

[0016] Preferably, the position prior embedding is used to construct a relative position encoding between pixels to be embedded in the rough feature representation local block, enhancing the sensitivity of the attention to the spatial distribution. The specific method is as follows: First, calculate the relative position offsets between pixels in the horizontal and vertical directions, and select the corresponding trainable vectors in the encoding bucket according to the offsets to obtain the relative position encoding. Finally, aggregate the relative position encoding by the input rough feature representation local block to obtain the position prior information.

[0017] Preferably, the specific calculation algorithm in the scene-aware attention sub-module is as follows:

[0018] First, perform 1×1 convolution on the input rough feature representation local block R l , local class center S l and global class center local block S g in the sub-module respectively to obtain three matrices Q, K, and V, and reshape their dimensions to (B′×hw×C), where B′ is B represents the Batch size of the input semantic segmentation model, H ′ , W′ are the height and width of the rough feature representation respectively, C is the number of feature channels of the rough feature representation, h and w are the height and width of each local block respectively; then construct a relative position encoding r for embedding position prior information into matrix Q. The dimension of r is (hw×hw×C), and the i-th hw×C matrix represents the relative position encoding of pixel i with all other pixels; then multiply the i-th row of matrix Q with the transpose of the relative position encoding r i to obtain the position prior information of pixel i Finally, concatenate the position priors of each pixel along the vertical direction to obtain the position prior information p of the local block, whose dimension is (B′×hw×hw); at the same time, construct a context diagonal matrix c with dimension (B′×C×C) for embedding context information into Q, and after contextualizing matrix Q using the context diagonal matrix c, aggregate matrix K to obtain the similarity matrix Its dimension is (B′×hw×hw); finally, based on the position prior information p and the similarity matrix S, the affinity matrix A = So / tmax(S + p) is calculated, and the matrix V is aggregated according to the affinity matrix A to obtain the enhanced feature representation after embedding pixel scene perception. And its dimension is reshaped to (B′×C×h×w).

[0019] Preferably, the backbone network is the HRNetv2-w32 model, and the pre-trained weights learned on the ImageNet dataset are loaded.

[0020] Preferably, the pre-classification operation is implemented by two consecutive 1×1 convolutions.

[0021] Preferably, the decoder cuts the rough feature representation and the global class center, and the local block sizes are both 4×4.

[0022] Preferably, the semantic segmentation model is pre-trained using the labeled training data before being used for actual semantic segmentation.

[0023] Preferably, the training data needs to be data-augmented, and the loss functions used for training the semantic segmentation model are all cross-entropy losses.

[0024] Preferably, the remote sensing image is a high-resolution remote sensing image with a spatial resolution below 1m.

[0025] The present invention has the following beneficial effects compared with the prior art:

[0026] The present invention discloses an image semantic segmentation method based on scene-aware class-level context aggregation. In view of the intrinsic spatial correlation of ground objects in high-resolution remote sensing images, the present invention embeds scene perception in the attention; and in view of problems such as complex background and large intra-class variance, local-global class attention is introduced. The present invention generates local class centers and global class centers through the class center generation sub-module, and designs a scene-aware attention sub-module to embed context information and position prior information for pixel feature representation. At the same time, by introducing the local class center as an intermediate perception element to indirectly associate with the global class center, not only the spatial correlation of ground objects in the remote sensing image is utilized to strengthen context modeling, but also the problems of a large amount of background noise and large intra-class variance are solved. The present invention combines scene perception and class-level context aggregation to provide a new solution for the high-resolution remote sensing image segmentation task, and can improve the performance of remote sensing image semantic segmentation. Description of the Drawings

[0027] Figure 1 It is the structure diagram of the SACANet model;

[0028] Figure 2Schematic diagram of the class center generation sub-module;

[0029] Figure 3 Schematic diagram of local-global class attention embedded with scene perception;

[0030] Figure 4 Flowchart of training and testing of the SACANet model in an embodiment of the present invention;

[0031] Figure 5 Test visualization results in an embodiment of the present invention. Detailed implementation manners

[0032] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention in conjunction with the accompanying drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined correspondingly without conflict.

[0033] Due to its ability to model long-range dependencies, the spatial attention mechanism has been widely used in remote sensing image semantic segmentation. Many methods using the spatial attention mechanism aggregate context information by using the direct relationship between pixels in the image, while ignoring the scene perception of pixels (that is, perceiving the global context of the scene where the pixels are located and perceiving their relative positions). Considering that scene perception helps to perform context modeling by using the spatial correlation of ground objects, the present invention designs a scene perception attention sub-module based on an improved spatial attention mechanism embedded with scene perception. In addition, aiming at the problem that the general attention mechanism introduces too much background noise and is difficult to solve the large intra-class variance in remote sensing images, the present invention proposes a local-global class attention mechanism. The core of the present invention is to propose a deep network based on scene perception class-level context aggregation, that is, SACANet. It should be noted that each module in this network has high portability and can be applied to most networks.

[0034] A method for remote sensing image semantic segmentation based on scene perception class attention provided by the present invention is specifically as follows: Input the image to be semantically segmented into the semantic segmentation model SACANet composed of an encoder module and a decoder module to obtain the semantic segmentation result. The image in the present invention is preferably a remote sensing image, and more preferably a high-resolution remote sensing image with a spatial resolution of less than 1m.

[0035] The following will describe the specific structure and principle of the above semantic segmentation model SACANet in detail.

[0036] In the encoder module of the above SACANet, first, feature extraction is performed through the backbone network, and the features output by the backbone network are used as rough feature representations.

[0037] The decoder module of the above SACANet mainly consists of a class center generation sub-module (CCG) and a scene-aware attention sub-module (SAA), and the decoder module uses the rough feature representation output by the encoder as input. When the decoder works, first, the rough feature representation obtained in the decoder is pre-classified to obtain a global class probability distribution, and then the rough feature representation and the global class probability distribution are used as inputs to the class center generation sub-module CCG to obtain global class centers, and the global class centers are cut along the spatial dimension to obtain multiple local blocks of the cut global class centers. Similarly, the decoder also performs the same cut on the rough feature representation and the global class probability distribution along the spatial dimension to obtain the same-sized multiple pairs of local blocks of the rough feature representation and local blocks of the global class probability distribution, and then each pair of corresponding local blocks of the rough feature representation and local blocks of the global class probability distribution are input into the class center generation sub-module CCG to obtain local class centers. Then, the cut local blocks of the rough feature representation, the cut local blocks of the global class centers, and the local class centers are simultaneously input into the scene-aware attention sub-module to obtain enhanced feature representations, and the enhanced feature representations of all local blocks are re-stitched according to the positions before cutting to restore the same spatial dimension as the rough feature representation. Finally, the rough feature representation and the stitched enhanced feature representations are stitched along the channel direction to obtain an output feature representation, and the output feature representation is upsampled to the size of the input remote sensing image to obtain the semantic segmentation result of the input remote sensing image.

[0038] Next, the specific structure of the SACANet model of the present invention will be described in detail. Figure 1 FIG. [FIG ID] is an overall structure diagram of the SACANet model, including an encoder module and a decoder module. The encoder module is used to extract semantic features, while the decoder is used to strengthen the semantic features obtained in the decoder and restore the image spatial resolution by embedding scene-aware local-global class context modeling, including CCG and SSA.

[0039] Please note that the "FIG. [FIG ID]" in ID=8 should be replaced with the actual figure number if available. Since it's not provided in the original text, I left it as is for the translation to be consistent with the original structure.Specifically, for the encoder module, its input is the image to be segmented, with a dimension of (B×3×H×W), where B is the batch size of the input, which depends on the number of samples in each batch during the training phase and can be set to 1 during the prediction phase, and H and W are the height and width of the original image respectively. First, in this embodiment, the HRNetv2-w32 model is used as the backbone network, and the HRNetv2-w32 is loaded with the pre-trained weights learned on the Image-Net dataset. The image to be segmented is input into the backbone network for feature extraction to obtain a relatively rough feature representation R, with a dimension of (B×C×H′×W′), where C is the number of feature channels of the rough feature representation, and H ′ and W′ are respectively and

[0040] When the decoder works, first, the rough feature representation R obtained in the decoder is pre-classified through two consecutive 1×1 convolutions to obtain the global class probability distribution <, with a dimension of (B×K×H′×W′), where K is the number of classes. Then, the rough feature representation R and the global class probability distribution < are simultaneously input into the CCG to obtain the global class center S, with a dimension of (B×C×H′×W′), and it is cut along the spatial dimension to obtain its local blocks S g , with a dimension of where h and w are the length and width of the local block. Similarly, the decoder cuts the rough feature representation R and the global class probability distribution < along the spatial dimension to obtain multiple local blocks R of the same size l and < l , and the corresponding R l and < l are input into the CCG to obtain the local class center S l , where R l and S l have a dimension of < l has a dimension of In this embodiment, the sizes of the local blocks obtained by cutting the rough feature representation and the global class center by the decoder can both be set to 4×4. Then, the cut rough feature representation R l , the cut global class center S g and the local class center S l are simultaneously input into the SSA to obtain the enhanced feature representation R a , and it is restored to the original spatial dimension (B×C×H′×W′). Finally, the rough feature representation and the enhanced feature representation are concatenated to obtain the output feature representation, and it is upsampled by a factor of four to obtain the semantic segmentation result of the input remote sensing image.

[0041] The class center generation sub-module (CCG) in the present invention is applied twice in the decoder, respectively for generating the global class center S and the local class center S l , and the same CCG module can be reused for both times of generating class centers. The input of the class center generation sub-module is the global or local feature representation and its corresponding class probability distribution, and the output is the global or local class center. Whether using the global or local feature representation and class probability distribution as the input, the process of generating class centers in this module is the same. Specifically: First, perform an affinity operation on the input class probability distribution and feature representation to obtain class representation information, then perform an argmax operation on the class representation information to obtain a pre-classification mask, and finally, according to the pre-classification mask, place the class representation information back to the corresponding pixel positions in the original rough feature representation, so as to obtain the class center. If the input is the global feature representation and class probability distribution, the output is also the global class center; if the input is the local feature representation and class probability distribution, the output is also the local class center.

[0042] In this embodiment, for the class center generation sub-module CCG, its main purpose is to replace the feature representation of pixels with a large amount of background noise with class representations that are richer in semantic information. As Figure 2 shown, taking the generation of the global class center as an example, the specific approach in CCG is: Reshape the dimensions of the input global class probability distribution < and the rough feature representation R to (B×K×N) and (B×C×N), where N is H×W. Then, along the channel direction, perform an affinity operation on the transposed matrices of < and R to obtain the global class representation information C g , that is with the dimension of (B×K×C). Each C-dimensional vector in C g is the feature representation of the corresponding class. To obtain the class representation to which each pixel belongs, calculate the pre-classification mask E = Argmax(C g ) along the channel direction, whose dimension is (B×1×H′×W′). The value of each pixel in this mask represents the subscript of the class representation information to which it belongs. Finally, according to the mask information, place the class feature representation back to the position of each pixel to obtain the global class center S, whose dimension is (B×C×H′×W′). The process of generating the local class center is the same as the above process, only need to change the input and output as well as the dimensions of each variable, that is, replace the input global class probability distribution < and the rough feature representation R with < l and R l , and the output becomes the local class center S l , and the dimensions of the remaining variables change accordingly.

[0043] The scene-aware attention submodule (SAA) in this invention introduces context information embedding and position prior embedding in the conventional attention operation to embed the scene perception of pixels. In addition, different from the general self-attention operation, this module introduces the local class center S l To indirectly associate the pixel feature representation R l and the global class center S g , which solves the problem of complex background and large intra-class variance in remote sensing images. In this submodule, the position prior information is first obtained by embedding the local block of the rough feature representation through the position prior, and the context information is embedded to construct the context diagonal matrix for the local block of the rough feature representation and contextualize it; then the contextualized feature representation first aggregates the local class center, and then adds the position prior information element by element to obtain the affinity matrix; finally, the global class center is aggregated according to the affinity matrix to obtain the enhanced feature representation after embedding the pixel scene perception.

[0044] In this embodiment, the scene perception attention submodule SSA is used to represent the local feature R l Embed context information and location prior information, and at the same time, introduce the local class center S l As an intermediate perception element to indirectly associate the global class center S g , thereby obtaining the enhanced feature representation R after embedding scene perception a .like Figure 3 As shown, in this embodiment, the scene perception attention submodule SSA uses an improved attention operation, and the specific method is as follows: first, R l , S l and S g Perform 1×1 convolution to obtain three matrices Q, K, and V, and reshape them into (B′×hw×C), where B′ is Then embed the position prior information into Q, that is, construct the relative position code r, whose dimension is (hw×hw×C), where the i-th hw×C matrix represents the relative position code of pixel i and all other pixels. i The transpose of the matrix multiplication is used to obtain the position prior information p of pixel i. i ,Right now Finally, we concatenate the position priors of each pixel in the vertical direction to obtain the global position prior information p of the local block, whose dimension is (B′×hw×hw). To embed the context information of Q, we also need to construct a context diagonal matrix c, whose dimension is (B′×C×C). After contextualizing Q using the context diagonal matrix c, we aggregate K to obtain the similarity matrix S, that is, Its dimension is (B′×hw×hw). Different from general attention, the affinity matrix in scene-aware attention simultaneously considers the relative positions between pixels and the similarity between pixel features, that is, the affinity matrix A = Softmax(S + p). Finally, aggregate V according to the affinity matrix A to obtain the feature representation R after embedding pixel scene awareness. a , that is and reshape its dimension to (B′×C×h×w). It should be noted that when reshaping the dimension in the present invention, functions such as Reshape in the Pytorch framework can be used to achieve it.

[0045] The position prior embedding used in this embodiment is to enable pixels to perceive the inherent distribution pattern followed by ground objects in the remote sensing image in space. Its focus is to construct the relative position encoding between pixels and embed it in the rough feature representation local block to enhance the sensitivity of attention to spatial distribution. The specific approach is as follows: First, calculate the relative position offsets between pixels in the horizontal and vertical directions, and select the corresponding trainable vectors in the encoding bucket according to the offset to obtain the relative position encoding. Finally, aggregate the relative position encoding by the input rough feature representation local block to obtain the position prior information. The present invention comprehensively considers the relative positions in the horizontal and vertical positions. Taking pixel i and pixel I as examples, their relative position encoding can be defined as where P is the encoding bucket storing a set of trainable vectors, and its dimension is ((2ξ + 1)×(2ξ + 1)×C), I x (i,j) = g(x i -x K ) and I y (i,j) = g(y i -X K ) are the offsets in the two directions respectively. represents selecting the corresponding encoding vector from the encoding bucket according to the offset. At the same time, in order to reduce the number of parameters and computational cost required for semantic segmentation of high-resolution remote sensing images, the present invention limits the offset within the maximum distance ξ, that is, uses the clipping function g(x) = 1ax(-ξ, 1in(x, ξ)) to map the offset to a finite set. According to this method, the dimension of the relative position encoding r finally obtained by the present invention is (hw×hw×C), where the vector r iK in the i-th row and j-th column represents the relative position encoding between pixel i and pixel j. Based on the relative position encoding r i aggregate to the matrix Q to form the position prior information p i The specific calculation formula is as described in the previous paragraph and will not be elaborated here.

[0046] The context information embedding used in this embodiment is to enable pixels to perceive the pairwise relationships between ground objects in different scenes in remote sensing images. The key point is to construct the context diagonal matrix c. The specific approach in context information embedding is as follows: First, perform parallel global average pooling and max pooling on the input rough feature representation local blocks through two branches, and use feature mapping twice for the pooling results of each branch to obtain context vectors. The feature mappings used by the two branches share the same weights. Finally, add the context vectors obtained by each branch element-wise, output through the Sigmoid function, and convert it into a context diagonal matrix, thereby contextualizing the rough features. In this embodiment, the specific calculation method of this context diagonal matrix can be expressed by the following formula: c = diag(σ(W1(W0(AvgPool(Q))) + w1(W0(EaxPool(Q)))), where σ is the sigmoid function, and are two feature mappings, and diag() maps a one-dimensional vector to the corresponding diagonal matrix, with its dimension being (B′×C×C).

[0047] It should be noted that before the above semantic segmentation model SACANet is used for actual semantic segmentation, it is pre-trained using the labeled training data. To expand the training samples, data augmentation can be performed on the training data. The loss functions used in the training of the semantic segmentation model are all cross-entropy losses. The specific training process can refer to the existing semantic segmentation model training methods and will not be elaborated here.

[0048] Next, the above remote sensing image semantic segmentation method based on scene-aware class attention will be applied to a specific embodiment to demonstrate the technical effects it can achieve.

[0049] Embodiment

[0050] The semantic segmentation model SACANet adopted in this embodiment has the specific network structure as described above and will not be elaborated here. As Figure 4 shown, the overall process of semantic segmentation of remote sensing images can be divided into three stages: data preprocessing, model training, and image prediction.

[0051] 1. Data preprocessing stage

[0052] For the obtained original remote sensing image (taking the LoveDA dataset as an example in this embodiment), perform image preprocessing. First, cut the image into a size of 512×512, and then perform operations such as random rotation and flipping on the cut image for data augmentation

[0053] 2. Model training

[0054] Step 1: Construct the training set data and batch the training data set according to a fixed batch size, with a total of N.

[0055] Step 2: Sequentially select a batch of training samples with index i from the training data set, where i ∈ {0, 1, …, N}. Use each batch of training samples to train the semantic segmentation model SACANet. During the training process, calculate the cross-entropy loss function for each training sample and adjust the network parameters in the entire model according to the total loss of all training samples in the batch until all batches of the training data set have participated in the model training. After reaching the specified number of iterations, the model converges and the training is completed.

[0056] 3. Image Prediction

[0057] Directly use the images in the test set as input through the trained semantic segmentation model SACANet. Finally, predict a probability vector for each pixel class, and select the class with the highest probability as the final result output through activation functions such as Sigmoid, thereby realizing semantic segmentation.

[0058] In this embodiment, the test visualization results are as Figure 5 shown, and the test quantification results are shown in Table 1:

[0059] Table 1 Test Quantification Results

[0060] Dataset Back Buil Road Water Barren Forest Agri mIoU LoveDA 47.6 59.1 58.4 80.5 17.8 46.7 67.1 53.9

[0061] As can be seen from Figure 5 and Table 1, the semantic segmentation model SACANet of the present invention can well process the segmentation results for remote sensing images. Relying on the improved attention module embedded with scene perception, it fully exploits the spatial correlation of ground objects in remote sensing images. At the same time, by introducing local-global attention, it effectively alleviates the problems of complex background noise interference and large intra-class variance, thereby improving the segmentation performance of remote sensing images and providing a new solution for the application of context modeling in the field of remote sensing image segmentation.

[0062] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A remote sensing image semantic segmentation method based on scene-aware class attention, characterized in that: Input the remote sensing image to be semantically segmented into a semantic segmentation model composed of an encoder module and a decoder module to obtain the semantic segmentation result; In the encoder module, first perform feature extraction through a backbone network, and use the features output by the backbone network as the rough feature representation; The decoder module includes a class center generation sub-module and a scene-aware attention sub-module. The decoder module takes the rough feature representation output by the encoder as the input. When the decoder works, first perform a pre-classification operation on the rough feature representation output by the encoder to obtain the global class probability distribution, and then use the rough feature representation and the global class probability distribution together as the input of the class center generation sub-module to obtain the global class center. Cut the global class center along the spatial dimension to obtain multiple cut global class center local blocks. At the same time, the decoder module performs the same cut on the rough feature representation and the global class probability distribution along the spatial dimension respectively to obtain the same-sized multiple pairs of rough feature representation local blocks and global class probability distribution local blocks, and input each pair of rough feature representation local blocks and global class probability distribution local blocks into the class center generation sub-module to obtain local class centers. Then, input the cut rough feature representation local blocks, the cut global class center local blocks, and the local class centers into the scene-aware attention sub-module at the same time to obtain the enhanced feature representation, and then splice and restore the enhanced feature representations of all local blocks to the same spatial dimension as the rough feature representation according to the position before cutting. Finally, splice the rough feature representation and the spliced enhanced feature representation along the channel direction to obtain the output feature representation, and perform upsampling on the output feature representation to obtain the semantic segmentation result of the input remote sensing image; The input of the class center generation sub-module is the global or local feature representation and its corresponding class probability distribution, and the output is the global or local class center. This module first performs an affinity operation on the input class probability distribution and feature representation to obtain the class representation information, then performs an Argmax operation on the class representation information to obtain the pre-classification mask, and finally puts the class representation information back to the corresponding pixel positions in the original rough feature representation according to the pre-classification mask, so as to obtain the class center; The scene-aware attention sub-module introduces context information embedding and position prior embedding in the attention operation to embed the scene awareness of pixels; In this sub-module, first obtain the position prior information through the position prior embedding according to the rough feature representation local block, and at the same time construct a context diagonal matrix for the rough feature representation local block through the context information embedding and contextualize it; then the contextualized feature representation first aggregates the local class centers, and then adds the position prior information element by element to obtain the affinity matrix; finally, aggregate the global class centers according to the affinity matrix to obtain the enhanced feature representation after embedding the scene awareness of pixels.

2. The remote sensing image semantic segmentation method based on scene-aware class attention according to claim 1, wherein The context information embedding is used to construct a context diagonal matrix so that the attention can be adjusted according to the given context. The specific method is as follows: First, global average pooling and max pooling are performed on the input rough feature representation local blocks in parallel through two branches, and two feature mappings are respectively used for the pooling results of the two branches to obtain context vectors. The feature mappings used by the two branches share the same weights. Finally, the context vectors obtained by the two branches are added element-wise, output through the Sigmoid function, and converted into a context diagonal matrix, thereby contextualizing the rough features.

3. The remote sensing image semantic segmentation method based on scene-aware class attention according to claim 1, characterized in that, The position prior embedding is used to construct the relative position encoding between pixels to be embedded in the rough feature representation local blocks, enhancing the sensitivity of the attention to the spatial distribution. The specific method is as follows: First, the relative position offsets between pixels are calculated in the horizontal and vertical directions, and the corresponding trainable vectors are selected in the encoding bins according to the offsets to obtain the relative position encoding. Finally, the relative position encoding is aggregated by the input rough feature representation local blocks to obtain the position prior information.

4. The remote sensing image semantic segmentation method based on scene-aware class attention according to claim 1, characterized in that, The specific calculation algorithm in the scene-aware attention sub-module is as follows: First, locally block the rough feature representation R input in the sub-module l , the local class center S l and the local block of the global class center S g are respectively subjected to 1×1 convolution to obtain three matrices Q, K, and V, and their dimensions are reshaped to (B′×hw×C), where B′ is B represents the Batch size of the input semantic segmentation model, H ′ , W′ are respectively the height and width of the rough feature representation, C is the number of feature channels of the rough feature representation, h and w are respectively the height and width of each local block; then a relative position encoding r for embedding position prior information for the matrix Q is constructed, the dimension of r is (hw×hw×C) and the i-th hw×C matrix represents the relative position encoding of pixel i with all other pixels; Then, perform matrix multiplication on the \(i\)-th row of matrix \(Q\) and the transpose of the relative position encoding \(r\) of pixel \(i\) to obtain the position prior information of pixel \(i\). i Finally, concatenate the position priors of each pixel along the vertical direction to obtain the position prior information \(p\) of the local patch, whose dimension is \((B' \times hw \times hw)\); meanwhile, embed context information into \(Q\), construct a context diagonal matrix \(c\) with dimension \((B' \times C \times C)\), after contextualizing matrix \(Q\) using the context diagonal matrix \(c\), then aggregate matrix \(K\) to obtain the similarity matrix. Its dimension is \((B' \times hw \times hw)\); finally, calculate the affinity matrix \(A = Sof0max(S + p)\) based on the position prior information \(p\) and the similarity matrix \(S\), and aggregate matrix \(V\) according to the affinity matrix \(A\) to obtain the enhanced feature representation after embedding pixel scene perception. And reshape its dimension to \((B' \times C \times h \times w)\).​ 5. The remote sensing image semantic segmentation method based on scene-aware class attention according to claim 1, wherein The backbone network is the HRNetv2-w32 model, and the pre-trained weights learned on the ImageNet dataset are loaded.

6. The method for remote sensing image semantic segmentation based on scene-aware class attention according to claim 1, wherein, The pre-classification operation is implemented by two consecutive 1×1 convolutions.

7. The remote sensing image semantic segmentation method based on scene-aware class attention according to claim 1, characterized in that, The local blocks obtained by cutting the rough feature representation and the global class center by the decoder both have a size of 4×4.

8. The remote sensing image semantic segmentation method based on scene-aware class attention according to claim 1, characterized in that Before being used for actual semantic segmentation, the semantic segmentation model is pre-trained using the labeled training data.

9. The method for remote sensing image semantic segmentation based on scene-aware class attention according to claim 8, wherein The training data needs to be data-augmented, and the loss functions used for training the semantic segmentation model are all cross-entropy losses.

10. The remote sensing image semantic segmentation method based on scene-aware class attention according to claim 1, characterized in that, The remote sensing image is a high-resolution remote sensing image with a spatial resolution of less than 1m.