An Image Semantic Segmentation Method Based on Relational Context Aggregation
By adopting a relational context aggregation method in image semantic segmentation, using a class-level multi-scale context generator and a relationship-level multi-scale context integrator, the problem of background noise interference in complex background remote sensing images is solved, and the segmentation effect is significantly improved.
Patent Information
- Application Number
- CN202211138798.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-09-19
AI Technical Summary
When processing remote sensing images of complex backgrounds, the prior art is difficult to effectively reduce background noise interference, resulting in poor semantic segmentation effect.
The image semantic segmentation method based on relational context aggregation is adopted to extract and aggregate the context information in the image through a class-level multi-scale context generator and a relation-level multi-scale context integrator to enhance the feature expression ability of each pixel.
Effectively reduce background noise interference and improve the performance of image semantic segmentation, especially in image segmentation tasks with complex background and front background imbalance.
Smart Images

Figure CN115512109B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and computer vision, and particularly relates to an image semantic segmentation method based on relational context aggregation. Background Art
[0002] In recent years, with the rapid development of optical imaging remote sensing technology, the quantity and quality of high-resolution remote sensing images have been greatly improved. By analyzing these remote sensing data, land cover information can be obtained, which helps urban management, planning, and detection. Remote sensing image segmentation, which assigns specific semantic classes to each pixel in the image, is a key step in remote sensing data analysis. Since the introduction of the CNN network into the field of semantic segmentation, semantic segmentation has developed rapidly. In recent years, the research on semantic segmentation has mainly focused on two aspects. One is how to improve the encoder structure so that the model can extract more robust feature representations for each pixel. The other is how to model the context so that the network can enhance the feature expression ability of each pixel by encoding the context information into the original feature representation. This is also the technical focus of the present invention.
[0003] Existing context aggregation methods are mainly divided into two categories: multi-scale context aggregation and relational context aggregation. The multi-scale context aggregation method aggregates the context information within a certain spatial range around the pixel. For example, PSPNet uses pyramid spatial pooling to aggregate multi-scale context information, and Deeplab introduces dilated convolution to capture multi-scale context information. When applied to remote sensing images, due to the characteristics of their complex backgrounds, the multi-scale context aggregation will absorb a large amount of background noise within the spatial range, thus affecting the segmentation effect. Relational context aggregation uses the relationships between pixels in the image or the relationships between pixels and category regions as weights for weighted aggregation of information. Regarding the relationships between pixels in the image, DANet proposes spatial self-attention and channel self-attention, and aggregates context information by calculating the spatial correlation and channel correlation between features. CCNet alleviates problems such as high computational complexity and high video memory requirements in the common spatial attention calculation process by introducing cross-shaped attention. Due to the characteristics of complex backgrounds and foreground-background imbalance in remote sensing images, using the relationships between pixels throughout the image for context modeling will cause foreground pixels to be associated with a large number of background pixels and be interfered by a large amount of background, resulting in poor segmentation effects. Regarding the relationships between pixels and category regions, OCRNet proposes dividing pixels into a group of regions and then performing weighted aggregation on the region representations. However, it ignores the significance of pixel representations within the same category, so objects within the category regions also absorb a lot of information from other categories. ISNet proposes semantic-level context information to alleviate this problem, but its semantic-level context information has a single scale and overly relies on the pre-classification process, resulting in poor segmentation performance for images with complex backgrounds.
[0004] Therefore, for the segmentation task of complex background images, designing a context aggregation module with superior performance to improve the performance of image semantic segmentation is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The technical problem to be solved by the present invention is how to make full use of the deep feature information in the deep learning network, enhance the feature expression ability of each pixel through context aggregation operation, and provide an image semantic segmentation method based on relational context aggregation.
[0006] The specific technical solutions adopted by the present invention are as follows:
[0007] A method for image semantic segmentation based on relational context aggregation, wherein the method comprises: inputting an image to be semantically segmented into a semantic segmentation model composed of an encoder module and a decoder module to obtain a semantic segmentation result;
[0008] In the encoder module, firstly, feature extraction is performed through the backbone network to obtain shallow feature representation and deep feature representation, then the deep feature representation is pre-classified to obtain a pre-classified representation, then the deep feature representation and the pre-classified representation are input into a category-level multi-scale context generator to obtain a category-level multi-scale context corresponding to the deep feature representation, then the pre-classified representation and the deep feature representation and their corresponding category-level multi-scale context are input into a relation-level multi-scale context integrator, context information aggregation is performed on the deep feature representation, and the aggregated features are fused with the deep feature representation to obtain an enhanced feature representation, and finally the encoder module outputs the enhanced feature representation, the shallow feature representation and the category-level multi-scale context to the decoder module as the module input of the decoder module;
[0009] The category-level multi-scale context generator takes the deep feature representation and the pre-classification representation as input, first obtains the category probability distribution of each pixel through the pre-classification representation, uses the category probability distribution to obtain the certainty that each pixel belongs to a different category, and then, for different scales, uses the certainty of the feature within the scale range as a weight to weight the feature, thereby obtaining category-level features with different semantics within each scale, and finally concatenates the category-level features within different scales, thereby obtaining the category-level multi-scale feature representation corresponding to each input feature representation;
[0010] The relationship-level multi-scale context integrator takes the pre-classification representation, the deep feature representation, and their corresponding class-level multi-scale feature representations as inputs. First, through the pre-classification representation, it calculates the similarity of the class distribution probabilities between pixels within each class. Then, through the deep feature representation and the class-level multi-scale feature representations, it calculates the similarity between each pixel in the feature representation and the class-level multi-scale feature representations, and weights this similarity by the similarity of the distribution probabilities, thereby obtaining the similarity between the integrated pixels and the class-level multi-scale feature representations. It multiplies this similarity with the class-level multi-scale feature representations and concatenates them in order along the channels with the original deep feature representation, thereby obtaining the enhanced feature representation;
[0011] In the decoder module, first, the enhanced feature representation is upsampled and fused with the shallow feature representation to obtain the fused feature representation. Then, an attention mechanism is applied between the fused feature representation and the class-level multi-scale feature representations to obtain the enhanced fused feature representation. Finally, through an upsampling operation on the enhanced fused feature representation, the semantic segmentation result of the image to be segmented is obtained.
[0012] Preferably, the backbone network is a Resnet-50 model with dilated convolutions and loaded with pre-trained weights learned on the image-net dataset.
[0013] Preferably, the deep feature representation is output from the fourth residual unit of Resnet-50, and the shallow feature representation is output from the first residual unit of Resnet-50.
[0014] Preferably, the pre-classification operation is implemented by two consecutive 1×1 convolutions.
[0015] Preferably, the number of scales in the class-level multi-scale context generator is fixed at 4.
[0016] Preferably, before being used for actual semantic segmentation, the semantic segmentation model is pre-trained using the labeled training data.
[0017] Preferably, the training data needs to be data-augmented.
[0018] Preferably, the loss functions used in the training of the semantic segmentation model are all cross-entropy losses.
[0019] Preferably, in the class-level multi-scale context generator, the difference obtained by subtracting the second-largest probability from the largest probability in the class probability distribution of each pixel point in the pre-classification representation is used as the certainty of the pixel point belonging to the class corresponding to the largest probability.
[0020] Preferably, the image is a remote sensing image.
[0021] The present invention has the following beneficial effects compared with the prior art:
[0022] The present invention discloses an image semantic segmentation method based on relational context aggregation. Aiming at the problems of complex background and large scale variation in high-resolution remote sensing images, based on the relational context aggregation mechanism, a category-level multi-scale feature extraction module and a relational-level multi-scale context aggregation module are introduced in the image semantic segmentation task. Through the category-level multi-scale feature extraction module, the semantic-level context information within each semantic range in the image is effectively extracted as the category expression, and through the relational-level multi-scale context aggregation module, combining the relationships between pixels within the semantics and the relationships between pixels and classes between semantics, dense and accurate context information is constructed for each pixel, thereby enhancing the feature expression ability of the pixels and greatly reducing the interference of background noise. The present invention provides a new solution for the context aggregation mechanism in the segmentation task of complex background images and can improve the performance of the corresponding image semantic segmentation. Brief Description of the Drawings
[0023] Figure 1 It is the structure diagram of the ReCAN model;
[0024] Figure 2 It is the schematic diagram of the category-level multi-scale context generation module;
[0025] Figure 3 It is the schematic diagram of the relational-level multi-scale context integration module;
[0026] Figure 4 It is the training and testing flow chart of the ReCAN model in the embodiment of the present invention;
[0027] Figure 5 It is the test visualization result in the embodiment of the present invention. Detailed Embodiments
[0028] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is provided in conjunction with the accompanying drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined correspondingly without conflict.
[0029] In the description of the present invention, it should be understood that when an element is considered to be "connected" to another element, it can be directly connected to the other element or indirectly connected, i.e., there is an intermediate element. On the contrary, when an element is called "directly" connected to another element, there is no intermediate element.
[0030] Context modeling is an important means for semantic segmentation. Current context modeling methods have also achieved remarkable results in the semantic segmentation of natural images. However, due to the characteristics of large natural scale variation, complex background, and imbalance between foreground and background in remote sensing images, the current context modeling methods result in poor segmentation effects due to absorbing more background noise. Therefore, it is necessary to establish a context modeling method suitable for complex background images. The core of the present invention is precisely to propose a new relationship-based context aggregation network, namely ReCAN. It should be noted that each module in this network has high portability and can be applied to most networks.
[0031] Therefore, a method for image semantic segmentation based on relational context aggregation provided by the present invention is specifically as follows: The image to be semantically segmented is input into the semantic segmentation model ReCAN composed of an encoder module and a decoder module to obtain a semantic segmentation result. The image in the present invention is preferably a remote sensing image.
[0032] In the encoder module, first, feature extraction is performed through a backbone network to obtain shallow feature representations and deep feature representations. Secondly, a pre-classification operation is performed on the deep feature representations to obtain pre-classification representations. Then, the deep feature representations and the pre-classification representations are input into a class-level multi-scale context generator to obtain the class-level multi-scale context corresponding to the deep feature representations. Next, the pre-classification representations, the deep feature representations, and their corresponding class-level multi-scale contexts are input into a relationship-level multi-scale context integrator to aggregate context information for the deep feature representations. After fusing the aggregated features with the deep feature representations, a strengthened feature representation is obtained. Finally, the encoder module outputs the strengthened feature representation, the shallow feature representation, and the class-level multi-scale context to the decoder module as the module input of the decoder module;
[0033] In the encoder module, first, feature extraction is performed through the backbone network to obtain shallow feature representations and deep feature representations. Second, a pre-classification operation is performed on the deep feature representations to obtain pre-classification representations. Then, the deep feature representations and the pre-classification representations are input into the class-level multi-scale context generator (CMCG) to obtain the class-level multi-scale feature representations corresponding to the deep feature representations. Next, the pre-classification representations, the deep feature representations, and their corresponding class-level multi-scale feature representations are input into the relation-level multi-scale context aggregator (RMCA) to aggregate the context information of the pixel representations in the deep feature representations. After fusing the aggregated features with the deep feature representations, enhanced feature representations are obtained. Finally, the enhanced feature representations, the shallow feature representations, and the class-level multi-scale feature representations are input into the decoder module as the module input of the decoder module. Subsequently, the decoder module can fuse the three input feature representations to form enhanced fused feature representations, and finally, through an upsampling operation on the enhanced fused feature representations, the semantic segmentation result of the output image is obtained.
[0034] The specific structure of the ReCAN model of the present invention will be described in detail below. Figure 1 As shown in the overall structure diagram of the ReCAN model, it includes an encoder module and a decoder module. The encoder module is used to extract semantic features, including CMCG and RMCA, while the decoder module is used to restore the image spatial resolution.
[0035] Specifically, for the encoder module, its input is the image to be segmented, with a dimension of (B×3×H×W), where B is the input batch size, which depends on the number of samples in each batch during the training phase and can be set to 1 during the prediction phase, and H and W are the height and width of the original image respectively. First, a Resnet-50 model with dilated convolutions is used as the backbone network, and the Resnet-50 is loaded with the pre-trained weights learned on the ImageNet dataset. The image to be segmented is subjected to feature extraction through the backbone network to obtain shallow feature representations and deep feature representations The dimension of the shallow feature representations is The dimension of the deep feature representations is C is the number of feature channels of the deep feature representations. Among them, the deep feature representations are output from the fourth residual unit of Resnet-50, and the shallow feature representations are output from the first residual unit of Resnet-50. Second, the deep feature representations obtain a pre-classification mask and pre-classification representations after a pre-classification operation implemented by two consecutive 1×1 convolutions K is the number of categories; then, the deep features are represented With pre-classification representation Input CMCG to get deep feature representation Corresponding category-level multi-scale features The dimension is (ε×K×C), ζ is the scale number; finally, the feature representation Pre-classification representation And category-level multi-scale features Input RMCA to get the enhanced feature representation The dimension is
[0036] The category-level multi-scale context generator (CMCG) in the present invention takes deep feature representation and pre-classification representation as input, first obtains the category probability distribution of each pixel through the pre-classification representation, and uses the category probability distribution to obtain the certainty that each pixel belongs to a different category, and then for different scales, the certainty of the feature within each scale range is used as a weight to weight the feature, thereby obtaining category-level features with different semantics within each scale, and finally concatenates the category-level features within different scales to obtain the category-level multi-scale feature representation corresponding to the feature representation of each input.
[0037] In this embodiment, for CMCG, the purpose is to represent the deep features of the input With pre-classification representation Get category-level multi-scale feature representation like Figure 2 As shown, the specific approach in CMCG is: first, through Get the pre-classified representation of each class and deep representation Right now in The dimension is (K×N k ), The dimension is (C×N k ), N kLet \(n_k\) be the number of pixels belonging to category \(k\), and \(i, j\) be the pixel coordinates. Then, for each pixel in the pre-classification representation, the difference between the maximum probability and the second maximum probability in the category probability distribution is taken as the certainty that the pixel belongs to the category corresponding to the maximum probability, and this is used as a weight to multiply with the deep representation of the pixel. The information of these pixels is integrated to obtain the category-level context representation for each category. Specifically, considering the large within-class differences in remote sensing images, the present invention sorts the weights and sets corresponding scales (e.g., the number of scales is 4). The points with weights in the top 25%, 50%, 75%, and 100% within each category are respectively used for information integration. That is, within each category, 5 context representations are obtained according to the method of obtaining the category-level context representation described above. The context representations under all scales and categories are concatenated to obtain the final category-level multi-scale feature representation.
[0038] The relationship-level multi-scale context aggregator (RMCA) in the present invention takes the pre-classification representation, the deep feature representation, and their corresponding category-level multi-scale feature representations as inputs. First, through the pre-classification representation, the similarity of the category distribution probabilities between pixels within each category is calculated. Then, through the deep feature representation and the category-level multi-scale feature representation, the similarity between each pixel in the feature representation and the category-level multi-scale feature representation is calculated, and this similarity is weighted by the similarity of the distribution probabilities, thereby obtaining the similarity between the integrated pixel and the category-level multi-scale feature representation. This similarity is multiplied with the category-level multi-scale feature representation and concatenated with the original deep feature representation along the channel direction to obtain the enhanced feature representation. Along the channel direction to obtain the enhanced feature representation.
[0039] In this embodiment, for RMCA, its purpose is to obtain the enhanced feature representation through the input deep feature representation Pre-classification representation And the category-level multi-scale context representation To obtain the enhanced feature representation Such as Figure 3 As shown, the specific operation in RMCA is as follows: First, using the mask generated in the pre-classification operation, the feature representation of each category And the pre-classification representation Then, for Perform a self-attention operation to obtain its affinity matrix That is Affinity matrix Represents the similarity of the category distribution probabilities between pixels within each category; then, perform a self-attention operation on And To obtain the affinity matrix That is Affinity matrix represents the similarity between each pixel in the feature representation and the class-level multi-scale feature representation; using to perform weighted multiplication to obtain the enhanced affinity matrix which is the affinity matrix represents the similarity between the integrated pixels and the class-level multi-scale feature representation; finally, the affinity matrix is multiplied by the class-level multi-scale feature representation to obtain the enhanced feature representation and the enhanced feature representation of each class is integrated through the mask generated in the pre-classification operation and then concatenated with the original deep feature representation along the channel dimension to obtain the enhanced feature representation which is ψ is a functional function that integrates the enhanced feature representation of each class using the mask.
[0040] In addition, in the decoder module, first, the enhanced feature representation is upsampled and fused with the shallow feature representation to obtain the fused feature representation, then the attention mechanism is applied between the fused feature representation and the class-level multi-scale feature representation to obtain the enhanced fused feature representation, and finally, the semantic segmentation result of the image to be segmented is obtained by performing an upsampling operation on the enhanced fused feature representation.
[0041] Specifically, in this embodiment, for the decoder module, its input is the shallow feature representation the enhanced feature representation and the class-level multi-scale feature representation First, the enhanced feature representation is subjected to a 2-fold upsampling operation and concatenated and fused with the enhanced feature representation extracted from the backbone network in the channel dimension to obtain the fused feature representation which is whose dimension is C1 is the number of feature channels of the fused feature representation; secondly, the fused feature representation and the class-level multi-scale feature representation are subjected to an attention operation to obtain the enhanced fused feature representation which is whose dimension is Finally, the segmented image is obtained by performing a 4-fold upsampling on the enhanced fused feature representation
[0042] It should be noted that before the above semantic segmentation model ReCAN is used for actual semantic segmentation, it is pre-trained using the labeled training data. The loss functions used in the training of the semantic segmentation model are all cross-entropy losses. The specific training process can refer to the training method of the semantic segmentation model in the prior art and will not be elaborated here.
[0043] Next, the above image semantic segmentation method based on relational context aggregation will be applied to a specific embodiment to demonstrate the technical effects it can achieve.
[0044] Embodiment
[0045] The specific network structure of the semantic segmentation model ReCAN used in this embodiment is as described above and will not be elaborated here. The number of scales in CMCG is set to 4. As Figure 4 shown, the overall process of semantic segmentation of remote sensing images can be divided into three stages: data preprocessing, model training, and image prediction, specifically as Figure 5 shown.
[0046] 1. Data preprocessing stage
[0047] For the obtained original remote sensing image (taking the ISPRS Vaihingen dataset as an example in this embodiment), image preprocessing is performed. First, the image is cut into a size of 768×768, and then operations such as random rotation and flipping are performed on the cut image for data augmentation
[0048] 2. Model training
[0049] Step 1, construct the training set data and batch the training dataset according to a fixed batch size, with a total of N.
[0050] Step 2, sequentially select a batch of training samples with index i from the training dataset, where i ∈ {0, 1,..., N}. Use each batch of training samples to train the semantic segmentation model ReCAN. During the training process, calculate the cross-entropy loss function for each training sample and adjust the network parameters in the entire model according to the total loss of all training samples in the batch until all batches of the training dataset have participated in the model training. After reaching the specified number of iterations, the model converges and the training is completed.
[0051] 3. Image prediction
[0052] Directly use the images in the test set as input through the trained semantic segmentation model ReCAN, and finally predict the probability of each pixel category. Select the category with the highest probability as the final result output through activation functions such as Sigmoid, thereby achieving semantic segmentation.
[0053] In this embodiment, the test visualization results are as follows Figure 5 shown, and the test data results are shown in Table 1:
[0054] Table 1 Test data results
[0055] Dataset Imp.Sur. Building Low Veg. Tree Car mIoU AF Vaihingen 81.05 88.57 65.46 75.86 67.44 68.96 80.31
[0056] From Figure 5 and Table 1, it can be seen that the semantic segmentation model ReCAN of the present invention can well process the segmentation results for remote sensing images. Relying on the extraction process of multi-scale class-level context information and the fusion process of the relationship between pixels within a class and the relationship between pixels and classes between classes, it can greatly avoid the interference of the background context of remote sensing images and provide an accurate and dense context information expression for each pixel in the image, thereby improving the segmentation performance of remote sensing images and providing a new solution for the application of context modeling in the field of remote sensing image segmentation.
[0057] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. An image semantic segmentation method based on relational context aggregation, characterized in that: Input the image to be semantically segmented into a semantic segmentation model composed of an encoder module and a decoder module to obtain a semantic segmentation result; In the encoder module, first perform feature extraction through a backbone network to obtain shallow feature representations and deep feature representations. Secondly, perform a pre-classification operation on the deep feature representations to obtain pre-classification representations. Then, input the deep feature representations and the pre-classification representations into a class-level multi-scale context generator to obtain the class-level multi-scale context corresponding to the deep feature representations. Next, input the pre-classification representations, the deep feature representations, and their corresponding class-level multi-scale contexts into a relation-level multi-scale context integrator to aggregate the context information of the deep feature representations. After fusing the aggregated features with the deep feature representations, obtain enhanced feature representations. Finally, the encoder module outputs the enhanced feature representations, the shallow feature representations, and the class-level multi-scale contexts to the decoder module as the module input of the decoder module; The class-level multi-scale context generator takes the deep feature representations and the pre-classification representations as inputs. First, obtain the class probability distribution of each pixel through the pre-classification representations, and use this class probability distribution to obtain the certainty of each pixel belonging to different classes. Then, for different scales, use the certainty of the features within each scale range as weights to weight the features within this scale range, so as to obtain the class-level features of different semantics within each scale. Finally, splice the class-level features within different scales to obtain the class-level multi-scale feature representations corresponding to each input feature representation; The relation-level multi-scale context integrator takes the pre-classification representations, the deep feature representations, and their corresponding class-level multi-scale feature representations as inputs. First, through the pre-classification representations, calculate the similarity of the class distribution probabilities between the pixels within each class. Then, through the deep feature representations and the class-level multi-scale feature representations, calculate the similarity between each pixel in the feature representations and the class-level multi-scale feature representations, and weight this similarity by the similarity of the distribution probabilities, so as to obtain the similarity between the integrated pixels and the class-level multi-scale feature representations. Multiply this similarity by the class-level multi-scale feature representations and concatenate them with the original deep feature representations along the channels in order to obtain enhanced feature representations; In the decoder module, first upsample the enhanced feature representations and fuse them with the shallow feature representations to obtain fused feature representations. Then, apply an attention mechanism between the fused feature representations and the class-level multi-scale feature representations to obtain enhanced fused feature representations. Finally, through an upsampling operation on the enhanced fused feature representations, obtain the semantic segmentation result of the image to be segmented.
2. The method for image semantic segmentation based on relational context aggregation according to claim 1, wherein The backbone network is a Resnet-50 model with dilated convolutions, and the pre-trained weights learned on the image-net dataset are loaded.
3. The method for image semantic segmentation based on relational context aggregation according to claim 2, wherein The deep feature representations are output from the fourth residual unit of Resnet-50, and the shallow feature representations are output from the first residual unit of Resnet-50.
4. The method for image semantic segmentation based on relational context aggregation according to claim 1, wherein The pre-classification operation is implemented by two consecutive 1×1 convolutions.
5. The method for image semantic segmentation based on relational context aggregation according to claim 1, wherein The number of scales in the class-level multi-scale context generator is fixed at 4.
6. The method for image semantic segmentation based on relational context aggregation according to claim 1, wherein Before being used for actual semantic segmentation, the semantic segmentation model is pre-trained using labeled training data.
7. The method for image semantic segmentation based on relational context aggregation according to claim 6, wherein The training data needs to be data-augmented.
8. The method for image semantic segmentation based on relational context aggregation according to claim 6, wherein The loss functions used in the training of the semantic segmentation model are all cross-entropy losses.
9. The method for image semantic segmentation based on relational context aggregation according to claim 1, wherein In the class-level multi-scale context generator, the difference obtained by subtracting the second-largest probability from the largest probability in the class probability distribution of each pixel point in the pre-classification representation is used as the certainty of the pixel point belonging to the class corresponding to the largest probability.
10. The method for image semantic segmentation based on relational context aggregation according to claim 1, wherein The image is a remote sensing image.