Image land coverage classification method based on spatial information enhancement
Through the improved ResNet network and spatial information enhancement method, the problem of insufficient fusion of optical images and SAR images is solved, and higher land cover classification accuracy and robustness are achieved, especially in complex scenarios of land type distinction.
Patent Information
- Application Number
- CN202510552827.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-01
AI Technical Summary
The existing multimodal land cover classification method fails to fully utilize the complementarity between optical images and SAR images, resulting in insufficient feature fusion, difficulty in distinguishing visual similarity categories, and low classification accuracy in complex scenarios.
Using an image land cover classification method based on spatial information enhancement, optical and SAR image features are extracted through the improved ResNet network, combined with multi-scale semantic feature modules and spatial information extraction modules, deep features and long-range spatial dependencies are captured, and dual-branch convolutional neural networks are designed for classification using adaptive fusion module weighted fusion features.
It improves the accuracy and robustness of land cover classification, especially in complex scenarios, which can better distinguish land objects, and improves the adaptability and classification accuracy of the model.
Smart Images

Figure CN120411641A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of land cover classification, and particularly relates to a method for classifying land cover of images based on enhanced spatial information. Background Art
[0002] In Earth observation missions, land cover classification is a core task in remote sensing data analysis and is widely used in crop monitoring, urban planning, and disaster response. Traditional methods rely on single-modal optical images. Although they have high resolution, they are easily affected by clouds and haze, and the imaging quality is unstable. SAR images have all-weather imaging capabilities and can provide surface morphology information, but they have low resolution and are significantly affected by speckle noise. The limitations of single-modal data make them perform poorly in complex scenarios. Multi-modal remote sensing technology provides a solution to this problem. The combination of optical images and SAR images can more comprehensively describe surface information. Optical images capture surface texture and color, while SAR images can effectively obtain surface information under extreme weather conditions. Although LiDAR and social perception data have advantages, their high cost and data quality limit their large-scale application. Therefore, the combination of optical and SAR images has become the focus of research. However, the large differences in the imaging mechanisms of optical and SAR images pose challenges to feature fusion. Existing fusion methods such as pixel-level, decision-level, and feature-level fusion have achieved multi-modal combination, but there are problems such as insufficient feature extraction and a large amount of redundant information. It is still difficult to distinguish between classes with similar visual features. For example, buildings and roads often appear similar in optical images, and traditional methods are difficult to effectively distinguish them. To solve these problems, introducing a spatial information module has become an effective strategy. By capturing the spatial distribution and context relationship of ground objects, the model can establish a clearer boundary between visually similar classes, thereby improving the classification accuracy. For example, the co-occurrence probability of rural areas and forests is higher than that of urban areas and forests. This spatial relationship helps to reduce visual ambiguity and improve the performance of the model. Capturing the spatial relationship between different classes, especially long-range dependence information, is the key to improving the performance of land cover classification.
[0003] The modal land cover classification method refers to the identification of ground object categories through single - type remote sensing data (such as optical images, SAR data, or LiDAR). These data are usually regarded as two - dimensional planes reflecting surface features, and each pixel contains the reflection information of a specific band. The classification process extracts spectral features or texture features from these data and uses different types of machine or deep - learning models for classification. For example, traditional methods include support vector machines, random forests, maximum likelihood classification, etc., which rely on manually designed features for classification. And single - modal methods based on deep - learning methods, such as convolutional neural network CNN and Transformer, classify by automatically learning feature representations and using the spectral information in the data. However, the single - modal method only uses optical images for classification, lacking supplementary information from other data sources. The classification model has limitations when dealing with complex weather scenarios such as cloud cover and rainy days, and performs poorly on ground objects with small spectral differences between different categories, easily leading to classification errors. LI et al. proposed a bilinear fusion network that fuses optical and SAR images. The semantic features of optical and SAR images are respectively extracted through a pseudo - twin convolutional neural network, and then the global average pooling and maximum pooling information are used to generate a channel attention map in a second - order statistical manner to highlight important features. However, this method only focuses on shallow - feature fusion, lacking the utilization of deep - level features and context information, and is vulnerable to image noise and inconsistent data, resulting in a decrease in classification accuracy.
[0004] In summary, in the existing multi - modal land cover classification methods, the complementarity between optical images and SAR images has not been fully utilized. Most methods adopt simple feature stitching or weighted fusion strategies, and such a fusion method will have problems of fusion or loss between feature information. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention proposes an image land cover classification method based on spatial - information enhancement, which includes: acquiring optical images and SAR images; performing enhancement processing on the optical images and SAR images; respectively inputting the enhanced optical images and SAR images into a two - branch convolutional neural network to obtain the land cover classification results; and using a comprehensive evaluation method to evaluate the classification results.
[0006] Training the dual-branch convolutional neural network includes: obtaining an original dataset, where the original dataset contains optical images and SAR images; performing enhancement processing on the images in the original image dataset; respectively inputting the enhanced optical images and SAR images into an improved ResNet network for feature extraction to obtain fusion features with optical image information and SAR image information; inputting the fusion features into a multi-scale semantic feature module to obtain multi-scale features; inputting the multi-scale features into a spatial information extraction module to obtain spatial information; inputting the spatial information into a classifier to obtain a classification result of land cover classification; calculating the loss function of the model according to the land cover classification result, adjusting the parameters of the model, and completing the training of the model when the loss function converges.
[0007] Advantages of the present invention:
[0008] 1. The present invention improves the network structure based on ResNet, fully considering the richness of feature acquisition; the present invention uses multiple adaptive fusion modules to interactively fuse the extracted optical image features and SAR image features, thereby improving the connection between optical images and SAR images. By designing a multi-scale feature extraction module, including atrous spatial pyramid pooling and multi-scale convolutional kernel modules, efficient capture of deep features is achieved. At the same time, low-level features are introduced through skip connections, enabling the network to fuse more surface detail information and combine it with high-level semantic features, ensuring the robustness of the model in complex scenarios.
[0009] 2. The present invention performs weighted fusion on the features of optical and SAR images through a feature attention mechanism. Using two fully connected layers and activation functions, the weights of optical and SAR features are adaptively adjusted according to different types of ground object features, thereby effectively suppressing redundant information in multi-modal data fusion and highlighting the features most helpful for classification.
[0010] 3. The present invention designs a spatial information extraction module to capture long-range spatial dependency relationships through multiple gated loops. The feature map is processed in four directions, enabling the network to more comprehensively obtain the spatial dependency information of surface targets. In addition, local features and global spatial features are combined for modeling to ensure the capture of global spatial relationships in complex scenarios while maintaining spatial details. Description of the Drawings
[0011] Figure 1 It is a structural diagram of the dual-branch convolutional neural network of the present invention;
[0012] Figure 2 It is a diagram of the adaptive fusion module of the present invention;
[0013] Figure 3 It is a diagram of the multi-scale semantic feature module of the present invention;
[0014] Figure 4 Spatial information extraction module diagram of the present invention;
[0015] Figure 5 BRNN spatial relationship extraction model of the present invention. Specific implementation manners
[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] The present invention designs a spatial information extraction module, uses a gated recurrent unit (GRU) to model long-range spatial relationships, processes the feature map in four directions, improves the adaptability to complex scenes, and enhances the classification accuracy. In addition, the present invention also adds a multi-scale semantic feature extraction module. After the high-level feature extraction is completed, atrous spatial pyramid pooling (ASPP) and a multi-scale convolutional kernel module are used to aggregate multi-scale context information, realizing the efficient capture of deep features. Through multi-scale feature extraction and spatial information learning, the model can better capture the global and local features of ground objects, improve the classification accuracy, and has higher robustness especially in complex ground object scenes.
[0018] An image land cover classification method based on spatial information enhancement, as Figure 1 shown, the method includes: obtaining optical images and SAR images; performing enhancement processing on the optical images and SAR images; respectively inputting the enhanced optical images and SAR images into a dual-branch convolutional neural network to obtain a land cover classification result; using a comprehensive evaluation method to evaluate the classification result; training the dual-branch convolutional neural network includes: obtaining an original data set, where the original data set contains optical images and SAR images; performing enhancement processing on the images in the original image data set; respectively inputting the enhanced optical images and SAR images into an improved ResNet network for feature extraction to obtain fused features with optical image information and SAR image information; inputting the fused features into a multi-scale semantic feature module to obtain multi-scale features; inputting the multi-scale features into a spatial information extraction module to obtain spatial information; inputting the spatial information into a classifier to obtain a classified land cover classification result; calculating a loss function of the model according to the land cover classification result, adjusting the parameters of the model, and when the loss function converges, completing the training of the model.
[0019] For the publicly available datasets WHU-OPT-SAR and DFC2020, effective fusion and classification of multimodal remote sensing images are achieved. The entire design method includes an input module (image cropping, cutting, and data augmentation), a neural network module (improved ResNet network), an adaptive fusion module, a multi-scale semantic feature module, a spatial information extraction module, and finally an upsampling output module.
[0020] In this embodiment, during the training phase, two datasets are obtained. Among them, WHU-OPT-SAR contains 38 pairs of SAR and optical images, each image being 5556 pixels wide and 3704 pixels high. For the convenience of experiments, we divide all images into images with a length and width of 256 pixels each, resulting in a total of 29,400 pairs of images; the DFC2020 data contains 6114 pairs of optical and SAR images of size 256x256, which are evenly distributed to the training set and the test set at a ratio of 4:1.
[0021] Data augmentation operations are performed on the datasets, using methods such as translation, horizontal flipping, and vertical flipping respectively to simulate various change scenarios of the same image, increase the diversity of samples, and improve the generalization ability of the model. These augmentation methods effectively expand the scale of the datasets, enabling the model to better adapt to different environmental and ground feature changes.
[0022] The data-augmented optical images and SAR images are respectively input into a dual-branch convolutional neural network. Using cross-entropy loss as the loss function, a parameter model for land cover classification is trained, and then the test set is used to verify the effect of the model. OA, Kappa coefficient, and mIoU are used as evaluation indicators, and the higher the value of each of them, the higher the land cover classification accuracy.
[0023] In this embodiment, the improved ResNet network includes two branches, which respectively process the optical images and SAR images; the processing of the input images by the improved ResNet network includes: the optical images and SAR images extract low-level features through ResNet0 to obtain OPT0 and SAR0, and then gradually extract high-level features through four residual encoding blocks. Before the input of the L-th layer, the optical feature OPT L-1 and the SAR feature SAR L-1 generate complementary features through the adaptive fusion module and are superimposed on the original branches:
[0024]
[0025] The output results OPT4 and SAR4 after passing through the ResNet4 layer are concatenated in channels and input into the multi-scale semantic feature module to extract multi-scale semantic features, and then the spatial information is extracted through the spatial information module. Finally, the resolution is restored through upsampling and the semantic segmentation result is output.
[0026] 5-layer Resnet structure: The feature extraction methods for optical images and SAR images are the same, both consisting of five residual convolutional encoding blocks, and each encoding block contains a convolutional layer, batch normalization, and ReLU activation function. Through these hierarchical feature extractions, features are gradually extracted from optical images and SAR images.
[0027] For the input of optical images and SAR images, first, low-level features are extracted through the preprocessing layer ResNet_pre, passing through a convolutional layer, batch normalization, and ReLU activation, and then through a max pooling layer. The mathematical representation of the preprocessing layer ResNet_pre:
[0028] X pre = MaxPool(ReLU(BN(Conv(X in ))))
[0029] Then, high-level semantic features are gradually extracted through ResNet_1 to ResNet_4. The basic form of each residual block is:
[0030] X out = Relu(X in + F(X in ,{W i}))
[0031] Among them, X in is the input of each residual block, F(X in ,{W i}) is the residual function, including convolution, batch normalization, and ReLU activation function, and {W i} is the set of weights of all convolutional layers within the residual block.
[0032] In this embodiment, as Figure 2 shown, the adaptive fusion module's fusion processing of the input features includes: analyzing the relationship between each element in the optical image features and SAR image features through multi-item bar pooling to obtain the adaptive fusion weights; fusing the optical image features and SAR image features according to the adaptive fusion weights to obtain the fused features.
[0033] By analyzing the relationship between the high-level features of optical images and the high-level features of SAR images among the elements in these two vectors, adaptive fusion weights are learned. First, through a concatenation operation, the number of channels is halved by a 1×1 convolutional layer to generate a four-dimensional tensor feature map, reducing the number of concatenated channels and the computational complexity while retaining key information:
[0034] F B = Conv(concat(F opt ,F sar ))
[0035] In the pooling stage, four strip pooling branches are designed, including two average pooling and two max pooling operations. Through these pooling operations, the original feature map is compressed into a single row or column, and more long-range dependence information is obtained through these long receptive fields. Then, the output of each pooling is subjected to a strip convolution operation with a specific kernel to further strengthen the feature transformation. After that, after an expansion process, the single-row or single-column feature information is restored to the size of H×W, F s1 ,F s2 ,F s3 and F s4 as the output of strip pooling:
[0036] F s = exp(Conv(Pool(F B )))
[0037] After fusing F s1 and F s2 , the feature map after average strip pooling is obtained. Similarly, by fusing F s3 and F s4 , the feature map after max strip pooling is obtained. Then, these two groups of feature maps are concatenated, and after convolution operations and non-linear calculations, the feature weight map F W is obtained:
[0038] F w = Sigmoid(Conv(ReLU(F s1 +F s2 ),ReLU(F s3 +F s4 )))
[0039] Finally, an element-wise multiplication operation is performed on F B and F W to obtain the fused feature fusion F fuse :
[0040] F fuse = F B *F W
[0041] In this embodiment, the multi-scale semantic feature module processes the fused features as follows: input the fused features into the ASPP module, perform dilated convolution operations on the fused features with multiple dilation rates to obtain multi-scale context features; perform dimensionality reduction on the multi-scale context features to obtain the dimensionality-reduced multi-scale context features; input the fused features into the MSCK module, extract multi-scale features from the fused features using convolutional kernels of different sizes, and perform dimensionality reduction on the extracted multi-scale features to obtain the dimensionality-reduced features; perform skip connection on the dimensionality-reduced multi-scale context features and the dimensionality-reduced features to obtain the fused multi-scale features.
[0042] As Figure 3 shown, the high-level fused feature F H obtained by fusion is first input into the ASPP module. Through dilated convolution operations with multiple dilation rates (6, 12, 18), multi-scale context features are obtained. These features include F p1 generated by ordinary convolution, F p2 obtained by convolution with different dilation rates, F p3 , F p4 , and the global feature F p5 obtained by pooling. After these features are dimensionally reduced through 1x1 convolution, they are finally fused into F ASPP .
[0043] Through the multi-scale convolution kernel module, four different sizes of convolution kernels (1x1, 3x3, 5x5, 7x7) are used to extract multi-scale features to handle target feature expressions of different scales. Multiple convolution kernels can capture more detailed information and global features, and the output features are represented as F MSCK , further enriching the diversity of features.
[0044] Finally, the high-level features F ASPP and F MSCK from the ASPP module and the MSCK module are fused with the low-level features and obtained by skip connection. In this way, the final multi-scale semantic feature representation F f is obtained, that is:
[0045]
[0046] This feature fusion strategy combines diverse features of different scales and different modalities, thus effectively improving the ability to distinguish complex land cover types in land cover classification.
[0047] In this embodiment, the spatial information extraction module processes the input data as follows: perform row division, column division, and convolution processing on the multi-scale features respectively to obtain row features, column features, and deep features; input the row features into the BRNN model to obtain horizontal spatial relationships; input the column features into the BRNN model to obtain vertical spatial relationships; input the deep features into the normalization layer for processing; fuse the horizontal spatial relationships, vertical spatial relationships, and the normalized features to obtain spatial information.
[0048] After multi-scale feature fusion, the obtained multi-scale features are input into the spatial information extraction module to model local spatial relationships and long-range spatial relationships, and the structure is as Figure 4 shown.
[0049] Local spatial relationships refer to the spatial relationships between ground objects within the same type of area. For example, in the classified area of a city, there is a certain spatial relationship between the residential area and the industrial area. To capture such local spatial relationships, for the constructed multi-scale semantic feature representation graph F f , apply a 3x3 convolution and obtain a new feature fusion graph F s of the same size through padding. The formula is as follows:
[0050]
[0051] where represents the convolution operation, W s represents the 3x3 filter, and b s is the bias of the same dimension as F f .
[0052] Long-range spatial relationships consider the spatial relationships between different areas, such as the spatial relationship between a village and a city. The present invention uses two Bidirectional RNNs (BRNNs) that can extract multi-directional information to model the long-range spatial relationships between different areas. The BRNN model consists of two independent processing streams that scan data in two directions. transmits information from left to right and from right to left, and BRNN ↑↓ transmits information from top to bottom and from bottom to top. By stacking and BRNN ↑↓ , BRNNnet can obtain the context information for transmitting the feature map in four directions.
[0053] Given the constructed feature map F f , using the information transmission along rows and columns in F f , through two BRNNs that sweep horizontally and vertically across F f , model the long-range spatial relationships. Taking <http: / / www.example.com>[http: / / www.example.com] as an example, asFigure 5 As shown, split F f into rows to obtain Fr. Each BRNN takes the first row features of the feature map Fr ([Fr 11 , Fr 12 ,..., Fr 1n ) as the data sequence, aggregates information through multiple gated recurrent units (GRUs), and converts it into two long-range dependencies to obtain S → and S ← . The BRNN model consists of two independent processing streams for scanning sequence data in two directions. In , the two directions are from left to right and from right to left, while in BRNN ↑↓ , the two directions are from top to bottom and from bottom to top.
[0054] By stacking and BRNN ↑↓ , the context of the feature map Ff can be transmitted in four directions, realizing two-dimensional long-range spatial information transmission in the feature map, which helps to identify large-sized UFZs (such as residential areas) or linear layouts (such as roads). The long-range relationship of the first row of F f can be described as:
[0055]
[0056] Finally, and S ↑↓ as well as the local spatial relationship F s are cascaded to obtain the enhanced feature F e . Input F e into the decoder for feature decoding and class prediction. The decoder consists of a convolutional layer that adjusts the number of channels of the feature map to the number of predicted classes. The decoded feature map contains the class prediction information for each pixel point. To restore to the original resolution of the input image, bilinear interpolation is used to upsample the decoded feature map, and finally the upsampled feature map of F out is output, which has the same spatial resolution as the input image and contains the class prediction results for each pixel point.
[0057] In this embodiment, the loss function of the model uses the cross-entropy loss function and the Dice loss function, and the weights of the cross-entropy loss function and the Dice loss function are both 0.5.
[0058] The comprehensive evaluation method is used to evaluate the classification results, including using the overall accuracy (OA), Kappa coefficient (Kappa), and mIoU as the main evaluation indicators:
[0059] Accuracy, OA), Kappa coefficient (Kappa), and mIoU as the main evaluation indicators:
[0060]
[0061] Among them, K is the total number of categories, which is equal to the width of the confusion matrix. p ij represents the number of pixels in which the i-th category is predicted as the j-th category. The higher the values of these three metrics, the better the result.
[0062] In this embodiment, the experiments in this paper were all carried out on a server equipped with NVIDIA A100 with a video memory of 40GB. The network architecture was built using PyTorch (version 1.12), and the operating system was Ubuntu 18.04. The Adam optimizer was used for parameter updates, and the weight decay rate was 0.0001. The step size of the StepLR scheduler was 20, and the decay rate was 0.1. The basic learning rate was set to 0.001. The training image size was 256x256 pixels. The batch size during training was set to 32.
[0063] The datasets used were the WHU-OPT-SAR and DFC2020 datasets, which contained 29,400 and 6,114 pairs of optical and SAR images respectively. The training set and test set were both allocated in a ratio of 4:1, and the number of channels for the multi-modal input of the two datasets was kept the same. A single-channel grayscale SAR image and a 4-channel (RGB and NIR) optical image were used. Among them, the multi-spectral image of DFC2020 had 13 bands, and B4, B3, B2, and B8 were selected to combine the corresponding RGBN as the input of the optical image. The evaluation metrics were OA, Kappa, and mIoU.
[0064] The model performing one pass of the gradient descent algorithm on all training data is called one epoch. In each epoch, the parameters of the model are updated, and the maximum number of epochs is set to 100. The learning rate is updated every 20 epochs. During the 100 epochs of training the model, the model and its parameters that achieved the best results on the test dataset are saved.
[0065] The above-mentioned embodiments further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image land cover classification method based on spatial information enhancement, characterized in that, Including: Obtain optical and shadow images and SAR images; perform enhancement processing on the optical and shadow images and SAR images; Input the enhanced optical and shadow images and SAR images into a dual-branch convolutional neural network respectively to obtain the land cover classification results; Use a comprehensive evaluation method to evaluate the classification results; Training the dual-branch convolutional neural network includes: obtaining an original dataset, where the original dataset contains optical and shadow images and SAR images; performing enhancement processing on the images in the original image dataset; inputting the enhanced optical and shadow images and SAR images into an improved ResNet network respectively for feature extraction to obtain fused features with optical and shadow image information and SAR image information; inputting the fused features into a multi-scale semantic feature module to obtain multi-scale features; inputting the multi-scale features into a spatial information extraction module to obtain spatial information; inputting the spatial information into a classifier to obtain the classified land cover classification results; calculating the loss function of the model according to the land cover classification results, adjusting the parameters of the model, and completing the training of the model when the loss function converges.
2. The method for classifying image land cover based on spatial information enhancement according to claim 1, wherein, Performing enhancement processing on the optical and shadow images and SAR images includes: segmenting the images to obtain optical and shadow images and SAR images with a size of 256×256; performing translation, horizontal flipping, and vertical flipping on the segmented images to obtain the enhanced optical and shadow images and SAR images.
3. A method for classifying image land cover based on spatial information enhancement according to claim 1, characterized in that, The improved ResNet network includes two branches, and the two branches process optical images and SAR images respectively; The improved ResNet network processes the input images as follows: the optical image and the SAR image extract low-level features through ResNet0 to obtain OPT0 and SAR0; the low-level features are gradually used to extract high-level features through four residual encoding blocks. Before the input of the L-th layer, the optical feature OPT L-1 and the SAR feature SAR L-1 generate complementary features through the adaptive fusion module, and the complementary features are correspondingly superimposed on the original branches; the output results OPT4 and SAR4 after the ResNet4 level are concatenated in channels to obtain the fusion features.
4. A method for classifying image land cover based on spatial information enhancement according to claim 3, characterized in that, The ResNet hierarchical structure includes a convolutional layer, batch normalization, and a ReLU activation function; its processed data includes using the convolutional layer to extract the features of the input image; performing batch normalization, ReLU activation, and max pooling layer processing on the extracted features to obtain low-level features.
5. A method for classifying image land cover based on spatial information enhancement according to claim 1, characterized in that, The adaptive fusion module performs fusion processing on the input features, including: concatenating the optical and shadow image features and the SAR image features, reducing the number of channels of the concatenated features by half through a 1×1 convolutional layer to generate a four-dimensional tensor feature map; inputting the four-dimensional tensor feature map into 4 bar pooling branches respectively for pooling operations, where the 4 bar poolings include two average pooling operations and two maximum pooling operations; performing convolution expansion operations on the features output by each bar pooling with the corresponding kernels to obtain four feature maps F s1 , F s2 , F s3 and F s4 ; fusing F s1 and F s2 to obtain the feature map after average bar pooling; fusing F s3 and F s4 to obtain the feature map after maximum bar pooling; concatenating the feature map after average bar pooling and the feature map after maximum bar pooling, performing convolution and non-linear calculations on the concatenated feature map to obtain a feature weight map; performing element-wise multiplication operations on the weight feature map and the four-dimensional tensor feature map to obtain the fused features.
6. A method for classifying image land cover based on spatial information enhancement according to claim 1, characterized in that, The multi-scale semantic feature module processes the fused features including: inputting the fused features into an ASPP module, performing dilated convolution operations on the fused features through multiple dilation rates to obtain multi-scale context features; performing dimensionality reduction processing on the multi-scale context features to obtain the dimensionality-reduced multi-scale context features; inputting the fused features into an MSCK module, performing multi-scale feature extraction on the fused features using convolutional kernels of different sizes, and performing dimensionality reduction on the extracted multi-scale features to obtain the dimensionality-reduced features; performing skip connection on the dimensionality-reduced multi-scale context features and the dimensionality-reduced features to obtain the fused multi-scale features.
7. A method for classifying image land cover based on spatial information enhancement according to claim 1, characterized in that, The spatial information extraction module processes the input data including: performing row division, column division, and convolution processing on the multi-scale features respectively to obtain row features, column features, and deep features; inputting the row features into a BRNN model to obtain the horizontal spatial relationship; inputting the column features into a BRNN model to obtain the vertical spatial relationship; inputting the deep features into a normalization layer for processing; fusing the horizontal spatial relationship, the vertical spatial relationship, and the normalized features to obtain spatial information.
8. A method for classifying image land cover based on spatial information enhancement according to claim 1, characterized in that The loss function of the model uses a cross-entropy loss function and a Dice loss function, and the weights of the cross-entropy loss function and the Dice loss function are both 0.
5.
9. A method for classifying image land cover based on spatial information enhancement according to claim 1, characterized in that, The evaluation of the classification results using a comprehensive evaluation method includes: using the overall accuracy OA, Kappa coefficient, and mIoU as evaluation indicators.