A lightweight land cover classification network model based on multi-scale and edge perception
By adopting the combination of ResNet-18 backbone network and multi-scale edge-aware loss function in remote sensing image geometry classification, the segmentation accuracy and speed of the lightweight CNN model are improved, and the problems of low segmentation accuracy and slow inference speed in remote sensing image geometry classification are solved, and are suitable for mobile devices and real-time response scenarios.
Patent Information
- Application Number
- CN202310588686.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-05-23
AI Technical Summary
The existing lightweight CNN models have low segmentation accuracy, poor generalization ability, and slow inference speed in remote sensing image geometry classification, which is difficult to meet the needs of real-time and storage resource-constrained scenarios.
ResNet-18 is used as the backbone network for feature extraction, and multi-scale information and edge-aware loss function are used to assist supervision in the decoder stage. Feature fusion is carried out through semantic information enrichment module, spatial feature refinement module and edge-aware module to improve the segmentation accuracy and speed of the model.
It realizes higher segmentation accuracy and faster inference speed in remote sensing image geometry classification, and is suitable for resource-constrained mobile devices and real-time response scenarios.
Smart Images

Figure CN116935094B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer image processing and realizes land cover classification based on a lightweight network model with multi-scale and edge awareness. Background Art
[0002] Remote sensing, as a comprehensive modern surveying and mapping technology, plays an important role in earth observation. Image semantic segmentation refers to assigning different labels to different pixels in an image to divide it into different semantic categories, achieving pixel-level classification and annotation. Using remote sensing images for efficient and accurate land cover classification is of great significance in practical applications such as environmental protection, urban planning, land use, and natural disaster prevention. In recent years, thanks to the powerful feature induction and learning ability, deep learning methods represented by Convolutional Neural Networks (CNNs) have greatly surpassed traditional methods in terms of accuracy, but they have also introduced too many parameters, hindering the deployment and application of the models.
[0003] While the land cover classification algorithm for remote sensing images based on deep learning achieves high accuracy, there are some real-time problems. As the complexity of the model increases, the computational overhead required for training and inference also increases, making it very difficult to deploy in resource-constrained scenarios. Especially in scenarios that require real-time response, the processing speed of the model must be fast enough, but currently, the models often have the problem of being too slow in processing speed. In addition, the volume of the model will also become large with the increase in complexity, which is not conducive to deployment in scenarios with limited storage capacity such as mobile devices. Therefore, on the premise of ensuring high accuracy, how to improve the real-time performance of the model, reduce the volume and computational overhead of the model has become a problem that needs to be solved by the current land cover classification algorithm for remote sensing images. It is necessary to further study the real-time problems of applying deep learning technology to land cover classification of remote sensing images to achieve the application and implementation of the algorithm.
[0004] Nowadays, the development of edge computing and the increase in the number of mobile terminal devices have made the lightweight of models become increasingly important. Therefore, lightweight neural network models designed specifically for mobile application terminals have received more and more attention. Compared with traditional neural networks, the network structure of lightweight models is simpler; it can not only train network parameters faster, but also occupies less computing resources of computer devices, and the trained neural network model requires less storage space. With the development of drones, the demand for using drones to process large samples of remote sensing images in real time will be greater and greater. With the development of edge computing and the improvement of image resolution, a large amount of computing and parameter limitations restrict real-time semantic segmentation. The emergence of lightweight CNN has accelerated the development of real-time semantic segmentation of ground object classification. However, nowadays, the spatio-temporal span of aerial images is getting larger and larger, resulting in the loss of details and blurred edges of lightweight CNN. Therefore, the existing lightweight CNN models have low segmentation accuracy and poor generalization ability in real-time ground object segmentation tasks. Therefore, the present invention improves the network to address these problems Summary of the Invention
[0005] In order to overcome the deficiencies of the above-mentioned prior art, the present invention discloses a lightweight ground object classification model based on multi-scale and edge perception to solve the problem that the feature recognition ability between categories near the fuzzy edges in the multi-classification of ground objects in remote sensing images is weak, resulting in incorrect segmentation in the fuzzy edge area and to improve the inference speed of the model. Taking the Feature Pyramid Network as the benchmark network, in the encoder stage, ResNet18, which is compatible with the training speed and network depth, is used as the backbone network to extract image features. In the decoder stage, the semantic information of the low-resolution feature map and the spatial information of the high-resolution feature map are fully utilized, and multi-task learning is used to combine the edge perception loss function to assist in supervising the learning of semantic segmentation. The proposed method can not only make full use of the semantic, spatial and edge information of the image, but also avoid the disadvantages of slow speed and high memory occupation during inference
[0006] The technical steps adopted by the present invention are as follows:
[0007] Step 1: The input image is subjected to feature extraction through the lightweight network ResNet-18 network to obtain feature maps X1, X2, X3, X4 and the feature map X5 after passing through the global average pooling layer;
[0008] Step 2: The feature maps X3, X4 and X5 form a feature map Y3 with rich semantic information through the semantic information enhancement module;
[0009] Step 3: The feature map X2 and Y3 obtain the refined feature map Y2 through the spatial feature refinement module;
[0010] Step 4: Feature map X1 and Y2 pass through the edge-aware attention module, and the loss function of edge detection is used to assist the loss function of semantic segmentation for training, and finally the semantic segmentation image of the image is extracted.
[0011] Further, the specific method of step 2 is as follows:
[0012] Step 2.1: Feature maps X4 and X5 pass through the feature fusion module. Specifically, the resolution of feature map X5 is upsampled to the same size as feature map X4, and then passed through two 3×3 convolutional layers with batch normalization (BN) and ReLU functions to obtain a new feature map Y4.
[0013] Step 2.2: Feature maps X3 and Y4 pass through the feature fusion module to obtain a feature map Y3 with rich semantic information.
[0014] Further, the specific method of step 3 is as follows:
[0015] Step 3.1: Feature map X2 is refined through the spatial attention mechanism module to obtain X2 ’ 。
[0016] Step 3.2: The feature map Y3 with rich semantic information is upsampled to the same dimension as the feature-refined feature map X2 ’ and multiplied pointwise with X2 ’ and then passed through a 3×3 convolutional layer with BN and ReLU functions to obtain a new feature map F1;
[0017] Step 3.3: Meanwhile, feature map X2 ’ is multiplied pointwise with feature map Y3 through a 3×3 convolutional layer with BN and ReLU functions and then passed through a 3×3 convolutional layer with BN and ReLU functions to obtain a new feature map F2;
[0018] Step 3.4: F2 and F1 pass through the feature fusion module to obtain the final fused feature map Y2.
[0019] Further, the specific method of step 4 is as follows:
[0020] Step 4.1: Feature map X1 and the feature map Y4 containing semantic information are fused to form an edge flow E t ;
[0021] Step 4.2: Set the feature map Y2 output in step 2 as the semantic flow S t ;
[0022] Step 4.3: Semantic flow S tFirst, it is upsampled to the same dimension as the edge flow through bilinear interpolation, and then the corresponding elements of the tensors are added to obtain the combined flow F t ;
[0023] Step 4.4: Semantic flow S t First, it is upsampled to the same dimension as the edge flow through bilinear interpolation, and then the corresponding elements of the tensors are added to obtain the combined flow F t ;
[0024] Step 4.5: Subsequently, the combined flow F t Passes through a gated attention mechanism. Specifically, first, it passes through a global average pooling layer and a 1×1 convolutional layer. Then, the module uses the Sigmoid function to limit the output within the range of [0,1];
[0025] Step 4.6: Finally, after passing through the channel attention mechanism, two 3×3 convolutional layers with BN and ReLU are used to refine the features. An edge-aware loss function is used to assist the main task of semantic segmentation, and finally, the ground object classification result is obtained.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0027] (1) The present invention designs a semantic information enrichment module to provide richer semantic information for the model to better understand the image.
[0028] (2) The present invention designs a spatial information enrichment module to achieve refined ground object classification of remote sensing images. While maintaining sufficient spatial information, the receptive field should be increased as much as possible to achieve high-precision real-time ground object classification of remote sensing images.
[0029] (3) The present invention designs an edge-aware attention module, which uses multi-task learning and combines an edge-aware loss function to assist in the learning of supervised semantic segmentation, improving the boundary effect of the network for ground object classification. Description of the Drawings
[0030] Figure 1 is the network structure diagram of the present invention
[0031] Figure 2 is the structure diagram of the feature fusion module of the present invention
[0032] Figure 3 is the structure diagram of the cross-fusion module of the present invention
[0033] Figure 4 is the structure diagram of the edge-aware attention module of the present invention
[0034] Specific implementation steps
[0035] The present invention will be further described below in conjunction with the accompanying drawings.
[0036] The present invention designs a lightweight ground object classification model based on multi-scale and edge perception to improve the segmentation speed of the current deep learning-based high-resolution remote sensing image ground object classification method and solve the weak feature recognition ability between categories near the fuzzy edges in the multi-classification of remote sensing images, resulting in incorrect segmentation in the fuzzy edge area. As Figure 1 shown, it is the network structure diagram of the present invention.
[0037] In the encoder stage, first, the lightweight network ResNet18 is used as the backbone feature extraction network. Feature maps X1, X2, X3, and X4 are obtained in Stage1, Stage2, Stage3, and Stage4 respectively, and the feature map X5 is obtained using the global average pooling layer. In the decoder stage, it is divided into a semantic information enrichment module, a spatial feature refinement module, and an edge perception module. In the semantic information enrichment module, the output feature map X3 of Stage3, the output feature map X4 of Stage4, and the global feature map X5 are used by the feature fusion module to form a feature map Y3 containing rich semantic information, as Figure 2 shown. In the spatial feature refinement module, as Figure 3 shown, the output feature map X2 of Stage2 is enhanced in spatial information through the spatial attention mechanism module and then cross-fused with the output feature map Y3 of the semantic information enrichment module to refine the spatial features of the feature map. As Figure 4 shown, in the edge perception module, the output feature map X1 of Stage1 and the feature map Y4 are feature-fused to extract the edge detection image of the image, and the semantic segmentation image of the image is extracted from the feature map Y2 after spatial feature refinement through the edge perception attention module. The loss function of edge detection is used to assist the loss function of semantic segmentation for training, and finally a network model with compatible accuracy and speed is obtained.
Claims
1. A lightweight feature classification network model based on multi-scale and edge perception, characterized in that It includes the following steps: Step 1: The input image is subjected to feature extraction through the lightweight network ResNet-18 network to obtain feature maps X1, X2, X3, X4, and the feature map X5 after passing through the global average pooling layer; Step 2: The feature maps X3, X4, and X5 form the feature map Y3 with rich semantic information through the semantic information enhancement module; Step 2.1: The feature maps X4 and X5 pass through the feature fusion module. Specifically, the resolution of the feature map X5 is upsampled to the same size as the feature map X4 and then passed through two 3×3 convolutional layers with batch normalization layer (Batch Normalization, BN) and ReLU function to obtain the new feature map Y4; Step 2.2: The feature maps X3 and Y4 pass through the feature fusion module to obtain the feature map Y3 with rich semantic information; Step 3: The feature map X2 and Y3 obtain the refined feature map Y2 through the spatial feature refinement module; Step 3.1: Refine the feature map X2 through the spatial attention mechanism module to obtain X2 ’ ; Step 3.2: Upsample the feature map Y3 with rich semantic information to the same dimension as the feature map X2 after feature refinement, and perform dot multiplication with X2, and then pass through a 3×3 convolutional layer with BN and ReLU functions to obtain a new feature map F1; ’ at the same dimension, and with X2 ’ perform dot multiplication, and then pass through a 3×3 convolutional layer with BN and ReLU functions to obtain a new feature map F1; Step 3.3: Meanwhile, the feature map X2 ’ is multiplied element-wise with the feature map Y3 through a 3×3 convolutional layer with BN and ReLU functions and then passed through another 3×3 convolutional layer with BN and ReLU functions to obtain a new feature map F2; Step 3.4: F2 and F1 pass through the feature fusion module to obtain the finally fused feature map Y2; Step 4: The feature map X1 and Y2 pass through the edge-aware attention module, and the loss function of edge detection is used to assist the loss function of semantic segmentation for training, and finally the semantic segmentation image of the image is extracted; Step 4.1: Feature fusion is performed on the feature map X1 and the feature map Y4 containing semantic information to form an edge flow E t ; Step 4.2: Set the feature map Y2 output in Step 2 as the semantic flow S t ; Step 4.3: Semantic flow S t First, it is upsampled to the same dimension as the edge flow through bilinear interpolation, and then the corresponding elements of the tensors are added together to obtain the combined flow F t ; Step 4.4: Semantic flow S t First, it is upsampled to the same dimension as the edge flow through bilinear interpolation, and then the corresponding elements of the tensors are added to obtain the combined flow F t ; Step 4.5: Subsequently, combine the flow F t through a gated attention mechanism, where first through a global average pooling layer and a 1×1 convolutional layer, and then, the module uses the Sigmoid function to limit the output within the range of [0, 1]; Step 4.6: Finally, after passing through the channel attention mechanism, two 3×3 convolutional layers with BN and ReLU are used to refine the features, and the edge-aware loss function is used to assist the main task of semantic segmentation, and finally the ground object classification result is obtained.
2. The method according to claim 1, characterized in that Step 2 designs the semantic information enrichment module, and uses feature fusion to obtain rich semantic information of the image.
3. The method according to claim 1, wherein Step 3 designs the spatial feature refinement module, and uses the cross-fusion module and spatial attention mechanism to obtain refined spatial feature information.
4. The method according to claim 1, wherein Step 4 introduces semantic features and edge features into the channel attention mechanism and gated attention mechanism, and uses multi-task learning to use the edge-aware loss function to assist in supervising the learning of the main task of semantic segmentation.
Citation Information
Patent Citations
High-resolution remote sensing image-oriented boundary enhanced semantic segmentation method
CN115049936A
Fully-supervised farmland plot extraction method under spatial constraint
CN115311575A