Remote sensing land segmentation method based on multi-scale attention and lightweight feature fusion
Through the remote sensing land segmentation method that integrates multi-scale attention and lightweight features, the problems of complex data features and large computing resource occupancy in remote sensing images are solved, and more accurate and efficient land recognition is achieved.
Patent Information
- Application Number
- CN202510545045.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-11
AI Technical Summary
In the land segmentation of remote sensing images, there are problems such as complex data characteristics, unbalanced samples, insufficient model receptive fields, blurred boundaries and large computing resources, resulting in low segmentation accuracy and high computing cost.
The remote sensing land segmentation method based on the fusion of multi-scale attention and lightweight features is adopted. The backbone feature extraction network, space pooling module and composite pooling module are combined with the multi-scale fusion upsampling structure to enhance feature extraction and context information capture, and the model training is optimized using the cross entropy loss function with weights.
It realizes more accurate and lighter land recognition for remote sensing images, improves segmentation accuracy and recognition efficiency, and reduces computing resource occupation.
Smart Images

Figure CN120298702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and remote sensing image processing, and particularly relates to a remote sensing land segmentation method based on multi-scale attention and lightweight feature fusion. Background Art
[0002] Remote sensing refers to a technology that uses sensors to obtain surface or atmospheric information from a distance. Remote sensing image processing tasks involve analyzing and processing the images obtained from remote sensors to extract useful information. Among them, remote sensing land segmentation is a key task in remote sensing image processing, which is a process of identifying and locating the types of ground objects in the scene using remote sensing images, classifying the ground objects in the original remote sensing data images to divide the land types. Good remote sensing land segmentation results can provide guiding opinions for environmental protection, agricultural production, and urban construction.
[0003] In recent years, in the field of remote sensing image land segmentation, deep learning has become increasingly mature and the detection accuracy has become higher and higher. Remote sensing images face many challenges in semantic segmentation. First, remote sensing images are different from ordinary images. In remote sensing image land segmentation, the data features are more complex and the sample distributions of each category are very uneven, which makes the deep learning method have a low generalization ability. Second, in the remote sensing image land segmentation task, the pixel range actually involved in the calculation when the model makes predictions. A smaller receptive field may not be able to cover the entire picture of the target object. Especially for complex and diverse-scale ground objects, the insufficient effective receptive field of the model may cause the model to be difficult to capture the large-scale context information, thereby affecting the accuracy of the segmentation result. Moreover, due to the large intra-class difference and small inter-class difference between samples, it is difficult for the model to accurately capture and distinguish the boundaries of ground objects, and the boundaries in the segmentation result are blurred, misclassified, and misaligned. With the development of software technology and the emergence of high-performance computer hardware, it provides strong support for the application of deep learning in remote sensing land segmentation. Using deep learning for remote sensing land segmentation can improve the classification accuracy and recognition efficiency. Therefore, it is of great significance to combine remote sensing image land segmentation and deep learning. Summary of the Invention
[0004] The purpose of the present invention is to propose a method that can achieve more accurate and lightweight recognition results for various types of land in remote sensing images in view of the technical problems in the research on land segmentation based on remote sensing images, such as many land categories, complex features, low accuracy of small targets, poor edge segmentation effect, and large consumption of computing resources.
[0005] A land segmentation method based on remote sensing images includes the following steps:
[0006] Step 1: Obtain remote sensing dataset images, screen the pictures and then label them, which altogether include six categories: water body, building, low vegetation, tree, car, background;
[0007] Step 2: Crop, scale, and perform data augmentation on the original image, and divide it into a training set, a validation set, and a test set according to 80%, 10%, and 10%.
[0008] Step 3: Put the training set in the prepared dataset into the feature extraction network and the multi-scale fusion upsampling structure for training to obtain a segmentation model.
[0009] Step 4: Set different training parameters to train the model to obtain an optimal model.
[0010] Step 5: Put the remote sensing image to be classified into the trained segmentation model to identify and obtain the best segmentation result.
[0011] In Step 1, obtain the original remote sensing image of urban and rural land, screen the pictures and label them. Specifically, it also includes the following sub-steps:
[0012] 1-1: Use a drone to take pictures of urban and rural areas to obtain remote sensing images.
[0013] 1-2: Perform image stitching, shadow removal, noise removal, etc.
[0014] 1-3: Crop the pictures, remove the pictures with a single category and those without the required categories in the images, and label the pictures containing water bodies, buildings, low vegetation, trees, cars, and backgrounds.
[0015] In Step 3, it specifically includes the following steps:
[0016] The first layer F1 of the backbone feature extraction network → the second layer F2 of the backbone feature extraction network → the third layer F3 of the backbone feature extraction network → the fourth layer F4 of the backbone feature extraction network → the fifth layer F5 of the backbone feature extraction network → the spatial pooling module;
[0017] The second layer of the backbone feature extraction network → the first composite pooling module, the third layer of the backbone feature extraction network → the second composite pooling module, the fourth layer of the backbone feature extraction network → the third composite pooling module;
[0018] The spatial pooling module and the composite pooling module MPPM are enhanced feature extraction modules.
[0019] The spatial pooling module is as follows: The fifth layer of the backbone is input into the spatial pooling module for multi-scale average pooling, and the bins are 1*1, 2*2, 3*3, 6*6, 12*12, and 24*24 respectively. Perform 1*1 convolution to adjust the channels, upsample to restore the size, and perform concat splicing to output the feature map F6.
[0020] The composite pooling module is as follows: The fourth, third, and second layers of the backbone network are input into the composite pooling module. The fourth layer performs strip pooling of H*1, 1*W, and H*W, followed by convolutional expansion and summation of the feature maps to output the feature map F5. After downsampling, the third layer is concatenated with the second layer and subjected to MPPM processing to output intermediate features. After downsampling, the second layer is concatenated with the third layer to output intermediate features.
[0021] In the decoder part, multi-scale feature fusion and upsampling are performed, which are divided into a pooling module fusion layer, a multi-scale fusion upsampling module, and final fusion and output.
[0022] The spatial pooling module F6 and the F5 output by the composite pooling module are fused through ADD to output the fused features.
[0023] The fourth layer F4 of the backbone → spatial attention SA → Conv+BN+ReLU → Add fusion with the output of the pooling module fusion layer 10 → output the feature map F10.
[0024] The third layer (F3) of the backbone → spatial attention SA → Conv+BN+ReLU → Add fusion with the output (F10) of the first module 11 → upsampling → output the feature map F11.
[0025] The second layer (F2) of the backbone → spatial attention SA → Conv+BN+ReLU → Add fusion with the output (F11) of the second module 12 → upsampling → output the feature map F12.
[0026] F12 (from the third module 13) + the first layer (F1) of the backbone are concatenated through Concat → 3×3 convolution to adjust the channels → upsampled to the original image size → output the segmentation result.
[0027] In step 3, the loss function for network training is as follows:
[0028] The loss function consists of the weighted cross-entropy loss Cross Entropy Loss, and its calculation formula is as follows:
[0029]
[0030] Among them, M represents the number of categories. In the present invention, M takes the value of 5, namely five categories of cultivated land, orchard, forest, building, and background. c is a variable value representing a certain category. yc is a one-hot vector, and its elements have only two values, 0 and 1. If the category is the same as the category of the sample, it takes 1, otherwise it takes 0. pc represents the probability that the predicted sample belongs to c, and wc represents the weight coefficient. Its calculation formula is:
[0031]
[0032] Where N represents the total number of pixels, and Nc represents the number of pixels with Ground truth category c.
[0033] In step 4, the remote sensing image of the area to be identified for land classification is cut into pictures of the input size required by the model and input into the trained segmentation model to obtain the segmentation result.
[0034] The present invention provides a method for identifying land category areas through remote sensing images; compared with general manual statistics, it is more cost-saving and has higher efficiency. Compared with traditional methods for identifying land classification areas on remote sensing images, the identification accuracy is higher, and it can identify and segment land on a larger scale;
[0035] In the construction of the network model of the present invention, an attention module is used to improve the segmentation accuracy of the model. At the same time, the enhanced feature extraction module is improved, the size of the square pooling is enriched, the segmentation effect of small targets is enhanced, and a composite pooling structure is newly added. Using feature maps at different levels, the model can better capture rich context information. At the same time, finally, a multi-layer upsampling and deep and shallow feature fusion structure is adopted to make up for the loss of edge information caused by the deepening of the network layers, optimize the segmentation effect, and obtain better results. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flow chart of the present invention;
[0037] Figure 2 is the overall framework diagram of the remote sensing image segmentation network model in the embodiment of the present invention;
[0038] Figure 3 is the structure diagram in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0039] The purpose of the present invention is to propose a method that can achieve more accurate and lightweight recognition results for various types of land in remote sensing images in view of the technical problems in the research on land segmentation based on remote sensing images, such as many land categories, complex features, low accuracy of small targets, poor edge segmentation effect, and large consumption of computing resources.
[0040] As Figure 1 shown, a land segmentation method based on remote sensing images includes the following steps:
[0041] Step 1: Obtain remote sensing dataset images, screen the pictures and mark them, which altogether include six categories: water body, building, low vegetation, trees, cars, and background;
[0042] Step 2: Crop, scale, and perform data augmentation on the original images, and divide them into a training set, a validation set, and a test set according to 80%, 10%, and 10%;
[0043] Step 3: Put the training set in the prepared dataset into the feature extraction network and the multi-scale fusion upsampling structure for training to obtain a segmentation model;
[0044] Step 4: Set different training parameters to train the model to obtain an optimal model.
[0045] Step 5: Put the remote sensing image to be classified into the trained segmentation model for recognition to obtain the best segmentation result;
[0046] In Step 1, obtain the original remote sensing image of urban and rural land, screen the pictures and mark them. Specifically, it also includes the following sub-steps:
[0047] 1-1: Use a drone to take pictures of urban and rural areas to obtain remote sensing images;
[0048] 1-2: Process the images such as stitching, shadow removal, and noise removal;
[0049] 1-3: Crop the pictures, remove the images with a single category and those without the required categories in the images, and label the pictures containing water bodies, buildings, low vegetation, trees, cars, and backgrounds.
[0050] In Step 3, the feature extraction network and the multi-scale fusion upsampling structure are composed of a backbone feature extraction network, a spatial pooling module, a composite pooling module, a pooling module fusion layer, and a multi-scale fusion upsampling module; specifically, it includes the following steps:
[0051] The first layer F1 of the backbone feature extraction network → the second layer F2 of the backbone feature extraction network → the third layer F3 of the backbone feature extraction network → the fourth layer F4 of the backbone feature extraction network → the fifth layer F5 of the backbone feature extraction network → the spatial pooling module;
[0052] The second layer of the backbone feature extraction network → the first composite pooling module, the third layer of the backbone feature extraction network → the second composite pooling module, the fourth layer of the backbone feature extraction network → the third composite pooling module;
[0053] The spatial pooling module and the composite pooling module MPPM are enhanced feature extraction modules.
[0054] The spatial pooling module is as follows: The fifth layer of the backbone is input into the spatial pooling module for multi-scale average pooling, and the bins are 1*1, 2*2, 3*3, 6*6, 12*12, and 24*24 respectively. Perform 1*1 convolution to adjust the channels, up-sample to restore the size, and perform concat splicing to output the feature map F6.
[0055] The composite pooling module is as follows: The fourth, third, and second layers of the backbone network are input into the composite pooling module. The fourth layer performs strip pooling of H*1, 1*W, and H*W, followed by convolutional expansion, summation of feature maps, and output of the feature map F5; After downsampling the third layer, it is concatenated with the second layer, and MPPM processing is performed to output intermediate features; After downsampling the second layer, it is concatenated with the third layer to output intermediate features.
[0056] The decoder part performs multi-scale feature fusion and upsampling, which is divided into a pooling module fusion layer, a multi-scale fusion upsampling module, and final fusion and output.
[0057] The spatial pooling module F6 and the F5 output by the composite pooling module are fused through ADD to output the fused feature.
[0058] The fourth layer of the backbone F4 → spatial attention SA → Conv+BN+ReLU → Add fusion with the output of the pooling module fusion layer 10 → output the feature map F10.
[0059] The third layer of the backbone F3 → spatial attention SA → Conv+BN+ReLU → Add fusion with the output of the first module 11 (F10) → upsampling → output the feature map F11.
[0060] The second layer of the backbone F2 → spatial attention SA → Conv+BN+ReLU → Add fusion with the output of the second module 12 (F11) → upsampling → output the feature map F12.
[0061] F12 (from the third module 13) + the first layer of the backbone (F1) are concatenated through Concat → 3×3 convolution to adjust the channels → upsampling to the original image size → output the segmentation result.
[0062] In step 3, the loss function for network training is as follows:
[0063] The loss function consists of weighted cross-entropy loss Cross Entropy Loss, and the calculation formula is as follows:
[0064]
[0065] Among them, M represents the number of categories, yc is a one-hot vector, and its elements can only take two values, 0 and 1. If the category is the same as the sample category, it takes 1, otherwise it takes 0. pc represents the probability that the predicted sample belongs to c, and wc represents the weight coefficient. Its calculation formula is:
[0066]
[0067] Among them, N represents the total number of pixels, and Nc represents the number of pixels with the Ground truth category of c.
[0068] In step 4, the remote sensing image of the area that needs to be identified for land classification is cut into pictures of the input size required by the model and input into the trained segmentation model to obtain the segmentation result. In step 3, specifically:
[0069] 1) Construct a feature extraction network to extract target features;
[0070] 2) Construct a multi-scale feature fusion upsampling structure to obtain prediction results using the features;
[0071] In step 1), a feature extraction network is constructed to obtain target features. The feature extraction network consists of a backbone feature extraction network and an enhanced feature extraction module, and specifically includes the following steps:
[0072] 1-1) Divide Resnet-50 into four parts, and then add channel and spatial attention to the first and last layers of the network: In the channel attention block, first perform global max pooling and average pooling on the feature map, then send them into a two-layer neural network, and then pass the obtained features through the Sigmoid activation function to obtain the weight coefficients; finally, multiply the weight coefficients by the input features to obtain the scaled new features. In the spatial attention block, first perform max pooling and average pooling in the channel dimension respectively to obtain two channel descriptions, and then concatenate these two descriptions together by channel. Then, through a 7×7 convolutional layer and the sigmoid activation function to obtain the weight coefficients, and finally multiply the weight coefficients by the input features to obtain the scaled new features.
[0073] 1-2) The feature map obtained by resnet-50 is input into the enhanced feature extraction module. The enhanced feature extraction module consists of spatial pooling and composite pooling. The first part divides the input feature map F4 into six regions of different sizes: 1×1, 2×2, 3×3, 6×6, 12×12, and 24×24, and then performs average pooling on each region to obtain feature layers of different sizes. After upsampling, they are concatenated with the input F4 to obtain the feature map F6. The second part first inputs F1 into the composite pooling module mppm, then downsamples the output result and concatenates it with F2 and then inputs it into the mppm module. The output result is downsampled and concatenated with F3 and then input into the mppm module to obtain the feature map F5. Among them, the composite pooling module mppm divides the feature map into two long-shaped regions of H×1 and 1×W and an original image region of H×W for strip pooling. Subsequently, the two output feature maps are expanded along the left and right and up and down respectively through convolution, and the sum of the corresponding positions of the expanded feature maps is obtained to obtain the feature map of H×W.
[0074] In step 2), it specifically includes the following steps:
[0075] 2-1): Obtain the F1, F2, and F3 layers in the backbone feature extraction network, and input them into a spatial attention respectively. The spatial attention performs average pooling and max pooling operations on the input feature maps F1, F2, and F3 in the channel dimension, and concatenates these two results to obtain a feature description. Then, after passing through a 7×7 convolutional layer and a Sigmoid function, a spatial weight coefficient is generated. Multiply the spatial weight coefficient with the original input feature map respectively to obtain the output results, and then input the output results into a convolutional, normalization, and activation function layer respectively to adjust the number of channels to obtain the feature maps F7, F8, and F9.
[0076] 2-2): Obtain the feature maps F5 and F6 output by the enhanced feature extraction module, fuse F5 and F6 by adding pixel points, and then perform upsampling to obtain the feature map F10. Then, first fuse F10 with F7 and perform upsampling to obtain the feature map F11, then fuse F11 with F8, perform upsampling to obtain the feature map F12, and finally fuse F12 with F9 and perform upsampling to output, which is the final pixel-level classification result.
[0077] In step 3, the loss function for network training is defined as follows:
[0078] The loss function consists of a weighted cross-entropy loss Cross Entropy Loss, and the calculation formula is as follows:
[0079]
[0080] Among them, M represents the number of categories, yc is a one-hot vector, and its elements only take two values, 0 and 1. If the category is the same as the sample category, take 1, otherwise take 0. pc represents the probability that the predicted sample belongs to c, and wc represents the weight coefficient, and its calculation formula is:
[0081]
[0082] Among them, N represents the total number of pixels, and Nc represents the number of pixels with the Ground truth category of c.
[0083] In step 4, the remote sensing image to be recognized needs to be cut into pictures of the input size required by the model, put into the model for recognition, and after the recognition is completed, the result picture is stitched together completely.
[0084] Example:
[0085] The method for land classification extraction of the present invention on remote sensing images is carried out in the following manner:
[0086] Step 1: The drone acquires remote sensing images, screens the remote sensing images, selects remote sensing images with diverse categories and rich features, and uses ENVI software to label the cultivated land, orchard, forest, and building areas in the remote sensing images. The areas that are not these regions are the background.
[0087] Step 2: Crop and data augment the labeled images. Cut the remote sensing images and label maps into three-channel images of 512×512 size to form a dataset. The main operations for data augmentation of the images are as follows:
[0088] (1) Rotation: Randomly rotate with the center of the image as the rotation center;
[0089] (2) Translation: The image is randomly translated along the X-axis or Y-axis;
[0090] (3) Scaling: The image is randomly enlarged or reduced according to a ratio to achieve multi-scale training;
[0091] (4) Color jitter: Randomly change the exposure, saturation, and hue of the image to form images under different illuminations and colors;
[0092] (5) Gaussian noise: Add noise to make the image blurred
[0093] Step 3: Input the images into the feature extraction network to obtain feature maps. Input F4 in the backbone feature extraction network into the spatial pooling module, and input F1, F2, and F3 into the composite pooling structure to strengthen feature extraction. Then fuse the results of the two parts.
[0094] Step 4: Pass the result after fusing the spatial pooling module and the composite pooling module through the multi-scale fusion upsampling structure to obtain the final prediction result.
[0095] Table 1 Results of different semantic segmentation methods on the remote sensing dataset
[0096]
[0097] The ablation experiment compared five commonly used algorithms for image segmentation and compared the algorithm of the present invention. The present invention has little difference in the IoU of water bodies and vegetation from other algorithms, but the IoU of the present invention is better than other algorithms for the two types of object categories with particularly complex scales and many small-scale targets, namely buildings and cars.
[0098] Step 5: When training the network, the loss function is defined as follows:
[0099] The loss function consists of a weighted cross-entropy loss Cross Entropy Loss, and the calculation formula is as follows:
[0100]
[0101] Among them, M represents the number of categories, yc is a one-hot vector, and its elements can only take two values, 0 and 1. If the category is the same as the sample category, it takes 1, otherwise it takes 0. pc represents the probability that the predicted sample belongs to c, and wc represents the weight coefficient. Its calculation formula is:
[0102]
[0103] Among them, N represents the total number of pixels, and Nc represents the number of pixels with the Ground truth category of c.
[0104] Step 6: Cut the remote sensing image of the area that needs to be classified and identified for land into pictures with the input size required by the model. Put the pictures to be identified into the model for identification. After the identification is completed, splice the result pictures into a complete one.
Claims
1. A remote sensing land segmentation method based on multi-scale attention and lightweight feature fusion, characterized in that It includes the following steps: S1. Design a remote sensing land use model based on a multi-scale attention module and a feature enhancement layer; S2. Process the publicly available dataset to establish a model dataset; S3. Method for validating and comparing the model using the dataset: Conduct a comparative experiment on the land use model proposed in S1 to verify the accuracy of the improved model.
2. The remote sensing land segmentation method based on multi-scale attention and lightweight feature fusion according to claim 1, wherein For the land use segmentation method ES-RT Net based on multiple attention semantic segmentation in S1, it specifically includes: S11. Construction of the basic framework Res-net model; S12. Construction of the global feature optimization module SA2SPP module; S13. Construction of the edge detection multi-scale module EDMS.
3. The remote sensing land segmentation method based on multi-scale attention and lightweight feature fusion according to claim 2, characterized in that, The global feature optimization module SA2SPP includes two components: the atrous spatial pyramid pooling ASPP and the Transformer encoder module.
4. The remote sensing land segmentation method based on multi-scale attention and lightweight feature fusion according to claim 1, characterized in that, S2 is specifically: Randomly crop and scale the original pictures into images of size 1024*1024 to make a dataset, and divide it into a training set, a validation set, and a test set according to a ratio.
5. The remote sensing land segmentation method based on multi-scale attention and lightweight feature fusion according to claim 4, wherein The datasets used in S2 are the ISPRS Vaihingen dataset and the Potsdam dataset. The images are all cropped into size 1024*1024, with a total of 6 land use types, namely water body, building, low vegetation, tree, car, and background, and the colors are respectively labeled as white, blue, cyan, green, yellow, and red; The cross-entropy loss function is used for training.
6. The remote sensing land segmentation method based on multi-scale attention and lightweight feature fusion according to claim 1, wherein The comparative models used in S3 are FCN, U-net, Seg-net, Deeplabv3+. All settings are the same, the batchsize is set to 4, and the cross-entropy loss function is used for training.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the steps in the land use classification method based on multiple attention semantic segmentation as described in any one of claims 1 to 6.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the land use classification method based on multiple attention semantic segmentation as described in any one of claims 1 to 6.