Image semantic segmentation method for air-ground inspection of unmanned aerial vehicle
By adopting a streamlined deep learning model structure and an optimized segmentation algorithm in the air-ground inspection of drones, combined with the encoder decoder architecture of ResNet and Transformer, the fast and accurate segmentation of air-ground inspection images is achieved, solving the shortcomings in real-time and accuracy of traditional models, and improving patrol efficiency and quality.
Patent Information
- Application Number
- CN202510269035.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional deep learning models are difficult to achieve efficient real-time image segmentation in drone air-ground inspections, resulting in excessive delay and calculation costs, limiting the real-time and accuracy of inspections.
By streamlining the model structure and optimizing the segmentation algorithm, a ResNet-based encoder and a Transformer-based decoder are used, combined with a global-local attention module and feature refinement module, to achieve fast and accurate segmentation of air-ground patrol images.
It realizes rapid and accurate segmentation of drone air-ground patrol images, reduces dependence on computing resources and storage space, improves patrol efficiency and quality, and is suitable for the actual deployment and application of drone air-ground patrols.
Smart Images

Figure CN120198665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) aerial and ground inspection, and particularly to an image semantic segmentation method for UAV aerial and ground inspection. Background Art
[0002] UAV aerial and ground inspection refers to using UAVs to conduct surface inspection and monitoring tasks. Compared with traditional ground inspections or inspections by manned aircraft, it has many advantages such as rapidity, flexibility, and safety. It can be deployed and execute tasks in a short time, avoiding direct contact of personnel with dangerous areas such as high altitudes, steep terrains, and toxic environments, thus ensuring the safety of inspection personnel. The semantic segmentation method for UAV aerial and ground inspection can be used to perform fine classification and recognition on the aerial and ground inspection images collected by UAVs, improve the intelligent level of UAV inspections, provide more refined and real-time image information for aerial and ground inspections, and assist in the efficient execution and decision-making support of inspection tasks.
[0003] In recent years, the continuous progress in the field of deep learning has greatly promoted the research on semantic segmentation. Convolutional neural network (CNN) is a neural network specifically designed to process data with grid structures, and has achieved extensive applications and remarkable results in the fields of image processing, computer vision, natural language processing, etc. It simulates the way the human brain processes visual information through multiple convolutional and pooling operations, automatically extracts various features such as edges, textures, and shapes in the image, thereby realizing the understanding and analysis of the image, and these features are used to distinguish different semantic regions in subsequent segmentation tasks. Fully Convolutional Networks (FCN) is based on CNN and realizes pixel-level semantic segmentation of the input image by replacing the fully connected layer with a convolutional layer, using transposed convolution for upsampling, introducing a skip connection structure, and optimizing the loss function, etc., greatly improving the segmentation accuracy and paving the way for deep learning-based semantic segmentation.
[0004] Due to the limited computing resources carried by UAVs, traditional deep learning models often have difficulty achieving efficient real-time processing while maintaining high-precision segmentation. This results in excessive delays and computational costs in image segmentation during actual inspections, restricting the real-time performance and accuracy of inspections. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides an image semantic segmentation method for UAV aerial and ground inspection, which realizes fast and accurate segmentation of aerial and ground inspection images by streamlining the model structure and optimizing the segmentation algorithm, while reducing the dependence on computing resources and storage space, promoting the intelligent and automated development of UAVs in the field of aerial and ground inspection, and improving the inspection efficiency and quality.
[0006] An image semantic segmentation method for UAV aerial and ground inspection, comprising the following steps:
[0007] Step 1: Preprocess the UAVid dataset for UAV remote sensing semantic segmentation;
[0008] Specifically: Cut the image data into small pieces of 1024*1024, and randomly divide them into a validation set and a training set.
[0009] Step 2: Use an encoder based on ResNet to extract image features;
[0010] Specifically: Use four residual blocks composed of convolution and pooling operations to gradually extract features of the image at different scales, and obtain four feature maps of different scales;
[0011] Step 3: Use a decoder based on Transformer to generate a segmentation map;
[0012] The decoder includes three global-local Transformer modules and a feature refinement module, where the global-local Transformer module is a global-local combined attention module, and the global-local combined attention module includes a global attention module and a local attention module; the feature refinement module uses multi-scale spatial feature extraction and context-aware attention mechanism to enhance the expression ability of important features.
[0013] Step 3.1: The global attention module expands the number of channels of the feature map output by the encoder to three times through 1×1 convolution, generates Query, Key and Value, and implements the Self-Attention operation of Transformer based on this, and then enhances the expression ability of global context information through the axial attention module; the local attention module uses two parallel convolution operations with sizes of 3×3 and 1×1 respectively to extract local context information, and attaches BatchNorm and ReLU activation functions after the two convolutions respectively, and finally fuses different local features to obtain a local feature map;
[0014] Step 3.2: Concatenate the outputs of the global attention module and the local attention module, and use 1×1 convolution to map the concatenated features to the output semantic features;
[0015] Step 3.3: The feature refinement module first processes the feature map generated by the first residual block of the encoder, uses three parallel convolution branches to extract spatial features, and each convolution branch is followed by Batch Norm and ReLU activation functions to enhance the non-linear expression ability. Concatenate the output features of the three branches to generate multi-scale spatial features that capture more spatial information.
[0016] Step 3.4: Add the semantic features obtained in Step 3.2 to the multi-scale spatial features to generate fused features; perform global average pooling and local depthwise separable convolution operations on the fused features, and use 1×1 convolution and Sigmoid activation function to generate global attention weights and local attention weights respectively, then add the global and local attention weights to generate context-aware attention weights; multiply the attention weights by the fused features to enhance the expression ability of important features, adjust the number of channels through 1×1 convolution, and upsample to generate the final segmentation map.
[0017] The beneficial effects produced by adopting the above technical solutions are as follows:
[0018] The present invention provides an image semantic segmentation method for UAV aerial-ground inspection. This method is used to segment the UAV aerial high-definition urban landscape dataset UAVid, and evaluate the segmentation accuracy, mIoU (i.e., the overlap degree between the predicted segmentation result and the actual label), segmentation efficiency, viewing experience, etc. of the segmentation result. The results show that compared with other lightweight segmentation models in recent years, the method of the present invention shows better performance in the accuracy of segmenting different entities, especially the segmentation effect of the regions of interest that need to be focused on during the inspection process, is suitable for actual deployment and application in the context of UAV aerial-ground inspection, and can meet the requirements for the segmentation accuracy of the model in the inspection task. Description of the Drawings
[0019] Figure 1 It is the decoder structure diagram of the present invention;
[0020] Figure 2 It is the global-local attention module structure diagram of the present invention;
[0021] Figure 3 It is the feature refinement module structure diagram of the present invention. Detailed Embodiments
[0022] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0023] An image semantic segmentation method for UAV aerial-ground inspection includes the following steps:
[0024] Step 1: Preprocess the UAV remote sensing semantic segmentation dataset UAVid;
[0025] Specifically: Cut the image data into 1024*1024 small pieces for model training, and randomly divide it into a validation set and a training set.
[0026] Step 2: Use an encoder based on ResNet to extract image features, such asFigure 1 as shown;
[0027] Specifically: Four residual blocks composed of convolutional and pooling operations are used to gradually extract features of the image at different scales, obtaining four feature maps at different scales, which serve as the input for the subsequent decoder;
[0028] Step 3: Use a Transformer-based decoder to generate a segmentation map;
[0029] The decoder includes three global-local Transformer modules and a feature refinement module. The global-local Transformer module is a global-local combined attention module as Figure 2 shown. The global-local combined attention module includes a global attention module and a local attention module; The feature refinement module is as Figure 3 shown, using multi-scale spatial feature extraction and context-aware attention mechanism to enhance the expression ability of important features.
[0030] The shallow features generated by the first residual block of the encoder retain rich spatial detail information, while the features output by the global-local Transformer module are rich in semantic information. Although directly summing these two features is fast, it will reduce the segmentation accuracy. The feature refinement module aims to fuse the feature map containing the spatial details of the image generated by the first residual block of the encoder with the feature map rich in semantic information obtained after deep encoding and decoding, so as to achieve the complementarity of spatial details and semantic information.
[0031] Step 3.1: The global attention module expands the number of channels of the feature map output by the encoder to three times through a 1×1 convolution, generates Query, Key, and Value, and based on this, implements the Self-Attention operation of the Transformer. Then, through the axial attention module, the expression ability of global context information is enhanced; The local attention module uses two parallel convolutional operations with sizes of 3×3 and 1×1 respectively to extract local context information, and attaches BatchNorm and ReLU activation functions after the two convolutions respectively. Finally, different local features are fused to obtain a local feature map;
[0032] Step 3.2: Concatenate the outputs of the global attention module and the local attention module, and use a 1×1 convolution to map the concatenated features to the output semantic features;
[0033] Step 3.3: The feature refinement module first processes the feature map generated by the first residual block of the encoder. It uses three parallel convolutional branches to extract spatial features. After each convolutional branch, there are Batch Norm and ReLU activation functions to enhance the non-linear expression ability. The output features of the three branches are concatenated to generate multi-scale spatial features that capture more spatial information.
[0034] Step 3.4: Add the semantic features obtained in Step 3.2 to the multi-scale spatial features to generate fused features; perform global average pooling and local depthwise separable convolution operations on the fused features, and use 1×1 convolution and Sigmoid activation function to generate global attention weights and local attention weights respectively. Then add the global and local attention weights to generate context-aware attention weights; multiply the attention weights by the fused features to enhance the expression ability of important features, adjust the number of channels through 1×1 convolution, and upsample to generate the final segmentation map.
[0035] In this embodiment, in order to significantly improve the efficiency and accuracy of semantic segmentation while ensuring the lightweight of the model, an encoder-decoder architecture is constructed, which combines the ResNet and Transformer methods. The model performance is further enhanced by introducing an attention mechanism; local fine features and global context features are applied to achieve a more accurate and detailed segmentation effect, so as to realize the efficient processing of video data. Both the encoder and the decoder contain four processing stages, corresponding to four residual blocks and four decoding processes respectively. The encoder is mainly responsible for feature extraction, and the decoder is responsible for feature fusion and segmentation map generation.
[0036] The encoder uses ResNet as the basic network, and gradually extracts image features through four residual blocks to generate four feature maps; the four feature maps are temporarily stored after convolution and pooling operations for feature fusion in the subsequent decoding process. This design draws on the idea of four feature fusions in U-net to ensure that feature information at different scales is fully utilized.
[0037] The decoder part corresponds to the encoder and also contains four decoding processes. In this embodiment, the decoder consists of three global-local Transformer modules and one feature refinement module, as Figure 1 shown. The global-local Transformer module is used to extract image features at different scales and fuse them with the feature maps generated by the encoder to achieve a comprehensive capture of global and local features. The feature refinement module uses the spatial details and semantic information of the image to generate a segmentation map through steps such as multi-scale spatial feature extraction, feature fusion, and depthwise separable convolution;
[0038] The global-local Transformer module introduces a global-local combined attention module, which not only retains the advantages of the Transformer in extracting global environmental information but also preserves the local spatial details of the image through convolutional operations. This design enables the model to simultaneously focus on global and local features in the segmentation task, thereby achieving better results in the segmentation task.
[0039] The global-local attention module is the core component of the global-local Transformer module, and its structure is as Figure 2 shown. This module is divided into a global feature attention module based on the Transformer and a local feature attention module based on convolution. The global feature attention module expands the number of channels through 1×1 convolution and uses self-attention operations to capture global context information. The local feature attention module uses two parallel 3×3 and 1×1 convolution operations to extract local context information. After concatenating the results of the global and local attention modules, the result of the global-local attention is output through a 1×1 convolutional layer.
[0040] The feature refinement module aims to fuse the feature map containing the spatial details of the image generated by the first residual block of the encoder with the feature map rich in semantic information obtained through deep encoding and decoding. This module first extracts multi-scale spatial features from the feature map generated by the first residual block of the encoder, fuses the results with the feature map obtained from the global-local Transformer module, then introduces a context-aware attention mechanism to focus on global semantic information and local detail information and enhance the expression ability of important features. Finally, the final segmentation map is generated through a 1×1 convolutional layer and upsampling. This design ensures that the model does not lose the spatial detail information of the image while focusing on global semantic information, thereby achieving a more accurate segmentation effect.
[0041] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An image semantic segmentation method for UAV air-ground inspection, characterized in that: The following steps are involved: Step 1: Preprocess the UAV remote sensing semantic segmentation dataset UAVid; Step 2: Extract image features using a ResNet-based encoder; Step 3: Generate segmentation map using Transformer-based decoder.
2. The image semantic segmentation method for UAV air-ground inspection according to claim 1 is characterized in that: The step 1 is specifically as follows: cutting the image data into small blocks of 1024*1024, and randomly dividing them into a validation set and a training set.
3. The image semantic segmentation method for UAV air-ground inspection according to claim 1 is characterized in that: The step 2 is specifically as follows: using four residual blocks consisting of convolution and pooling operations to gradually extract features of the image at different scales to obtain feature maps of four different scales.
4. The image semantic segmentation method for UAV air-ground inspection according to claim 1 is characterized in that: The decoder in step 3 includes three global-local Transformer modules and one feature refinement module.
5. The image semantic segmentation method for UAV air-ground inspection according to claim 4 is characterized in that: The global-local Transformer module is a global-local combined attention module, which includes a global attention module and a local attention module; the feature refinement module uses multi-scale spatial feature extraction and context-aware attention mechanism to enhance the expressiveness of important features.
6. The image semantic segmentation method for UAV air-ground inspection according to claim 5 is characterized in that: The step 3 comprises the following steps: Step 3.1: The global attention module expands the number of channels of the feature map output by the encoder to three times through 1×1 convolution, generates Query, Key and Value, and implements the Self-Attention operation of Transformer based on this, and then enhances the expression ability of global context information through the axial attention module; the local attention module uses two parallel convolution operations of size 3×3 and 1×1 to extract local context information, and adds Batch Norm and ReLU activation functions after the two convolutions, and finally fuses different local features to obtain local feature maps; Step 3.2: Concatenate the outputs of the global attention module and the local attention module, and use 1×1 convolution to map the concatenated features into the output semantic features; Step 3.3: The feature refinement module first processes the feature map generated by the first residual block of the encoder and uses three parallel convolution branches to extract spatial features. Each convolution branch is followed by a batch norm and a ReLU activation function to enhance the nonlinear expression capability. The output features of the three branches are concatenated to generate multi-scale spatial features that capture more spatial information. Step 3.4: Add the semantic features obtained in step 3.2 and the multi-scale spatial features obtained in step 3.3 to generate fused features; perform global average pooling and local depth-separable convolution operations on the fused features to generate the final segmentation map.
7. The image semantic segmentation method for UAV air-ground inspection according to claim 6 is characterized in that: As described in step 3.4, the fused features are subjected to global average pooling and local depth-separable convolution operations, specifically: 1×1 convolution and Sigmoid activation function are used to generate global attention weights and local attention weights respectively, and then the global and local attention weights are added together to generate context-aware attention weights; the attention weights are multiplied by the fused features to enhance the expressiveness of important features, the number of channels is adjusted by 1×1 convolution, and the final segmentation map is generated by upsampling.