A method, medium and device for real-time semantic segmentation of unmanned aerial vehicle remote sensing images
By designing lightweight feature extraction backbone network, multiple semantic association activation module and difficult category semantic enhancement module on the drone platform, the problems of low accuracy and poor real-time segmentation of the drone remote sensing image are solved, and efficient and flexible real-time semantic segmentation effect is achieved.
Patent Information
- Application Number
- CN202210778558.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-06-30
AI Technical Summary
The semantic segmentation accuracy of drone remote sensing images is low, poor real-time, and the model is difficult to deploy on a drone computing platform with low memory/low computing power.
A real-time semantic segmentation method for remote sensing images of drones was designed, including building a lightweight feature extraction backbone network, multiple semantic association activation modules, and difficult-class semantic enhancement modules, extracting multi-scale output feature maps and performing semantic association activation, and finally performing real-time semantic segmentation on edge computing devices.
It realizes efficient real-time semantic segmentation of drone remote sensing images on drone platforms with low memory and limited computing power, improving the agility and deployment flexibility of the model, while ensuring high segmentation accuracy and real-timeness.
Smart Images

Figure CN115100552B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video image processing, and in particular to a method, medium and device for real-time semantic segmentation of unmanned aerial vehicle remote sensing images. Background Art
[0002] Semantic segmentation is an important task in the field of computer vision. The goal of semantic segmentation is to predict pixel-level labels based on the semantic information represented by the pixels of an image, which can be considered as a dense classification problem. In recent years, convolutional neural networks have made great progress and have been proven to have significant effects in semantic segmentation tasks.
[0003] The semantic association technology applied in convolutional neural networks aims to mine and extract the implicit information in images or feature maps to assist in tasks such as semantic segmentation.
[0004] The application of remote sensing images taken by drones is of great significance. With the help of photography equipment installed on drones, low-altitude high-resolution aerial images can be collected more conveniently and economically. Drones can fly close to the ground to improve the resolution of the objects photographed. These features enable drone remote sensing images to distinguish detailed objects such as non-motor vehicles and pedestrians.
[0005] Real-time semantic segmentation technology applied on UAV platforms refers to the use of edge computing devices on UAV platforms close to the data source to provide computing power for real-time interpretation of remote sensing images and other needs, perform real-time semantic segmentation, achieve faster response, and meet the needs of subsequent UAV autonomous reconnaissance, autonomous action and other tasks.
[0006] However, due to the need to output pixel-by-pixel semantic labels, existing high-resolution semantic segmentation technologies consume a large amount of computing power, resulting in poor real-time performance. At the same time, larger models make it difficult to deploy models on drone computing platforms with low memory or computing power. If smaller models are deployed, the accuracy will be low. Summary of the invention
[0007] In order to solve the problems existing in the above-mentioned prior art, the present invention provides a method, medium and equipment for real-time semantic segmentation of UAV remote sensing images, which is intended to solve the problems of low accuracy, poor real-time performance and difficult model deployment of UAV remote sensing image segmentation.
[0008] A real-time semantic segmentation method for UAV remote sensing images comprises the following steps:
[0009] Step 1: Construct a UAV remote sensing image semantic segmentation training dataset and annotate the UAV remote sensing images in the dataset with semantic labels;
[0010] Step 2: Build a lightweight feature extraction backbone network for the semantic segmentation model from image to pixel-by-pixel category label, and extract multi-scale output feature maps of UAV remote sensing images based on the lightweight feature extraction backbone network;
[0011] Step 3: construct a multiple semantic association activation module of the semantic segmentation model, wherein the multiple semantic association activation module includes a spatial semantic association activation unit and a channel semantic association activation unit; connect the spatial semantic association activation unit and the channel semantic association activation unit in parallel to the lightweight feature extraction backbone network; connect the multi-scale output feature map output by the backbone network and the feature map output by the multiple semantic association activation module in parallel to obtain a discriminant feature map;
[0012] Step 4: Construct a semantic label discrimination module of the semantic segmentation model, and construct a difficult category semantic enhancement module after the semantic label discrimination module; the semantic label discrimination module discriminates the semantic label of each pixel in the discrimination feature map to obtain a semantic segmentation label; then, based on the difficult category semantic enhancement module, difficult semantic category enhancement is performed to obtain an optimized semantic segmentation discrimination result;
[0013] Step 5: Use the UAV remote sensing image semantic segmentation training dataset to train the semantic segmentation model and obtain the initial weight of the semantic segmentation model;
[0014] Step 6: Input the UAV remote sensing image. First, preprocess the input UAV remote sensing image on the edge computing device, and then input the preprocessed UAV remote sensing image into the semantic segmentation model for a forward propagation to obtain the pixel-by-pixel semantic category of the UAV remote sensing image.
[0015] Preferably, the lightweight feature extraction backbone network includes 6 convolution combinations connected to each other in a cascade manner, each convolution combination includes 4 convolution layers, 1 inactivation layer, 1 batch normalization layer, and 1 linear rectification layer, and each convolution combination is connected to a maximum pooling layer for reducing the spatial scale.
[0016] Preferably, the specific steps of extracting the multi-scale output feature map of the UAV remote sensing image based on the lightweight feature extraction backbone network are as follows:
[0017] After the UAV remote sensing image is input into the lightweight feature extraction backbone network, it is processed by 6 cascaded convolution combinations in sequence, and the feature maps of the fourth convolution combination, the fifth convolution combination and the sixth convolution combination are upsampled by 2 times, 4 times and 8 times respectively through successive deconvolution; the feature maps output by the upsampled fourth convolution combination, the fifth convolution combination and the sixth convolution combination are spliced in parallel with the feature map output by the third convolution combination to obtain a multi-scale output feature map.
[0018] Preferably, the spatial semantic association activation unit is composed of a convolution operation and a softmax function; the calculation process includes feature matrix deformation, feature matrix multiplication, pixel-to-pixel semantic association extraction, and spatial semantic association enhancement.
[0019] Preferably, the channel semantic association activation unit is composed of a convolution operation and a softmax function; the calculation process includes feature matrix deformation, feature matrix multiplication, inter-channel semantic feature extraction and channel semantic association enhancement.
[0020] Preferably, the specific process of the semantic label discrimination module for discriminating the semantic label of each pixel in the discriminant feature map is as follows:
[0021] The discriminant feature map is input into the fully connected layer of the semantic label recognition module to obtain the discriminant feature map. For each pixel of the discriminant feature map, the softmax function is used to calculate the posterior probability p of the pixel belonging to each category i. i :
[0022]
[0023] Where: x i is the posterior probability that pixel x belongs to the i-th category; take the category with the largest posterior probability As the semantic segmentation label of the pixel position.
[0024] Preferably, the specific steps of the difficult category semantic enhancement module for performing difficult semantic category enhancement are as follows:
[0025] The semantic segmentation labels are counted through the difficult category semantic enhancement module to obtain the area ratios of different categories of objects. The categories that account for less than 20% of the total image area in the statistical results are regarded as rare difficult categories.
[0026] According to the discriminant feature map, the difficulty coefficient is calculated for the surrounding pixels of the difficult category:
[0027]
[0028] In the formula: avg ,s max ,s min are the average area occupied by each category, the area occupied by the largest category, and the area occupied by the rarest category; a i is the difficulty coefficient; s i is the area occupied by the ith category;
[0029] Correct the discrimination result of pixels at the edge of the difficult category to The optimized semantic segmentation discrimination result map is obtained, and then the semantic segmentation discrimination result map is upsampled n times through the interpolation method to restore the same size as the UAV remote sensing image. The semantic segmentation label is matched one by one with the pixels of the UAV remote sensing image to realize the semantic segmentation of the UAV remote sensing image.
[0030] Preferably, the step 5 comprises the following steps:
[0031] The semantic segmentation model is pre-trained based on the UAV remote sensing image semantic segmentation training dataset. After the pre-training is completed, the pre-trained semantic segmentation model is fine-tuned based on the aerial photography UAV dataset related to the task. After the fine-tuning is completed, the semantic segmentation model structure file and weight value file are generated.
[0032] A storage medium for storing computer instructions, wherein the computer instructions are used to enable the computer to execute a semantic segmentation model constructed by any of the methods described above.
[0033] An electronic device comprises at least one processor, in which a semantic segmentation model constructed by any one of the above methods is deployed, so that the at least one processor can execute the semantic segmentation model constructed by any one of the above methods.
[0034] The beneficial effects of the present invention include:
[0035] 1. The image semantic segmentation model of the present invention reduces the computational complexity and parameter quantity of the model through a lightweight backbone network design, thereby ensuring the real-time performance of the semantic segmentation while still ensuring the high performance of the model and improving the agility of the model, so that it can be more flexibly deployed on edge computing devices with low memory and limited computing power on drone platforms, thereby realizing real-time semantic segmentation of drone remote sensing images.
[0036] 2. The spatial semantic association activation module introduced in the present invention can extract the spatial semantic association between each position in the feature map and other positions, characterize the spatial semantic relationship by assigning different weights, and map the association back to the original feature map, so that the positions with high spatial semantic association in the feature map tend to maintain similar features, ensuring the high integrity of the segmentation results and avoiding unreasonable or fragmented segmentation results.
[0037] 3. The channel semantic association activation module introduced in the present invention can extract the semantic association between feature map channels, extract the association between channel features, express it in the form of weights, and map it back to the original feature map, thereby introducing channel semantic association into the feature map, mining better feature descriptions, and ensuring the accuracy of the segmentation results.
[0038] 4. The difficult category semantic enhancement module introduced in the present invention can count and analyze the rare difficult categories in the discrimination results, find relatively rare but more important difficult samples, perform more refined semantic segmentation on such objects, and add the loss to the back propagation, so that the model focuses on important samples and ensures the completeness and credibility of the segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flow chart of the overall method of the present invention;
[0040] Figure 2 This is a structural diagram of the backbone network for extracting image semantic segmentation features of the present invention;
[0041] Figure 3 It is a structure diagram of the spatial semantic association activation unit of the present invention;
[0042] Figure 4 It is a structural diagram of the channel semantic association activation unit of the present invention;
[0043] Figure 5 This is a structural diagram of the difficult category semantic enhancement module of the present invention;
[0044] Figure 6 It is a structural diagram of the multiple semantic association activation and semantic label discrimination module of the present invention;
[0045] Figure 7 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0046] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0047] The following is combined with Figure 1 To the attached Figure 7 The embodiments of the present invention are further described in detail:
[0048] like Figure 1 As shown, a real-time semantic segmentation method for UAV remote sensing images includes the following steps:
[0049] Step 1: Construct a UAV remote sensing image semantic segmentation training dataset and annotate the UAV remote sensing images in the dataset with semantic labels;
[0050] Step 2: Build a lightweight feature extraction backbone network for the semantic segmentation model from image to pixel-by-pixel category label, and extract multi-scale output feature maps of UAV remote sensing images based on the lightweight feature extraction backbone network;
[0051] like Figure 2 As shown, the lightweight feature extraction backbone network includes 6 convolution combinations connected to each other in a cascade manner, each convolution combination includes 4 convolution layers, 1 inactivation layer, 1 batch normalization layer, and 1 linear rectification layer, and each convolution combination is connected to a maximum pooling layer for reducing the spatial scale.
[0052] The specific steps of extracting the multi-scale output feature map of the UAV remote sensing image based on the lightweight feature extraction backbone network are as follows:
[0053] After the UAV remote sensing image is input into the lightweight feature extraction backbone network, it is processed by 6 cascaded convolution combinations in sequence, and the feature maps of the fourth convolution combination, the fifth convolution combination and the sixth convolution combination are upsampled by 2 times, 4 times and 8 times respectively through successive deconvolution; the feature maps output by the upsampled fourth convolution combination, the fifth convolution combination and the sixth convolution combination are spliced in parallel with the feature map output by the third convolution combination to obtain a multi-scale output feature map.
[0054] Step 3: construct a multiple semantic association activation module of the semantic segmentation model, wherein the multiple semantic association activation module includes a spatial semantic association activation unit and a channel semantic association activation unit; connect the spatial semantic association activation unit and the channel semantic association activation unit in parallel to the lightweight feature extraction backbone network; connect the multi-scale output feature map output by the backbone network and the feature map output by the multiple semantic association activation module in parallel to obtain a discriminant feature map;
[0055] like Figure 3 As shown, the spatial semantic association activation unit is composed of a convolution operation and a softmax function; the calculation process includes feature matrix deformation, feature matrix multiplication, pixel semantic association extraction, and spatial semantic association enhancement, as described below:
[0056] Assume that the number of channels of the multi-scale output feature map is c, the spatial scale is h×w, and the multi-scale output feature map of size c×h×w is deformed into Then transpose to get the size The feature map of The multi-scale output feature map of size and size is The feature maps of c×h×w are matrix-multiplied, and after the softmax function, a spatial semantic feature map of size (h×w)×(h×w) is obtained, which represents the semantic association of each pixel with other pixels. The multi-scale output feature map of size c×h×w is multiplied with the spatial semantic feature map of size (h×w)×(h×w) to enhance the spatial semantic association, and finally a new set of feature maps of size c×h×w is obtained.
[0057] like Figure 4 As shown, the channel semantic association activation unit is composed of a convolution operation and a softmax function; the calculation process includes feature matrix deformation, feature matrix multiplication, inter-channel semantic feature extraction, and channel semantic association enhancement, as described below:
[0058] Transpose the c×h×w multi-scale output feature map to a h×w×c feature map. Perform matrix multiplication of the c×h×w multi-scale output feature map and the h×w×c feature map. After the softmax function, obtain a c×c channel semantic feature map to represent the semantic association between each channel and other pixels. Multiply the c×h×w multi-scale output feature map with the c×c channel semantic feature map to enhance the spatial semantic association, and finally obtain a new set of c×h×w feature maps.
[0059] Since the multiple semantic association activation module includes a spatial semantic association activation unit and a channel semantic association activation unit, the discriminant feature map is generated by connecting in parallel the multi-scale output feature map output by the backbone network, the feature map output by the spatial semantic association activation module (a feature map of size c×h×w), and the feature map output by the channel semantic association activation module (a feature map of size c×h×w); that is, the size of the discriminant feature map obtained after the three sets of feature maps are connected in parallel becomes (3c)×h×w;
[0060] Step 4: Construct a semantic label discrimination module of the semantic segmentation model, and construct a difficult category semantic enhancement module after the semantic label discrimination module; the semantic label discrimination module discriminates the semantic label of each pixel in the discrimination feature map to obtain a semantic segmentation label; then, based on the difficult category semantic enhancement module, difficult semantic category enhancement is performed to obtain an optimized semantic segmentation discrimination result;
[0061] The specific process of the semantic label discrimination module for discriminating the semantic label of each pixel in the discriminant feature map is as follows:
[0062] The discriminant feature map of size (3c)×h×w is input into the fully connected layer of the semantic label discrimination module to obtain a discriminant feature map of size n×h×w, where n is the number of categories that may appear in the image; then, for each pixel of the discriminant feature map of size n×h×w, the softmax function is used to calculate the posterior probability p of the pixel belonging to each category i. i :
[0063]
[0064] Where: x i is the posterior probability that pixel x belongs to the i-th category; take the category with the largest posterior probability As the semantic segmentation label of the pixel position.
[0065] The semantic segmentation labels are counted through the difficult category semantic enhancement module to obtain the area ratio S = {s1, s2, ..., s n}, where n is the total number of categories, and the categories that account for less than 20% of the total image area in the statistical results are considered rare and difficult categories;
[0066] According to the discriminant feature map, the difficulty coefficient is calculated for the surrounding pixels of the difficult category:
[0067]
[0068] Where: s avg ,s max ,s min are the average area occupied by each category, the area occupied by the largest category, and the area occupied by the rarest category; a i is the difficulty coefficient; s i is the area occupied by the ith category;
[0069] Correct the discrimination result of pixels at the edge of the difficult category to The optimized semantic segmentation discrimination result map is obtained, and then the semantic segmentation discrimination result map is upsampled 8 times through the interpolation method to restore the same size as the UAV remote sensing image. The semantic segmentation label is matched one by one with the pixels of the UAV remote sensing image to realize the semantic segmentation of the UAV remote sensing image.
[0070] Step 5: Use the UAV remote sensing image semantic segmentation training dataset to train the semantic segmentation model and obtain the initial weight of the semantic segmentation model;
[0071] The UAV remote sensing image semantic segmentation training data set includes the image data set ImageNet, the UAV semantic segmentation data set UAVid and UDD6;
[0072] The step 5 comprises the following steps:
[0073] First, the semantic segmentation model is pre-trained using the image dataset ImageNet, and then further trained using the drone semantic segmentation datasets UAVid and UDD6. After the training is completed, the pre-trained semantic segmentation model is fine-tuned based on the task-related aerial photography drone dataset; after the fine-tuning is completed, the semantic segmentation model structure file and weight value file are generated.
[0074] Step 6: Input the UAV remote sensing image. First, pre-process the input UAV remote sensing image on the edge computing device, and then input the pre-processed UAV remote sensing image into the semantic segmentation model for a forward propagation to obtain the pixel-by-pixel semantic category of the UAV remote sensing image.
[0075] The semantic segmentation model trained in step 5 is deployed in the edge computing device on the drone. After the imaging device takes the remote sensing image, it is input into the semantic segmentation model for forward propagation, and the pixel-by-pixel semantic category of the image is output. The specific process is as follows: first, the input UAV remote sensing image is preprocessed, the size is adjusted to 512*512, and geometric correction and image defogging are performed; then the image is input into the feature extraction backbone network, and the outputs of the 3rd, 4th, 5th, and 6th convolution combinations are extracted. After 1x (unprocessed), 2x, 4x, and 8x deconvolution, the feature maps of size 64*64 are uniformly adjusted. After combination, they are input into the multiple semantic association activation module, and the spatial semantic association activation unit and the channel semantic association activation unit are respectively performed. The feature maps output by the two units are combined with the multi-scale output feature maps, and input into the semantic label discrimination module. After the output convolution layer, the fully connected layer, and the softmax function, the discrimination result of semantic segmentation is obtained, and then it is input into the difficult category semantic enhancement module, the discrimination result of difficult ground object samples is optimized, and 8x upsampling is performed to make the semantic segmentation result correspond to each pixel of the original image, and the final semantic segmentation result map is obtained, thereby realizing the semantic segmentation of UAV remote sensing images.
[0076] A storage medium for storing computer instructions, wherein the computer instructions are used to enable the computer to execute the semantic segmentation model constructed by the method described above.
[0077] An electronic device comprises at least one processor, on which a semantic segmentation model constructed by the above method is deployed, so that the processor can execute the semantic segmentation model constructed by the above method.
[0078] For example, the edge computer (edge computing module) of the drone platform is used as the processor; the specific implementation method is as follows:
[0079] The drone platform is used to carry imaging modules and edge computing modules for aerial photography. It is required to be able to carry the weight of other modules and adjust its position and posture at will in the air to meet aerial photography needs.
[0080] The aerial imaging module is used to obtain remote sensing images of the ground during the flight of the UAV. It has a controllable gimbal and autofocus device, which can remotely control its shooting angle and perform autofocus according to imaging needs.
[0081] The edge computing module is used to decode, preprocess and semantically segment the aerial images in real time. The edge computing module includes a video decoder, a neural computing processor and a memory. The edge computing module deploys a deep learning-based drone remote sensing semantic segmentation model, divides the decoded aerial video into images, runs the semantic segmentation program and obtains the semantic segmentation results.
[0082] The communication module receives flight control commands, lens adjustment commands, and shooting commands, and sends the semantic segmentation results to the ground control console after the semantic segmentation calculation is completed. The communication module maintains two-way uninterrupted communication with the ground control console, receives flight control commands, lens adjustment commands, and shooting commands issued by the ground control console, commands the UAV platform to fly, adjust the lens posture, and shoot remote sensing images; after the edge computing module completes the calculation, the semantic segmentation results are sent to the ground control console.
[0083] The power module provides power for other modules. It is composed of high-capacity and high-discharge rate model aircraft batteries. It supplies power to the drone platform, aerial imaging module, edge computing module, and communication module through power lines.
[0084] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A real-time semantic segmentation method for UAV remote sensing images, characterized in that: The following steps are involved: Step 1: Construct a UAV remote sensing image semantic segmentation training dataset and annotate the UAV remote sensing images in the dataset with semantic labels; Step 2: Build a lightweight feature extraction backbone network for the semantic segmentation model from image to pixel-by-pixel category label, and extract multi-scale output feature maps of UAV remote sensing images based on the lightweight feature extraction backbone network; Step 3: construct a multiple semantic association activation module of the semantic segmentation model, wherein the multiple semantic association activation module includes a spatial semantic association activation unit and a channel semantic association activation unit; The spatial semantic association activation unit and the channel semantic association activation unit are connected in parallel to the lightweight feature extraction backbone network; the multi-scale output feature map output by the backbone network and the feature map output by the multiple semantic association activation module are connected in parallel to obtain a discriminative feature map; Step 4: Construct a semantic label discrimination module of the semantic segmentation model, and construct a difficult category semantic enhancement module after the semantic label discrimination module; the semantic label discrimination module discriminates the semantic label of each pixel in the discrimination feature map to obtain a semantic segmentation label; then, based on the difficult category semantic enhancement module, difficult semantic category enhancement is performed to obtain an optimized semantic segmentation discrimination result; Step 5: Use the UAV remote sensing image semantic segmentation training dataset to train the semantic segmentation model and obtain the initial weight of the semantic segmentation model; Step 6: Input the UAV remote sensing image. First, pre-process the input UAV remote sensing image on the edge computing device, and then input the pre-processed UAV remote sensing image into the semantic segmentation model for a forward propagation to obtain the pixel-by-pixel semantic category of the UAV remote sensing image. The spatial semantic association activation unit is composed of a convolution operation and a softmax function; the calculation process includes feature matrix deformation, feature matrix multiplication, pixel semantic association extraction and spatial semantic association enhancement; The channel semantic association activation unit is composed of a convolution operation and a softmax function; the calculation process includes feature matrix deformation, feature matrix multiplication, inter-channel semantic feature extraction, and channel semantic association enhancement; The specific process of the semantic label discrimination module for discriminating the semantic label of each pixel in the discriminant feature map is as follows: The discriminant feature map is input into the fully connected layer of the semantic label recognition module to obtain the discriminant feature map. For each pixel of the discriminant feature map, the softmax function is used to calculate the posterior probability p of the pixel belonging to each category i. i : Where: x i is the posterior probability that pixel x belongs to the i-th category; take the category with the largest posterior probability As the semantic segmentation label of the pixel position; The specific steps of the difficult category semantic enhancement module for performing difficult semantic category enhancement are as follows: The semantic segmentation labels are counted through the difficult category semantic enhancement module to obtain the area ratios of different categories of objects. The categories that account for less than 20% of the total image area in the statistical results are regarded as rare difficult categories. According to the discriminant feature map, the difficulty coefficient is calculated for the surrounding pixels of the difficult category: Where: s avg ,s max ,s min are the average area occupied by each category, the area occupied by the largest category, and the area occupied by the rarest category; a i is the difficulty coefficient; s i is the area occupied by the ith category; Correct the discrimination result of pixels at the edge of the difficult category to The optimized semantic segmentation discrimination result map is obtained, and then the semantic segmentation discrimination result map is upsampled n times through the interpolation method to restore the same size as the UAV remote sensing image. The semantic segmentation label is matched one by one with the pixels of the UAV remote sensing image to realize the semantic segmentation of the UAV remote sensing image.
2. The method for real-time semantic segmentation of UAV remote sensing images according to claim 1, characterized in that: The lightweight feature extraction backbone network includes 6 convolution combinations connected to each other in a cascade manner, each convolution combination includes 4 convolution layers, 1 inactivation layer, 1 batch normalization layer, and 1 linear rectification layer, and each convolution combination is connected to a maximum pooling layer for reducing the spatial scale.
3. The method for real-time semantic segmentation of UAV remote sensing images according to claim 2, characterized in that: The specific steps of extracting the multi-scale output feature map of the UAV remote sensing image based on the lightweight feature extraction backbone network are as follows: After the UAV remote sensing image is input into the lightweight feature extraction backbone network, it is processed by 6 cascaded convolution combinations in sequence, and the feature maps of the fourth convolution combination, the fifth convolution combination and the sixth convolution combination are upsampled by 2 times, 4 times and 8 times respectively through successive deconvolution; the feature maps output by the upsampled fourth convolution combination, the fifth convolution combination and the sixth convolution combination are spliced in parallel with the feature map output by the third convolution combination to obtain a multi-scale output feature map.
4. The method for real-time semantic segmentation of UAV remote sensing images according to claim 1, characterized in that: The step 5 comprises the following steps: The semantic segmentation model is pre-trained based on the UAV remote sensing image semantic segmentation training dataset. After the pre-training is completed, the pre-trained semantic segmentation model is fine-tuned based on the aerial photography UAV dataset related to the task. After the fine-tuning is completed, the semantic segmentation model structure file and weight value file are generated.
5. A storage medium, characterized in that: Used to store computer instructions, wherein the computer instructions are used to enable the computer to execute the semantic segmentation model constructed by the method described in any one of claims 1 to 4.
6. An electronic device, characterized in that: It includes at least one processor, in which a semantic segmentation model constructed by the method described in any one of claims 1 to 4 is deployed, so that the at least one processor can execute the semantic segmentation model constructed by the method described in any one of claims 1 to 4.