A crop segmentation method for multispectral UAV remote sensing images

By improving the DeepLabv3+ model, adding vegetation index channels and dual-channel attention modules, the problems of low recognition accuracy and fuzzy edge recognition in the crop classification of drone remote sensing images are solved, and higher classification accuracy and edge recognition effects are achieved.

CN116503590BActive Publication Date: 2025-05-23NORTHWEST A & F UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310132989.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-17
Publication Date
2025-05-23
Estimated Expiration
2043-02-17

AI Technical Summary

Technical Problem

The existing crop classification methods for remote sensing images of drone crops have problems with low recognition accuracy and fuzzy edge recognition.

Method used

The multi-spectral UAV remote sensing image crop segmentation method is used to improve the input layer structure of the DeepLabv3+ model, add NDVI and OSAVI vegetation index channels, and add a dual-channel attention module to the hollow space pyramid pooling module to enhance the model's utilization of spectral and spatial information.

Benefits of technology

The recognition accuracy and edge recognition effect of crop classification are improved, and the defects of the deep learning model insufficient utilization of spectral feature information and less attention to image channel information and position information are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116503590B_ABST
    Figure CN116503590B_ABST
Patent Text Reader

Abstract

The present invention provides a method for crop segmentation of multispectral UAV remote sensing images, including: step one, multispectral remote sensing image data processing. Step two, vegetation index extraction. Step three, building a data set. Step four, building a deep semantic segmentation model: adding a convolutional attention module to the Deeplabv3+ model to obtain a deep semantic segmentation model. Step five, model training and preservation. Step six, crop segmentation. In terms of spectral information, the present invention improves the input layer structure of the DeepLabv3+ model, adds NDVI normalized vegetation index and OSAVI optimized soil regulated vegetation index channels, and increases the spectral differences between different crops. In terms of spatial information, the present invention adds a dual-channel attention module including channel attention and spatial attention to the void space pyramid pooling module, expands the model receptive field while increasing the correlation of pixel information, optimizes the boundaries of classification results, and thus improves model accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of agricultural remote sensing technology, relates to crop intelligent classification technology, and specifically relates to a multi-spectral unmanned aerial vehicle remote sensing image crop segmentation method. Background Art

[0002] Rapid crop identification based on remote sensing images plays an important role in estimating the area of ​​planting areas and disaster forecasting. At present, the use of remote sensing technology to identify crops mainly includes two methods: satellites and drones. However, high-altitude satellites have a long re-entry period and a lot of noise, and must be supplemented by ground sampling surveys to achieve fine extraction of crop classification. Low-altitude drone remote sensing has the characteristics of flexibility, strong timeliness and high resolution, and has been widely used to obtain agricultural information at the farmland level.

[0003] Traditional crop remote sensing classification methods mainly use machine learning methods such as maximum likelihood, decision trees, support vector machines, and random forests. In addition to conventional K-means clustering and maximum likelihood methods, remote sensing segmentation methods have also developed machine learning methods such as random forests and support vector machines. Machine learning methods for farmland crop identification are mainly based on manually selected features, and have a high dependence on feature channels. Manually selected features are difficult, so this method has certain limitations. In recent years, with the development of deep learning semantic segmentation technology, it has made great breakthroughs in the field of image segmentation, and deep learning has also been widely used in the agricultural field. However, the current classification of crops in drone remote sensing images still has low recognition accuracy and blurred edge recognition. Therefore, it is urgent to propose a crop classification model with high recognition accuracy. There are several shortcomings in the existing research and technology:

[0004] (A) In terms of band information, less attention is paid to and less attention is paid to the full extraction of multispectral band information in the extraction of land objects. However, the addition of vegetation index and the fusion of vegetation index and other features in network iteration can highlight the differences between different land objects.

[0005] (B) In terms of spatial information, deep learning models require a deeper hierarchical structure to expand the model's receptive field. Building a model with higher accuracy requires a large increase in the number of model parameters. If an attention mechanism is added to the deep learning model, the global information of the image can be selectively aggregated to some important information, thereby increasing the weight of important information.

[0006] (C) The existing methods of remote sensing impact segmentation have the problem of low detection accuracy, which is difficult to meet the accuracy requirements of land survey methods. Some algorithms cannot achieve high detection accuracy while reducing the number of parameters. Summary of the invention

[0007] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a multispectral UAV remote sensing image crop segmentation method to solve the technical problems of low recognition accuracy and blurred edge recognition in the prior art UAV remote sensing image crop classification method.

[0008] In order to solve the above technical problems, the present invention adopts the following technical solutions to achieve the above problems:

[0009] A method for crop segmentation of multispectral UAV remote sensing images, the method comprising the following steps:

[0010] Step 1: Multispectral remote sensing image data processing.

[0011] Step 2: Extract vegetation index.

[0012] Step 3: Build the dataset.

[0013] Step 4: Build a deep semantic segmentation model:

[0014] Add a convolutional attention module to the DeeplabV3+ model to obtain a deep semantic segmentation model.

[0015] Step 5: Model training and saving:

[0016] The deep semantic segmentation model obtained in step 4 is trained using the training sample dataset obtained in step 3 and then saved.

[0017] Step 6: Crop segmentation:

[0018] The deep semantic segmentation model trained in step 5 is used to segment crops in the remote sensing image data obtained by multispectral drone photography.

[0019] Compared with the prior art, the present invention has the following technical effects:

[0020] (I) In terms of spectral information, the present invention improves the input layer structure of the DeepLabv3+ model, adds NDVI normalized vegetation index and OSAVI optimized soil adjusted vegetation index channels, and increases the spectral differences between different crops.

[0021] (II) In terms of spatial information, the present invention adds a dual-channel attention module including channel attention and spatial attention to the atrous spatial pyramid pooling module, thereby expanding the receptive field of the model while increasing the correlation of pixel information, optimizing the boundaries of the classification results, and thus improving the model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a multispectral drone remote sensing image.

[0023] Figure 2 It is a model structure diagram.

[0024] Figure 3 It is the network structure diagram of the convolutional attention module.

[0025] Figure 4 This is the experimental result diagram.

[0026] The specific contents of the present invention are further explained in detail below in conjunction with embodiments. DETAILED DESCRIPTION

[0027] It should be noted that, unless otherwise specified, all software, modules and layers in the present invention are software, modules and layers known in the prior art.

[0028] The present invention specifically relates to a farmland crop classification model based on multispectral unmanned aerial vehicle remote sensing images. By improving the model input structure, a DeeplabV3+ model suitable for including multispectral channels and vegetation index channels is constructed, and an attention mechanism is added to the model to improve the classification accuracy of the model, thereby overcoming the defects of the deep learning model in insufficient utilization of spectral feature information and less attention to channel information and position information of the image.

[0029] Based on the DeepLabv3+ semantic segmentation network, the present invention proposes a method for crop recognition in multispectral UAV remote sensing images, which solves the problems of low recognition accuracy and blurred edge recognition in previous crop classification of UAV remote sensing images. The deep learning model is improved to focus on and fully extract the information of each multispectral band, and integrate features such as vegetation index in network iteration to highlight the differences between different objects. Adding an attention mechanism to the deep learning model can selectively aggregate the global information of the image to some important information, increase the weight of important information, overcome the defects of the deep learning model that does not fully utilize the spectral feature information and pays less attention to the channel information and position information of the image, and increase the segmentation accuracy of the object.

[0030] In accordance with the above technical scheme, specific embodiments of the present invention are given below. It should be noted that the present invention is not limited to the following specific embodiments, and all equivalent changes made on the basis of the technical scheme of this application fall within the protection scope of the present invention.

[0031] Example:

[0032] This embodiment provides a method for crop segmentation of multispectral UAV remote sensing images, which includes the following steps:

[0033] Step 1: Multispectral remote sensing image data processing:

[0034] Step 101, data capture:

[0035] Remote sensing image data obtained by shooting with a multispectral drone.

[0036] In this embodiment, the multispectral UAV flies at an altitude of 130 meters (426 ft), with a heading overlap rate of 85% and a lateral overlap rate of 70%. It takes pictures according to the predetermined flight trajectory points to obtain the following Figure 1 The remote sensing image shown.

[0037] Step 102, two-dimensional data reconstruction:

[0038] The remote sensing image data obtained in step 101 is imported into Pix4Dmapper software for two-dimensional reconstruction, and after the reconstruction is completed, image data of five bands are obtained respectively.

[0039] The five bands are blue, green, red, red edge and near infrared.

[0040] Step 2: Extract vegetation index:

[0041] The vegetation index includes the normalized vegetation index and the optimized soil-adjusted vegetation index, and the calculation formula is as follows:

[0042] NDVI=(NIR-R) / (+R)

[0043] OSAVI=(NIR–R) / (NIR+R+0.16)

[0044] Where:

[0045] NDVI stands for Normalized Difference Vegetation Index;

[0046] OSAVI stands for Optimized Soil Adjusted Vegetation Index;

[0047] NIR stands for near infrared band;

[0048] R represents the red band.

[0049] The calculated normalized vegetation index, the optimized soil adjusted vegetation index and the image data of the five bands obtained in step 102 are fused to obtain multispectral UAV remote sensing image data of seven channels, so that the data contains data of more bands.

[0050] Step 3: Build the dataset:

[0051] According to the drone images of the shooting location and the local crop planting distribution, the crop categories are divided according to actual needs; based on the actual ground survey and visual interpretation, the remote sensing images are manually vectorized and annotated in ArcGIS10.5 software to generate labeled images with the same size and spatial resolution as the original images; the labeled images are used through data augmentation technology to obtain sufficient training sample data sets, and the training sample data sets are divided into training set, test set and validation set in a ratio of 6:2:2.

[0052] In this embodiment, in order to improve the convergence speed of the model and reduce the amount of calculation, the above drone images and label data are cropped into training data images of 512 pixels × 512 pixels at an overlap rate of 0.5, and the drone images and label images are named the same. Data augmentation technology is used to generate differentiated copies, while the enhanced features remain unchanged, more valid data is generated, and the images are rotated, flipped, mirrored, translated, and other data augmentation methods are used to obtain sufficient training samples.

[0053] In this embodiment, the training set is used to train the model, and the test set and the validation set are used to validate the generalization ability of the model.

[0054] Step 4: Build a deep semantic segmentation model:

[0055] Add a convolutional attention module to the DeeplabV3+ model to obtain a deep semantic segmentation model.

[0056] DeeplabV3+ model description:

[0057] The Deeplab series was first proposed by Chen et al. in 2015. It was improved on the basis of the VGG-16 skeleton network. DeeplabV1 proposed convolution with holes. The hole convolution can capture image context information at multiple scales, which can improve the receptive field while maintaining a high spatial resolution, and is more conducive to semantic segmentation tasks. Since then, the DeepLab series has continuously integrated new achievements in the field of semantic segmentation. Inspired by the Pyramid Scene Parsing Network (PSPNet), DeepLabv2 was proposed on the ResNet skeleton network. The Astrous Spatial Pyramid Pooling (ASPP) was used to sample the input layer with hole convolutions at different sampling rates to capture image context at multiple scales. Inspired by the encoder-decoder structure of networks such as U-Net, a decoder was added to DeepLabv3 to refine the segmentation results, and deep separable convolutions were applied to ASPP and the decoder to obtain DeepLabv3+. The model structure diagram is shown below. Figure 2 shown.

[0058] The DeepLabv3+ model uses atrous convolutions and atrous spatial pyramid pooling with different dilation rates to obtain more scale information. It uses a simple and effective decoder module to improve the model's ability to process image boundaries and better preserve the target's edge detail information. It is suitable for solving the problem of different sizes of farmland and poor edge effects. Therefore, this network is selected as the basis for designing a basic model for multispectral UAV image crop classification.

[0059] In this embodiment, the model is improved by introducing a convolutional attention mechanism module and a vegetation index channel into the DeepLabv3+ model. The convolutional attention module is an attention mechanism module that combines the two dimensions of the channel attention module and the spatial attention module, referred to as the CBAM model. The model structure diagram is shown in Figure 3 As shown in the figure, the channel attention module and the spatial attention module are used to adaptively select and enhance the input feature map in the channel direction and the spatial direction respectively. It models the correlation between feature channels, strengthens the main features, and suppresses irrelevant features, so that the network pays more attention to the content information and location information of the target to be detected, thereby improving the detection accuracy of the network.

[0060] The description of the convolutional attention module is as follows:

[0061] The convolutional attention model obtains the feature map F after the input multispectral remote sensing image is fused with vegetation index and processed by the backbone network. Then, the feature map F is subjected to maximum pooling and average pooling respectively to generate two different spatial information description feature maps F. max and F avg , and then describe the spatial information feature map F max and F avg The shared network compresses the spatial dimension of the input feature map, sums and merges the elements, and then calculates the channel attention map At through the ReLU activation function. The specific calculation formula of the channel attention map At is shown below:

[0062] At(F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))

[0063] =σ(W1(W0(F avg ))+W1(W0(F max )))

[0064] Where:

[0065] σ represents the ReLU activation function;

[0066] AvgPool represents the average pooling operation;

[0067] MaxPool represents the maximum pooling operation;

[0068] MLP means the shared network consists of multi-layer perceptrons;

[0069] W0, W1 represent the weights of the multi-layer perceptron.

[0070] Then the input feature map F is fused with the channel attention map At to obtain the feature channel map F with channel attention weights T , F T In the spatial attention module, the channel information of the function map is first aggregated through average pooling and maximum pooling operations to generate two 2D maps. The two 2D maps are then subjected to 1×1 convolution operations and ReLU activation functions to obtain the spatial attention map A. K , spatial attention map A K The specific calculation formula is as follows:

[0071] A k =σ(f7×7(AvgPool(F);MaxPool(F)))

[0072] =σ(f7×7(F avg ; F max ))

[0073] Where:

[0074] f7×7 represents a 7×7 convolution operation.

[0075] Finally, the feature channel map F T With the spatial attention map A K The output feature map F including channel weights and position weights is obtained by fusion R , the overall process is shown as follows:

[0076]

[0077]

[0078] Where:

[0079] Represents element-wise multiplication.

[0080] This step makes full use of the spectral feature information and increases attention to the channel information and position information of the image, making an important contribution to improving the algorithm's recognition accuracy.

[0081] Step 5: Model training and saving:

[0082] The deep semantic segmentation model obtained in step 4 is trained using the training sample dataset obtained in step 3 and then saved.

[0083] Step 501, model training:

[0084] The deep learning in this experiment runs under the Windows 10 operating system. The computer is equipped with two NVIDIA GeForceRTX 2080Ti graphics cards. The network is built using the Tensorflow 1.14.0 deep learning framework. The software environment is Anaconda (python 3.6). The constructed multispectral drone remote sensing image dataset is input into the deep learning model for training and learning. Different learning rates, batch sizes, loss functions, and optimizer hyperparameters are set to optimize the model. Early stopping technology can be used in training. If the accuracy of the training set does not improve within 30 rounds, or the accuracy of the validation set decreases, the model will automatically stop the training process when overfitting occurs.

[0085] Step 502, model saving:

[0086] Select multiple models with high accuracy and verify them through validation set data respectively. Select the models with better generalization results, higher precision and clearer boundaries. Save the models and apply them directly to actual production to provide guiding suggestions for agricultural planning.

[0087] Step 6: Crop segmentation:

[0088] The deep semantic segmentation model trained in step 5 is used to segment crops in the remote sensing image data obtained by multispectral drone photography.

[0089] Application examples:

[0090] This application example provides a multispectral UAV remote sensing image crop segmentation method based on the above embodiment, which specifically includes the following contents.

[0091] Experimental part:

[0092] (1) Experimental data:

[0093] The test site is located in the Yangling Demonstration Zone of Shaanxi Province, which is located in the central part of the Guanzhong Plain of Shaanxi Province. The main high-quality crops in the demonstration zone include soybeans, wheat, corn and kiwifruit. In this experiment, a multispectral version of the DJI Phantom 4RTK drone is used, equipped with 6 1 / 2.9-inch CMOS, including 1 color sensor for visible light imaging and 5 monochrome sensors for multispectral imaging. The spectral range of the monochrome sensor is shown in Table 1. The effective pixels of each single sensor are 2.08 million (total pixels are 2.12 million), and the maximum resolution of the photo is 1600×1300 (4:3.25). In mid-September 2021, the area near the Innovation Park of the Yangling Modern Agriculture Demonstration Zone was photographed. The drone flew at an altitude of 130 meters (426FT), with a heading overlap rate of 85% and a lateral overlap rate of 70%. The photos were taken according to the predetermined flight trajectory points. A total of 2574 original images were captured, with a total of 429 shooting points. Each shooting point includes a visible light color photo and 5 single-channel images in different bands. For the original image, the positions of the six camera sensor centers relative to the phase center of the airborne D-RTK antenna have been compensated in the photo EXIF ​​coordinates, and then spliced ​​using Pix4Dmapper software to obtain five bands of data images. The five bands of data obtained are fused to obtain five channels of drone data.

[0094] Table 1 Spectral range of monochromatic sensors for drones

[0095]

[0096] The experiment uses three drone remote sensing images as shown in the study area map. The three experimental data are roughly cut into training sets and test sets in an 8:2 ratio. The training set is used to train the model, and the test set is used to verify the generalization ability of the model. After cutting, the image size of 5800×5000 is obtained as the test set.

[0097] According to the drone images of the shooting location and the local crop planting distribution, the crops are divided into three types, namely soybean, corn, and kiwifruit. Other crops, roads, weeds, bare land, houses and other features are uniformly classified as background. According to the actual ground survey combined with visual interpretation, the remote sensing images are manually vectorized and annotated in ArcGIS10.5, with soybean fields annotated as 3, corn fields annotated as 2, kiwifruit fields annotated as 1, and other backgrounds annotated as 0, to generate label images with the same size and spatial resolution as the original image. According to the map of the study area, it can be seen that there are few corn data samples. In order to ensure the balance of sample data, some corn images are selected for replication in the experiment. In order to improve the convergence speed of the model and reduce the amount of calculation, the above drone images and label data are cropped into training data images of 512 pixels × 512 pixels with an overlap rate of 0.5, and the drone images and label images are named the same, forming a total of 672 image-label pairs. In order to obtain enough training samples, data augmentation technology is used to generate differentiated copies, which generates more effective data while maintaining the enhanced features. The images were rotated, flipped, mirrored, translated, and other data enhancement methods were used to obtain a total of 5460 image-label pairs. In order to make the images seen by the model random when training data, 5460 non-repeating random numbers were generated using Python, and the original image-label was renamed as a random number, and then divided into a training set and a validation set in a 9:1 ratio, where the training set contains 4800 image-label pairs and the validation set contains 660 image-label pairs. The multispectral data set was converted to form a standard data set.

[0098] (2) Experimental plan:

[0099] The deep learning in this paper is run under the Windows 10 operating system. The computer is equipped with two NVIDIA GeForce RTX2080Ti graphics cards. The network is built using the Tensorflow1.14.0 deep learning framework. The software environment is Anaconda (python3.6). The initial learning rate of the model is 0.01. If the accuracy does not improve within 30 rounds during training, the learning rate is reduced to 0.01 times the original. The Adam optimizer is used with a batch size of 6. The early stopping technique is used in training. If the accuracy of the training set does not improve within 25 rounds, or the accuracy of the validation set decreases, and overfitting occurs, the model automatically stops the training process.

[0100] Comparative Example 1:

[0101] This comparative example provides a crop segmentation method for UAV remote sensing images. The difference between this method and the above application example is that this comparative example removes the vegetation index extraction module in step 2 and the convolutional attention mechanism module in step 4 in the above application example, and the other steps are the same as the above application example.

[0102] Comparative Example 2:

[0103] This comparative example provides a crop segmentation method for UAV remote sensing images. The difference between this method and the above application example is that this comparative example removes the vegetation index extraction module in step 2 of the above application example, and the other steps are the same as the above application example.

[0104] Comparative Example 3:

[0105] This comparative example provides a crop segmentation method for UAV remote sensing images. The difference between this method and the above application example is that this comparative example removes the convolutional attention mechanism module in step four of the above application example, and the other steps are the same as the above application example.

[0106] The crop segmentation effect of the multispectral UAV remote sensing imagery crop segmentation method given in the above application example of the present invention is compared with the semantic segmentation algorithm given in Comparative Examples 1 to 3 for UAV imagery.

[0107] Table 2 UAV image crop segmentation results of different comparison methods

[0108]

[0109] Table 2 shows the experimental results of Comparative Example 1, Comparative Example 2, Comparative Example 3 and the example of the present invention. According to the experimental results of Comparative Example 1, Comparative Example 3 and the example of the present invention in Table 2, it can be seen that after deleting the introduced vegetation index structure, the model will identify some corns with similar texture features to kiwifruit as kiwifruit, while the crop classification results of the model without deleting the vegetation index module are more accurate. Because the vegetation index channel can amplify the differences between different crops and between crops and weeds, and resist the interference of texture features, the model after adding the vegetation index module can more accurately distinguish different crops, making the classification accuracy of the model higher. According to the experimental results of Comparative Example 1, Comparative Example 2 and the example of the present invention in Table 2, it can be seen that the ASPP structure of the DeepLabv3+ model will cause discontinuous model segmentation, incomplete farmland segmentation, and hollow phenomena in some areas when using different expansion rates. The experimental effect of the DeepLabv3+ model after introducing the convolutional attention mechanism is better, and the farmland is more complete. Because the convolutional attention model can adaptively learn different crop characteristics, strengthen the characteristics of different crop feature categories, highlight the interconnected feature maps and promote specific farmland characteristics. At the same time, the convolutional attention mechanism can also adaptively select and enhance the feature maps in terms of channels and positions, selectively aggregate the global information of the image to some important information, and increase the weight of important information, thereby overcoming the incomplete farmland segmentation phenomenon and improving the classification accuracy of the model.

[0110] Comparative Example 4:

[0111] Compared with other methods, the machine learning model of the support vector machine model is compared with the deep learning models of the U-Net model and the SegNet model. The experiments all use multispectral UAV remote sensing images of Yangling District. The crop classification method of the support vector machine uses the supervised learning SVM (Support Vector Machine Classification) function provided by the ENVI software to manually select the region of interest to ensure that the regions of interest corresponding to the three crops represent each crop as much as possible. For other categories, the regions of interest should involve various landforms, such as houses, weeds, bare land, and paths. In addition, try to keep the regions of interest between categories balanced. For the five deep learning models, the standard data sets created as described above are used, and the input layers of the existing SegNet model and U-Net model are modified so that the models can input five-channel multispectral UAV data, and then the model training is performed.

[0112] Table 3 Comparative Example 4 UAV image crop segmentation results of different methods

[0113]

[0114] Table 3 and Figure 4 The crop segmentation effects of different methods in Example 4 are shown in FIG. Figure 4 From the overall experimental results on the test set and the visualization results of three different scenes, it can be seen that the classification results of the SVM algorithm are relatively messy, and it cannot identify and classify crops well. The segmentation of sparsely planted farmland in the study area is incomplete, and weeds are easily regarded as crops. Deep learning can effectively extract the texture features of the image through operations such as convolution and pooling, and the farmland in the segmentation map is relatively complete. Among them, the U-Net model and SegNet can only segment crops with small reflectivity differences in the image. When the reflectivity difference inside the planting area increases, more holes appear, and the effect of extracting the planting edge is poor, such as Figure 4 As shown in (e) and (f) in the figure, the reflectivity difference in the crop planting area is large, which will cause holes in the segmentation results. At the same time, there are many prediction errors at the edge of the farmland, and the edge extraction effect is fuzzy. Figure 4In the three scenes (c) and (d), the edge extraction effect of DeepLabv3+ model and improved DeepLabv3+ model is better than SegNet, because both models use dilated convolution and multi-scale strategies, which have a larger receptive field. The large receptive field enables the deep learning model to recognize simple targets and classify sparse feature targets well. In addition, the large receptive field can better restore the edge information of objects, making farmland identification more complete. Deeper layers can extract high-level semantic features, and thus better represent targets than simple artificial features, but will mispredict some kiwifruit planting areas, such as Figure 4 As shown in the kiwifruit growing areas in scenes (1) and (3) in (d), the top of the kiwifruit growing area is not fully recognized. The improved DeepLabv3+ model proposed in this paper achieves the best results in the crop classification task, as shown in Figure 2. Figure 4 As shown in (c), the improved DeepLabv3+ achieves better classification results in most areas by utilizing more detailed information in the image and the spatial correlation of pixels in a large range. The vegetation index module and convolutional attention module introduced in the model effectively increase the differences between objects and resist the interference of texture features. The farmland is relatively complete and unified, the edge recognition effect is slightly improved, and there are fewer missed and misclassified phenomena. Although the false detection rate of most texture-similar areas is reduced, some areas with large spectral information differences will be missed, such as Figure 4 As shown in scene (2) in (c), the soybean in the lower right corner cannot be completely identified, but the overall segmentation accuracy of the improved model is improved.

Claims

1. A method for crop segmentation of multispectral UAV remote sensing images. The following steps are involved: Step 1: Multispectral remote sensing image data processing: Step 101, data capture: Remote sensing image data obtained by shooting with multispectral drones; Step 102, two-dimensional data reconstruction: The remote sensing image data obtained in step 101 is imported into Pix4Dmapper software for two-dimensional reconstruction, and after the reconstruction is completed, image data of five bands are obtained respectively; The five bands are blue, green, red, red edge and near infrared; Step 2: Extract vegetation index: The vegetation index includes the normalized vegetation index and the optimized soil-adjusted vegetation index, and the calculation formula is as follows: Where: represents the normalized difference vegetation index; represents the optimized row soil conditioned vegetation index; represents the near infrared band; represents the red band; The calculated normalized difference vegetation index, the optimized soil adjusted vegetation index and the image data of the five bands obtained in step 102 are fused to obtain multispectral UAV remote sensing image data of seven channels; Step 3: Build a data set; Step 4: Build a deep semantic segmentation model: Add a convolutional attention module to the DeeplabV3+ model to obtain a deep semantic segmentation model; The description of the convolutional attention module is as follows: The convolutional attention model obtains the feature map of the input multispectral remote sensing image after vegetation index fusion and backbone network processing. F , then the feature map F Perform maximum pooling and average pooling respectively to generate two different spatial information description feature maps F max and F avg , and then describe the feature map with spatial information F max and F avg The shared network compresses the spatial dimension of the input feature map, sums and merges the elements, and then calculates the channel attention map through the ReLU activation function. At , channel attention map At The specific calculation formula is shown below: Where: ReLU activation function. represents the average pooling operation; Represents the maximum pooling operation; MLP Represents a shared network consisting of multi-layer perceptrons; W0, W1 represent the weights of the multilayer perceptron; Then fuse the input feature map 𝐹 with the channel attention map At to obtain a feature channel map with channel attention weights F T , F T In the spatial attention module, the channel information of the function map is first aggregated through average pooling and maximum pooling operations to generate two 2D maps. The two 2D maps are then subjected to 1×1 convolution operations and ReLU activation functions to obtain the spatial attention map. A K , spatial attention map A K The specific calculation formula is as follows: Where: f 7×7 represents a 7×7 convolution operation; Finally, the feature channel map F T With spatial attention map A K The output feature map including channel weight and position weight is obtained by fusion F R , the overall process is shown as follows: Where: Represents element-wise multiplication.

2. The multispectral UAV remote sensing image crop segmentation method according to claim 1, It is characterized in that The specific process of step three includes: According to the drone images of the shooting location and the local crop planting distribution, the crop categories are divided according to actual needs; based on the actual ground survey and visual interpretation, the remote sensing images are manually vectorized and annotated in ArcGIS10.5 software to generate labeled images with the same size and spatial resolution as the original images; the labeled images are used through data augmentation technology to obtain sufficient training sample data sets, and the training sample data sets are divided into training set, test set and validation set in a ratio of 6:2:

2.

3. The multispectral UAV remote sensing image crop segmentation method according to claim 1, It is characterized in that It also includes step 5, model training and saving: The deep semantic segmentation model obtained in step 4 is trained using the training sample dataset obtained in step 3 and then saved.

4. The multispectral UAV remote sensing image crop segmentation method according to claim 3, It is characterized in that It also includes step six, crop segmentation: The deep semantic segmentation model trained in step 5 is used to segment crops in the remote sensing image data obtained by multispectral drone photography.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on pyramid segmentation attention module

    CN113807210A

  • Crop planting information extraction method based on convolutional neural network

    CN115331104A