A method and system for classifying objects using hyperspectral remote sensing technology from unmanned aerial vehicles
By constructing a spatial spectral attention feature extraction network model, the problems of insufficient computational efficiency and memory storage of UAV hyperspectral remote sensing images are solved, and efficient ground object classification and low-cost processing are achieved, which is suitable for embedded systems.
Patent Information
- Application Number
- CN202310950554.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-07-31
AI Technical Summary
Existing technologies cannot meet the computational efficiency and memory storage requirements of UAV hyperspectral remote sensing images, resulting in low floating-point operation efficiency per second, which in turn affects the accuracy of ground object classification.
A spatial-spectral attention feature extraction network model is constructed, using a lightweight spatial feature extraction convolution layer and a spectral attention mechanism module. Through partial convolution and channel attention mechanism, redundant calculations and memory access are reduced, and the feature extraction and classification accuracy are improved.
While maintaining the classification accuracy of hyperspectral remote sensing images, the computational cost and energy consumption are significantly reduced, making it possible to quickly process the classification on a GPU or CPU, making it suitable for embedded systems.
Smart Images

Figure CN116958696B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) hyperspectral remote sensing image data processing technology, and more particularly to a method and system for classifying land objects in UAV hyperspectral remote sensing images. Background Art
[0002] Rich spatial and spectral information greatly enhances image perception, making hyperspectral remote sensing technology widely used in various fields such as military reconnaissance, resource exploration, environmental monitoring, and disaster assessment. Compared with traditional remote sensing imagery, unmanned aerial vehicle (UAV) hyperspectral (HSI) remote sensing imagery has higher spectral resolution and is usually presented in the form of a data cube. UAV hyperspectral remote sensing imagery has higher spectral resolution than traditional remote sensing imagery and is usually presented in the form of a data cube. UAV hyperspectral remote sensing imagery has the advantages of short cycle time, wide coverage, rich data sources, rapid and repeatable deployment, and time and labor saving. It can obtain high spectral resolution, high spatial resolution, and multi-temporal data in mesoscale areas, balancing the requirements of accuracy and efficiency. It is an unparalleled data source for studying fine-scale land cover and will play an important role in monitoring grassland degradation in relatively small areas. It is becoming a superior supplement to traditional ground-based monitoring and aerial and satellite remote sensing.
[0003] Classifying objects in drone-generated hyperspectral remote sensing images involves labeling each pixel based on its spectral and spatial information. This is one of the most critical techniques in hyperspectral data analysis. Each hyperspectral image scene consists of hundreds of narrow, continuous spectral bands, making extracting information from HSI data a complex and computationally challenging task. Convolutional neural networks (CNNs) are currently widely used for HSI classification.
[0004] However, due to the large number of internal network parameters in hyperspectral images, the computational efficiency and memory storage requirements cannot be met, resulting in low floating-point operations per second (FLOPs) efficiency under frequent memory access and a large number of computing operations, which in turn leads to poor hyperspectral land object classification accuracy. Summary of the Invention
[0005] In response to the problems existing in the above-mentioned fields, the present invention provides a method and system for unmanned aerial vehicle hyperspectral remote sensing object classification, which can solve the technical problem that the computing efficiency and memory storage requirements cannot be met, resulting in low floating-point operations per second (FLOPs) efficiency under frequent memory access and a large number of computing operations, and thus leading to poor hyperspectral object classification accuracy.
[0006] To solve the above technical problems, the present invention discloses a method for classifying objects using hyperspectral remote sensing by an unmanned aerial vehicle, comprising the following steps:
[0007] Acquire remote sensing images of the object to be measured and classified, pre-process the acquired remote sensing images, and obtain processed hyperspectral remote sensing image data;
[0008] Construct a spatial spectral attention feature extraction network model, input the trained hyperspectral remote sensing image data into the spatial spectral attention feature extraction network model, extract the deep semantic feature map; based on the extracted deep semantic feature map, obtain the deep semantic information of the hyperspectral remote sensing image;
[0009] The spatial-spectral attention feature extraction network model includes a lightweight spatial feature extraction convolutional layer and a spectral attention mechanism module. The lightweight spatial feature extraction convolutional layer uses partial convolution to extract spatial features to obtain a feature map of deep spatial semantic information of the hyperspectral remote sensing image. The spectral attention mechanism module captures the nonlinear information between channels in the feature map of deep spatial semantic information.
[0010] According to the deep semantic information of the obtained hyperspectral remote sensing image, the deep semantic feature map of the remote sensing image for the object classification is classified, and the object classification result of the hyperspectral remote sensing image is output.
[0011] Preferably, the preprocessing of the acquired remote sensing image comprises the following steps:
[0012] Determine the attribute information and size information of the UAV hyperspectral remote sensing data to be extracted;
[0013] According to the attribute information and size information of the UAV hyperspectral remote sensing data, the hyperspectral data is unified into the MAT format, and the image is cropped and set to the ratio consistent with the training data.
[0014] Preferably, the cropped image size includes 7 pixels×7 pixels, 9 pixels×9 pixels, and 11 pixels×11 pixels.
[0015] Preferably, obtaining the processed hyperspectral remote sensing image data comprises the following steps:
[0016] Normalize the hyperspectral remote sensing image data in the training set;
[0017] The deep semantic feature maps of the normalized UAV hyperspectral remote sensing image data in the training set are extracted and input into the spatial spectral attention feature extraction network to obtain the deep semantic information of the hyperspectral remote sensing images in the training set.
[0018] Preferably, obtaining a feature map of deep spatial semantic information of a hyperspectral remote sensing image comprises the following steps:
[0019] After preprocessing and normalization, the hyperspectral remote sensing data F m, input it into the spatial feature extraction module, and use partial convolution to perform regular convolution on only part of the input channels to extract spatial features, while the remaining channels are not processed;
[0020] To facilitate sequential or regular memory access, place the first or last consecutive c p The channel is calculated as a substitute for the complete feature map, and the floating-point operation number of the network is:
[0021]
[0022] Among them, h and w correspond to the height and width of the feature map, k represents the size of the convolution kernel, and c p Indicates the number of channels selected when implementing partial convolution;
[0023] Set the partial ratio r = c p / c, where c p Represents the number of channels of partial convolution, c represents the number of channels of traditional convolution, r is taken as 1 / 4, and after using partial convolution, the floating-point operation of the network is reduced to 1 / 16 of the traditional convolution;
[0024] Partial convolution requires fewer memory accesses, which is:
[0025]
[0026] When r = 1 / 4, the memory access amount is 1 / 4 of the traditional convolution; partial convolution only performs operations on c p Extract spatial features from the channel and convert the remaining channels cc p reserve.
[0027] Preferably, the method of capturing nonlinear information between channels in a feature map of deep spatial semantic information comprises the following steps:
[0028] Global average pooling and global maximum pooling operations are used to aggregate the spatial information of the feature map to generate two different spatial context descriptors. and Represent global average pooling and global maximum pooling respectively;
[0029] The two descriptors are combined through a 2×1 convolution;
[0030] To improve the generalization ability, a multi-layer perceptron is added to learn the final channel attention map M. To reduce the parameter overhead, the size of the hidden layer is set to C / γ, where γ is the compression ratio.
[0031] The enhanced channel attention computation is expressed as:
[0032]
[0033] Where W0∈R (C / γ)×C ,W1∈R C ,W0 and W1 are the weights of the multilayer perceptron, f 2×1 Indicates that the size of the convolutional layer filter is 2×1.
[0034] Preferably, the classifying of the deep semantic feature maps of the remote sensing images for classifying the measured objects includes inputting the extracted deep semantic feature maps into the trained classifier model, classifying the deep semantic feature maps of the hyperspectral remote sensing images in the training set, and obtaining classification results.
[0035] Preferably, the method further includes calculating the value of the loss function according to the classification results, updating the parameters of the spatial spectral attention feature extraction network through back propagation, and completing the training of the network.
[0036] Preferably, a UAV hyperspectral remote sensing ground object classification system includes:
[0037] The data acquisition module is used to obtain remote sensing images of the object to be measured and classified, and pre-process the obtained remote sensing images to obtain processed hyperspectral remote sensing image data;
[0038] The model building module is used to build a spatial spectral attention feature extraction network model, input the trained hyperspectral remote sensing image data into the spatial spectral attention feature extraction network model, extract the deep semantic feature map; based on the extracted deep semantic feature map, the deep semantic information of the hyperspectral remote sensing image is obtained;
[0039] The spatial-spectral attention feature extraction network model includes a lightweight spatial feature extraction convolutional layer and a spectral attention mechanism module. The lightweight spatial feature extraction convolutional layer uses partial convolution to extract spatial features to obtain a feature map of deep spatial semantic information of the hyperspectral remote sensing image. The spectral attention mechanism module captures the nonlinear information between channels in the feature map of deep spatial semantic information.
[0040] The data output module is used to classify the deep semantic feature map of the remote sensing image for the object classification based on the deep semantic information of the hyperspectral remote sensing image obtained by the model building module, and output the object classification result of the hyperspectral remote sensing image.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] The present invention classifies objects in hyperspectral remote sensing images by constructing a spatial spectral attention feature extraction network model. The lightweight spatial feature extraction convolution layer set uses partial convolution to effectively extract spatial features. The partial convolution operation is used to selectively operate on some input channels, reducing redundant calculations in spatial feature extraction. The spectral attention mechanism module captures the nonlinear information between channels in the feature map from the feature map after spatial feature extraction to the maximum extent, improves the feature extraction and generalization capabilities of the deep neural network, inputs the effective spectral feature map into the feature extraction network, enables the network to focus on more important spectral information, and improves the accuracy of object classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 Schematic diagram of the overall method of the present invention;
[0044] Figure 2 Schematic diagram of the process of the hyperspectral remote sensing image object classification method of the present invention;
[0045] Figure 3 Schematic diagram of the structure of the spatial spectrum feature extraction network model of the present invention;
[0046] Figure 4 Schematic diagram of the structure of the spatial feature extraction module of the present invention;
[0047] Figure 5 Schematic diagram of the structure of the spectral attention mechanism module of the present invention. DETAILED DESCRIPTION
[0048] The following is a combination of the embodiments of the present invention Figure 1-5 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms used in the present invention are only used to describe specific implementation methods and are not intended to limit the present invention.
[0049] Example
[0050] like Figure 1 As shown, the embodiment of the present invention provides a method and system for classifying objects using hyperspectral remote sensing of unmanned aerial vehicles.
[0051] Hyperspectral remote sensing imaging technology is rapidly developing, making it easier to acquire hyperspectral data. For example, it can be collected directly using drone-mounted equipment. Compared to previous methods requiring satellite-based or airborne hyperspectral data acquisition, the cost of acquiring hyperspectral data directly using drone-mounted equipment is significantly lower. However, rapidly processing this massive amount of data in a cost-effective manner remains a challenge.
[0052] To address the above issues, embodiments of the present invention disclose a method for object classification based on hyperspectral remote sensing images from unmanned aerial vehicles (UAVs). This method can quickly and accurately extract effective features from hyperspectral remote sensing images and achieve precise object classification. Furthermore, while maintaining reliable classification results, it significantly reduces computational costs and can be rapidly processed on GPUs and even CPUs, making it a promising option for embedded systems. Furthermore, optimizing global computational costs reduces the need for computing power and memory, resulting in lower energy consumption. This cost reduction helps improve the performance of deep learning models.
[0053] like Figure 1-2 As shown, the present invention provides a method for classifying objects in remote sensing images, comprising the following steps:
[0054] Step S1: Obtain the remote sensing image to be extracted, and extract the deep semantic feature map in the remote sensing image through a feature extractor.
[0055] In step S1, before extracting the deep semantic feature map from the remote sensing image, the remote sensing image to be extracted is preprocessed, including the following steps:
[0056] Determine the attribute information and size information of the hyperspectral remote sensing data to be extracted; convert the hyperspectral data into MAT format according to the attribute information and size information of the hyperspectral remote sensing data, and crop the image; set the proportion of the data used for training.
[0057] The cropped size includes but is not limited to 7 pixels × 7 pixels, 9 pixels × 9 pixels, 11 pixels × 11 pixels, etc., as long as the size is an odd number.
[0058] Step S2: Construct a spatial-spectral attention feature extraction network model, which includes a lightweight spatial feature extraction convolutional layer and a spectral attention mechanism module.
[0059] In step S2, the spatial-spectral attention feature extraction network is trained, including the following steps:
[0060] Step S21: using a portion of the pre-processed hyperspectral remote sensing image to be extracted as a training set;
[0061] Step S22: normalizing the hyperspectral remote sensing image data in the training set;
[0062] Normalized remote sensing image data is used as training data for the spatial-spectral attention feature extraction network, making the network training process more stable. An optional normalization method is batch normalization, which can accelerate the convergence of model training, making the model training process more stable and avoiding gradient explosion or vanishing.
[0063] Step S23: extracting the deep semantic feature map of the normalized remote sensing image data in the training set, and inputting it into the feature extraction network to obtain the deep semantic features of the hyperspectral remote sensing data in the training set;
[0064] Step S24: Use a classifier to classify the extracted deep semantic features to obtain the predicted category to which the input data belongs; calculate the loss based on the label value and the predicted category, and update the network parameters through back propagation to complete the training of the feature extraction network.
[0065] Step S3: Input the pre-processed hyperspectral data into the lightweight spatial feature extraction convolution layer, and use partial convolution to effectively extract spatial features. The spectral attention mechanism module captures the nonlinear information between channels in the feature map; input the effective spectral feature map into the feature extraction network, so that the network focuses on more important spectral information, such as Figure 3 shown.
[0066] In step S3, the spatial feature extraction network consists of two stages:
[0067] (1) Spatial feature extraction stage: Partial convolution (PConv) is used as a means to effectively reduce computational redundancy and optimize memory access, thereby enhancing cost optimization. It only performs regular convolution on some input channels to extract spatial features, while the remaining channels are not processed. In order to facilitate sequential or regular memory access, the first or last consecutive c p It is computed as a proxy for the full feature map. This is a common assumption since the number of channels of the input and output feature maps is similar. As a result, PConv's floating-point operations (FLOPs) are reduced to 1 / 16 of the original.
[0068] (2) Spectral attention feature extraction stage: The attention mechanism is used to capture important nonlinear information between channels in the feature map.
[0069] The traditional attention mechanism is a compression and excitation channel attention method. The compression component uses global average pooling (GAP) to convert the feature map into a one-dimensional vector. The excitation component uses two fully connected (FC) layers to determine the weight of each channel, effectively promoting cross-channel information interaction.
[0070] In this paper, lightweight convolutional layers and a channel attention mechanism are used to enhance the extraction of spatial and spectral information in the network. The proposed spatial-spectral attention feature extraction network significantly reduces computational cost while maintaining reliable classification results. It can be processed quickly on GPUs and even CPUs, making it a promising choice for embedded systems.
[0071] Step S4: Input the extracted deep semantic information into the trained classifier model, classify the deep semantic feature map of the hyperspectral remote sensing image in the training set, and obtain the ground object classification map of the hyperspectral remote sensing image.
[0072] The present invention classifies objects in hyperspectral remote sensing images by constructing a spatial spectral attention feature extraction network, wherein the spatial spectral attention feature extraction network includes a lightweight spatial feature extraction convolution layer and a spectral attention mechanism module; the lightweight spatial feature extraction convolution layer uses partial convolution to effectively extract spatial features, and uses partial convolution operations to selectively operate on some input channels, thereby reducing redundant calculations in spatial feature extraction. The spectral attention mechanism module captures the nonlinear information between channels in the feature map from the feature map that has undergone spatial feature extraction to the maximum extent possible. The captured nonlinear information provides the classifier with a more expressive and discriminative feature representation, thereby enhancing the classifier's ability to learn complex data distributions and abstract features, and improving the accuracy and generalization performance of classification. Inputting the effective spectral feature map into the feature extraction network allows the network to focus on more important spectral information, thereby improving the accuracy of object classification.
[0073] The captured nonlinear information provides the classifier with more expressive and discriminative feature representation, thereby enhancing the classifier's ability to learn complex data distributions and abstract features, and improving classification accuracy and generalization performance.
[0074] Figure 4 A schematic diagram of an architecture for spatial feature extraction is given. To design fast neural networks, much work has focused on reducing the number of floating-point operations (FLOPs). However, this reduction in FLOPs does not necessarily lead to a similar reduction in latency. This is mainly due to inefficient floating-point operations per second (FLOPS). The low FLOPS is mainly due to frequent memory accesses during convolution operations, especially depthwise convolutions. Therefore, we use a new PConv to more efficiently extract spatial features by reducing redundant computations and simultaneous memory accesses.
[0075] After preprocessing and normalization, the hyperspectral remote sensing data F m, input it into the spatial feature extraction module, and use PConv to perform conventional convolution on only some input channels to extract spatial features, while the remaining channels are not processed. In order to facilitate sequential or conventional memory access, the first or last consecutive c p Computed as a proxy for the complete feature map.
[0076]
[0077] Among them, h and w correspond to the height and width of the feature map, k represents the size of the convolution kernel, and c p Indicates the number of channels selected when implementing PConv. By using a regular partial ratio of 1 / 4 (r = c p / c, where c represents the number of channels of the original convolution), the FLOPs of PConv is reduced to only 1 / 16 of that of regular convolution. In addition, PConv requires less memory access.
[0078]
[0079] When r=1 / 4, the memory access is only 1 / 4 of that of regular convolution.
[0080] PConv will only be used for c p The spatial features are extracted from the channel and the remaining channels (cc p ) are retained because they are useful for the subsequent Conv1×1 layers. This approach allows feature information to propagate through all channels, keeping the design simple and lightweight, and making the overall architecture hardware-friendly. To comprehensively combine the extracted features, further feature processing is required to complete the classification task. The PConv layer is followed by two Conv1×1 layers, forming a scheme similar to an inverted residual block. Here, the intermediate layers have more channels.
[0081] Figure 5 A schematic diagram of the channel attention mechanism module is presented. When each channel of the feature map is considered a feature detector, these channels focus on what is meaningful in the input image. To efficiently compute channel attention, the channel attention mechanism module first compresses the spatial dimensions of the input feature map. Experiments demonstrate that both average pooling and max pooling effectively improve the network's representational capabilities. Unlike the convolutional block attention module, to effectively combine two different features, a 2×1 convolutional layer is used to combine the two pooled features. Dimensionality reduction and enhancement modules are then used to obtain per-channel attention.
[0082] Spectral attention feature extraction first aggregates the spatial information of the feature map using global average pooling and global maximum pooling operations to generate two different spatial context descriptors and Denote global average pooling and global maximum pooling respectively. The two descriptors are combined through 2×1 convolution.
[0083] To improve the generalization ability, a multi-layer perceptron (MLP) is added to learn the final channel attention map M. To reduce the parameter overhead, the size of the hidden layer is set to C / γ, where γ is the compression rate.
[0084] The enhanced channel attention calculation is expressed as:
[0085]
[0086] Where W0∈R (C / γ)×C ,W1∈R C ,W0 and W1 are the weights of the multilayer perceptron, f 2×1 Indicates that the size of the convolutional layer filter is 2×1.
[0087] The present invention also provides a ground object classification system for UAV hyperspectral remote sensing images, comprising:
[0088] The data acquisition module is used to obtain remote sensing images of the object to be measured and classified, and pre-process the obtained remote sensing images to obtain processed hyperspectral remote sensing image data;
[0089] The model building module is used to build a spatial-spectral attention feature extraction network model, using a portion of the processed hyperspectral remote sensing image data as a training set. The hyperspectral remote sensing images in the training set are input into the model, and the deep semantic feature map is extracted through the spatial-spectral attention mechanism. Based on the deep semantic feature map, the deep semantic information of the hyperspectral remote sensing image is obtained.
[0090] The data output module is used to classify the deep semantic feature maps of the hyperspectral remote sensing images in the training set according to the deep semantic information of the hyperspectral remote sensing images obtained, and output the ground object classification results of the hyperspectral remote sensing images.
[0091] The spatial-spectral attention feature extraction network consists of a lightweight spatial feature extraction convolutional layer and a spectral attention mechanism module. The lightweight spatial feature extraction convolutional layer uses partial convolution to effectively extract spatial features. The spectral attention mechanism module captures the nonlinear information between channels in the feature map. The effective spectral feature map is input into the feature extraction network, allowing the network to focus on more important spectral information.
[0092] To extract spatial features, partial convolution (PConv) is used to perform regular convolution on only some input channels to extract spatial features, while the remaining channels are not processed. In order to facilitate sequential or regular memory access, the first or last consecutive c pIt is calculated as a substitute for the complete feature map and can extract the deep semantic information of hyperspectral spatial data using fewer parameters.
[0093] Experimental analysis
[0094] In order to highlight the advantages of the present invention, the present invention uses three methods for comparative experiments, namely DFFN, DHCNet and SSRN.
[0095] Table 1 Comparison of parameters and calculation amount of different classification methods
[0096]
[0097] Table 2 Comparison of Pavia University data classification accuracy (%)
[0098]
[0099] As shown in Table 1, when the input is a 7*7 pixel block, the method proposed in the present invention is compared with the other three methods in terms of the number of parameters and the amount of calculation.
[0100] Table 2 shows the comparison of classification accuracy and inference time on GPU and CPU using Pavia University data.
[0101] From the analysis of the results in the table, it can be seen that compared with DFFN, DHCNet, and SSRN, the method proposed in this invention has the least number of parameters and computational complexity while maintaining the highest classification accuracy and the shortest inference time.
[0102] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
[0103] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.
Claims
1. A method for classifying objects using hyperspectral remote sensing by unmanned aerial vehicles, characterized in that: The following steps are involved: Acquire a UAV hyperspectral remote sensing image, preprocess the acquired UAV hyperspectral remote sensing image, and obtain a processed UAV hyperspectral remote sensing image; Extract deep semantic feature maps from processed UAV hyperspectral remote sensing images; Construct a spatial-spectral attention feature extraction network model, input the extracted deep semantic feature map into the spatial-spectral attention feature extraction network model, and obtain the deep semantic information of the UAV hyperspectral remote sensing image; Based on the deep semantic information of the obtained UAV hyperspectral remote sensing image, the deep semantic feature map of the UAV hyperspectral remote sensing image to be tested is classified, and the ground object classification result of the UAV hyperspectral remote sensing image is output; The spatial-spectral attention feature extraction network model includes a lightweight spatial feature extraction convolutional layer and a spectral attention mechanism module. The lightweight spatial feature extraction convolutional layer uses partial convolution to extract spatial features to obtain a feature map of deep spatial semantic information of the UAV hyperspectral remote sensing image. The spectral attention mechanism module captures the nonlinear information between channels in the feature map of deep spatial semantic information. The method of obtaining a feature map of deep spatial semantic information of a UAV hyperspectral remote sensing image comprises the following steps: The deep semantic feature map is input into the lightweight spatial feature extraction convolution layer, and partial convolution is used to perform regular convolution on only part of the input channels to extract spatial features, while the remaining channels are not processed; The first or last consecutive c p The channel is calculated as a substitute for the complete feature map, and the floating-point operation number of the network is: Among them, h and w correspond to the height and width of the feature map, k represents the size of the convolution kernel, and c p Indicates the number of channels selected when implementing partial convolution; Set the partial ratio r = c p / c, where c p represents the number of channels of partial convolution, and c represents the number of channels of traditional convolution; Partial convolution requires fewer memory accesses, which is: Partial convolution is only performed on c p Extract spatial features from the channel and convert the remaining channels cc p reserve; The spectral attention mechanism module captures the nonlinear information between channels in the feature map of deep spatial semantic information, including the following steps: Global average pooling and global maximum pooling operations are used to aggregate the spatial information of the feature map to generate two different spatial context descriptors. and Represent global average pooling and global maximum pooling respectively; The two descriptors are combined through a 2×1 convolution; A multi-layer perceptron is added to learn the final channel attention map M; the size of the hidden layer is set to c / γ, where γ is the compression rate; Where W0∈R (c / γ)×c ,W1∈R c ,W0 and W1 are the weights of the multilayer perceptron, f 2×1 Indicates that the size of the convolutional layer filter is 2×1.
2. The method for classifying objects using hyperspectral remote sensing by an unmanned aerial vehicle according to claim 1, wherein: The preprocessing of the obtained UAV hyperspectral remote sensing image includes the following steps: Determine the attribute information and size information of UAV hyperspectral remote sensing images; According to the attribute information and size information of the UAV hyperspectral remote sensing images, the UAV hyperspectral remote sensing images are unified into the MAT format, and the images are cropped and set to the ratio consistent with the training data.
3. The method for classifying objects using hyperspectral remote sensing by an unmanned aerial vehicle according to claim 2, wherein: The extraction of the deep semantic feature map of the processed UAV hyperspectral remote sensing image includes the following steps: A portion of the processed UAV hyperspectral remote sensing image data is used as a training set; Normalize the UAV hyperspectral remote sensing image data in the training set; Extract the deep semantic feature map of the UAV hyperspectral remote sensing image data in the normalized training set.
4. The method for classifying objects using hyperspectral remote sensing by an unmanned aerial vehicle according to claim 1, wherein: The method of classifying the deep semantic feature map of the UAV hyperspectral remote sensing image to be tested includes inputting the extracted deep semantic feature map into the trained classifier model, classifying the deep semantic feature map of the UAV hyperspectral remote sensing image in the training set, and obtaining a classification result.
5. The method for classifying objects using hyperspectral remote sensing by an unmanned aerial vehicle according to claim 4, wherein: It also includes calculating the value of the loss function based on the classification results, updating the parameters of the spatial spectral attention feature extraction network model through back propagation, and completing the training of the network model.
6. A UAV hyperspectral remote sensing object classification system, characterized by: include: A data acquisition module is used to acquire UAV hyperspectral remote sensing images, preprocess the acquired UAV hyperspectral remote sensing images, and obtain processed UAV hyperspectral remote sensing images; The model building module is used to extract the deep semantic feature map of the processed UAV hyperspectral remote sensing image; build a spatial spectral attention feature extraction network model, input the extracted deep semantic feature map into the spatial spectral attention feature extraction network model, and obtain the deep semantic information of the UAV hyperspectral remote sensing image; The spatial-spectral attention feature extraction network model includes a lightweight spatial feature extraction convolutional layer and a spectral attention mechanism module. The lightweight spatial feature extraction convolutional layer uses partial convolution to extract spatial features to obtain a feature map of deep spatial semantic information of the UAV hyperspectral remote sensing image. The spectral attention mechanism module captures the nonlinear information between channels in the feature map of deep spatial semantic information. The method of obtaining a feature map of deep spatial semantic information of a UAV hyperspectral remote sensing image comprises the following steps: The deep semantic feature map is input into the lightweight spatial feature extraction convolution layer, and partial convolution is used to perform regular convolution on only part of the input channels to extract spatial features, while the remaining channels are not processed; The first or last consecutive c p The channel is calculated as a substitute for the complete feature map, and the floating-point operation number of the network is: Among them, h and w correspond to the height and width of the feature map, k represents the size of the convolution kernel, and c p Indicates the number of channels selected when implementing partial convolution; Set the partial ratio r = c p / c, where c p represents the number of channels of partial convolution, and c represents the number of channels of traditional convolution; Partial convolution requires fewer memory accesses, which is: Partial convolution is only performed on c p Extract spatial features from the channel and convert the remaining channels cc p reserve; The spectral attention mechanism module captures the nonlinear information between channels in the feature map of deep spatial semantic information, including the following steps: Global average pooling and global maximum pooling operations are used to aggregate the spatial information of the feature map to generate two different spatial context descriptors. and Represent global average pooling and global maximum pooling respectively; The two descriptors are combined through a 2×1 convolution; A multi-layer perceptron is added to learn the final channel attention map M; the size of the hidden layer is set to c / γ, where γ is the compression rate; Where W0∈R (c / γ)×c ,W1∈R c ,W0 and W1 are the weights of the multilayer perceptron, f 2×1 Indicates that the size of the convolutional layer filter is 2×1; The data output module is used to classify the deep semantic feature map of the UAV hyperspectral remote sensing image to be tested according to the deep semantic information of the UAV hyperspectral remote sensing image obtained, and output the ground object classification result of the UAV hyperspectral remote sensing image.