Hyperspectral image and LiDAR data collaborative classification system based on combination of frequency domain feature learning and CNN

Through the method of frequency domain feature learning combined with CNN, multi-scale decomposition and feature fusion of hyperspectral images and lidar data is solved, and the problems of insufficient spectral-spatial feature coupling and loss of high-frequency detail information in remote sensing classification are achieved, and efficient multimodal data collaborative classification is achieved.

CN120279330APending Publication Date: 2025-07-08HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510431731.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the existing multimodal remote sensing classification method, insufficient spectral-space feature coupling, low cross-modal fusion efficiency and loss of high-frequency details are problems, especially in the coordinated classification scenarios of hyperspectral images and lidar data.

Method used

The method of frequency domain feature learning combined with convolutional neural network is adopted to perform multi-scale decomposition of hyperspectral images and lidar data. The spectral-spatial features are extracted through a hybrid convolutional architecture, and a dynamic channel weighting mechanism is designed to fuse multi-scale elevation features, and a frequency domain-space dual-current feature interaction module is built to realize cross-modal feature alignment and adaptive fusion.

Benefits of technology

It significantly improves the accuracy of remote sensing classification, solves the problems of insufficient spectral-space feature coupling and loss of high-frequency details, and realizes efficient coordinated classification of HSI and LiDAR data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005348391290000071
    Figure BDA0005348391290000071
  • Figure BDA0005348391290000091
    Figure BDA0005348391290000091
  • Figure BDA0005348391290000092
    Figure BDA0005348391290000092
Patent Text Reader

Abstract

The invention discloses a frequency domain feature learning and CNN combined hyperspectral image and LiDAR data collaborative classification system, and belongs to the technical field of remote sensing image classification. According to the method, a spectrum-space feature coupling mechanism in an existing multi-modal remote sensing classification method is optimized, the cross-modal fusion efficiency is improved, and the obtaining capability of high-frequency detail information is enhanced. According to the method, low-frequency spectral features and high-frequency spatial details of hyperspectral data are decomposed through wavelet transform, and hierarchical feature extraction is realized by adopting a hybrid convolution architecture; meanwhile, a dynamic channel weighting mechanism is designed to fuse the multi-scale elevation features of the LiDAR, and the detail extraction precision of the topographic features is enhanced through a residual attention mechanism; and finally, constructing a frequency domain-space domain double-flow feature interaction module, and realizing cross-modal complementary feature alignment and adaptive fusion. Through a series of innovative technical means, the classification performance of the multi-modal remote sensing data is remarkably improved, and an efficient, accurate and high-robustness solution is provided for the technical field of multi-source remote sensing image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image classification, and particularly relates to a collaborative classification system for hyperspectral images and lidar data that combines frequency-domain feature learning with a convolutional neural network. Background Art

[0002] In the current field of remote sensing data processing, the performance of deep learning largely depends on the scale of the dataset. However, obtaining sufficient and high-quality remote sensing data to support model training often faces significant challenges. These challenges include difficulties in fusing heterogeneous data and significant differences between different data sources, especially in the collaborative classification scenario of hyperspectral images (HSIs) and lidar (light detection and ranging, LiDAR) data.

[0003] HSIs can accurately reflect the spectral characteristics of ground objects with their rich spectral resolution and detailed spectral properties, but they also have a spectral redundancy problem due to their high-dimensional characteristics. LiDAR data effectively complements the deficiencies of HSIs by providing rich elevation information. However, due to the significant differences in data types, dimensions, and characteristics between the two types of data, the collaborative utilization of multimodal data has become an important research topic. Traditional feature extraction and classification methods often rely on a single data type or simple feature fusion techniques and cannot fully utilize the complementary characteristics of HSIs and LiDAR data.

[0004] To this end, the present invention proposes a collaborative classification system for HSIs and LiDAR data that combines frequency-domain feature learning with a convolutional neural network (CNN). The system performs multi-scale decomposition on HSIs and LiDAR data through discrete wavelet transform (DWT), and uses a hybrid convolutional architecture to hierarchically extract spectral-spatial features of HSIs; at the same time, a dynamic channel weighting mechanism is designed to fuse the multi-scale elevation features of LiDAR data, and residual attention is used to enhance the retention of terrain details; finally, a frequency-domain-spatial-domain two-stream feature interaction module is constructed to achieve cross-modal complementary feature alignment and adaptive fusion.

[0005] The system effectively solves the problems of insufficient coupling of spectral-spatial features, low cross-modal fusion efficiency, and loss of high-frequency detail information in existing multimodal remote sensing classification methods, significantly improves the classification accuracy, and provides a new solution for the collaborative classification of HSIs and LiDAR data. Summary of the Invention

[0006] The object of the present invention is to solve the problems of insufficient coupling of spectral-spatial features, low cross-modal fusion efficiency, and loss of high-frequency detail information in existing multi-modal remote sensing classification methods, and to propose a collaborative classification system for HSI and LiDAR data that combines frequency-domain feature learning with CNN.

[0007] The technical solution adopted by the present invention to solve the above technical problems is as follows: a collaborative classification system for HSI and LiDAR data that combines frequency-domain feature learning with CNN, and the method specifically includes the following steps:

[0008] Step 1: Obtain Light Detection and Ranging-Digital Surface Model (LiDAR-DSM) image data and HSI data from the dataset.

[0009] Step 2: Slice the LiDAR-DSM image data and the hyperspectral image data after dimensionality reduction respectively, and then divide the sliced LiDAR-DSM image data and hyperspectral image data into two parts: a training set and a test set.

[0010] Step 3: Construct a collaborative classification system that combines frequency-domain feature learning with CNN, and use the training set to train the constructed fusion classification network until the maximum number of training times set is reached or the classification accuracy of the classification network on the test set no longer improves, and then stop training to obtain a trained collaborative classification system.

[0011] The collaborative classification system that combines frequency-domain feature learning with CNN includes an HSI data processing branch and a LiDAR-DSM data processing branch. The working process of the collaborative classification system that combines frequency-domain feature learning with CNN is as follows:

[0012] Take the HSI data as the input of the HSI data processing branch. In the HSI data processing branch, the input HSI first undergoes three-dimensional discrete wavelet transform (3D-Discrete Wavelet Transform, 3D-DWT).

[0013] Take the high-frequency output after 3D-DWT as the input of the first convolutional layer with a convolutional kernel size of 1×1×1.

[0014] Take the low-frequency output after 3D-DWT as the input of the Spectral-Spatial DualFusion GraphConv Module (SSDFGM).

[0015] Take the HSI data as the input of the HSI data processing branch. Inside the HSI data processing branch, the input HSI first undergoes a two-dimensional discrete wavelet transform (2D-Discrete Wavelet Transform, 2D-DWT).

[0016] Take the high-frequency output after 2D-DWT as the input of a convolutional layer with a second convolutional kernel size of 1×1.

[0017] Take the low-frequency output after 2D-DWT as the input of a third convolutional layer with a convolutional kernel size of 3×3; then take the output of the third convolutional layer as the input of a fourth convolutional layer with a convolutional kernel size of 3×3, and then take the output of the fourth convolutional layer as the input of the spatial attention mechanism.

[0018] Take the LiDAR-DSM image data as the input of the LiDAR-DSM data processing branch. Inside the LiDAR-DSM data processing branch, the input LiDAR-DSM image first undergoes 2D-DWT.

[0019] Take the high-frequency output after 2D-DWT as the input of a fifth convolutional layer with a convolutional kernel size of 1×1, and then take the output of the fifth convolutional layer as the input of the channel attention mechanism.

[0020] Take the low-frequency output after 2D-DWT as the input of an AsymResConv Block (ARCB), and then take the output of the AsymResConv Block as the input of a Dynamic Multi-Scale Fusion Module (DMSF).

[0021] Take the output after ARCB and the output after DMSF through a residual connection as the input of a Feature Enhancement Module (FEM).

[0022] Merge the output after the first convolutional layer and the output after the SSDFGM module into output A.

[0023] Merge the output after the second convolutional layer and the output after the spatial attention mechanism into output B.

[0024] Merge the output after the channel attention mechanism and the output after the feature enhancement module into output C.

[0025] Weightedly sum output A and output B to get output D, and weightedly sum output C and output D to get output E.

[0026] After flattening the output E, the output F is obtained through a position encoder, a frequency encoder, and a frequency domain attention-based encoder (FAE).

[0027] The output F is passed through a classifier to obtain the classification result of the trained model.

[0028] Step 4: Use the trained data collaborative classification system to jointly process the HSI and LiDAR-DSM images of the area to be classified to obtain the classification result.

[0029] The beneficial effects of the present invention are as follows:

[0030] The present invention can decompose the spectral-spatial joint frequency domain features by implementing 3D discrete wavelet transform on the HSI branch, enhance the spatial-spectral correlation modeling module using the spatial-spectral dual-fusion graph convolutional module, and simultaneously implement 2D discrete wavelet transform to extract spatial high-frequency details. The local texture features are captured through cascaded convolution and spatial attention mechanism. The LiDAR-DSM data is processed by an asymmetric residual convolution block and a dynamic multi-scale fusion module to capture terrain elevation details and multi-scale geometric features, construct a complementary feature expression system, break through the limitations of single-modal features, and is optimized by a frequency domain attention encoder, significantly improving the classification accuracy. Brief Description of the Drawings

[0031] Figure 1 is a training flowchart of a collaborative classification system for HSI and LiDAR data combining frequency domain feature learning and CNN;

[0032] Figure 2 is the SSDFGM structure diagram;

[0033] Figure 3 is the ARCB structure diagram;

[0034] Figure 4 is the DMSF structure diagram;

[0035] Figure 5 is the spatial distribution and ground object category information of the Trento dataset;

[0036] Figure 6 is the spatial distribution and ground object category information of the Houston 2013 dataset;

[0037] Figure 7 is the spatial distribution and ground object category information of the Augsburg dataset;

[0038] Figure 8 is the Trento dataset;

[0039] In the figure, (a) is a hyperspectral false color image, (b) is a DSM grayscale image, and (c) is a true value image;

[0040] Figure 9 is the Houston 2013 dataset;

[0041] In the figure, (a) is a hyperspectral false color image, (b) is a DSM grayscale image, and (c) is a true value image;

[0042] Figure 10 is the Augsburg dataset;

[0043] In the figure, (a) is a hyperspectral false color image, (b) is a DSM grayscale image, and (c) is a true value image;

[0044] Figure 11 is the subjective classification result of the fusion classification network based on feature confidence for different datasets;

[0045] In the figure, (a) is Trento, (b) is Houston 2013, and (c) is Augsburg. Detailed implementation manners

[0046] Detailed implementation manner 1: Combined with Figure 1 to illustrate this implementation manner. A hyperspectral and lidar data fusion classification method based on feature confidence described in this implementation manner specifically includes the following steps:

[0047] Step 1: Obtain LiDAR-DSM image data and HSI data from the dataset;

[0048] For the HSI obtained from the dataset, LiDAR-DSM image data corresponding to the HSI region was obtained simultaneously;

[0049] Step 2: Slice the LiDAR-DSM image data and the hyperspectral image data after dimensionality reduction respectively, and then divide the sliced LiDAR-DSM image data and hyperspectral image data into two parts: a training set and a test set;

[0050] When dividing, the LiDAR-DSM image data corresponding to the same region and the hyperspectral image data after dimensionality reduction need to be divided into the training set or the test set simultaneously, so that when inputting into the model later, the LiDAR-DSM image data corresponding to the same region and the hyperspectral image data after dimensionality reduction can be input into the model simultaneously;

[0051] Step 3: Construct a collaborative classification system that combines frequency-domain feature learning with CNN. Use the training set to train the constructed fusion classification network until the maximum number of training times is reached or the classification accuracy of the classification network on the test set no longer improves, and then stop training to obtain a trained collaborative classification system;

[0052] The collaborative classification system that combines frequency-domain feature learning with CNN includes an HSI data processing branch and a LiDAR-DSM data processing branch. The working process of the collaborative classification system that combines frequency-domain feature learning with CNN is as follows:

[0053] Take the HSI data as the input of the HSI data processing branch. In the HSI data processing branch, the input HSI first undergoes 3D-DWT;

[0054] Take the high-frequency output after 3D-DWT as the input of the first convolutional layer with a convolutional kernel size of 1×1×1;

[0055] Take the low-frequency output after 3D-DWT as the input of SSDFGM;

[0056] Take the HSI data as the input of the HSI data processing branch. In the HSI data processing branch, the input HSI first undergoes 2D-DWT;

[0057] Take the high-frequency output after 2D-DWT as the input of the second convolutional layer with a convolutional kernel size of 1×1;

[0058] Take the low-frequency output after 2D-DWT as the input of the third convolutional layer with a convolutional kernel size of 3×3; then take the output of the third convolutional layer as the input of the fourth convolutional layer with a convolutional kernel size of 3×3, and then take the output of the fourth convolutional layer as the input of the spatial attention mechanism;

[0059] Take the LiDAR-DSM image data as the input of the LiDAR-DSM data processing branch. In the LiDAR-DSM data processing branch, the input LiDAR-DSM image first undergoes 2D-DWT;

[0060] Take the high-frequency output after 2D-DWT as the input of the fifth convolutional layer with a convolutional kernel size of 1×1, and then take the output of the fifth convolutional layer as the input of the channel attention mechanism;

[0061] Take the low-frequency output after 2D-DWT as the input of ARCB, and then take the output of the asymmetric residual convolutional block as the input of the DMSF module;

[0062] Take the output after the ARCB module and the output after DMSF as the input of FEM through residual connection;

[0063] Combine the output of the first convolutional layer and the output of the SSDFGM to obtain Output A;

[0064] Combine the output of the second convolutional layer and the output of the spatial attention mechanism to obtain Output B;

[0065] Combine the output of the channel attention mechanism and the output of the feature enhancement module to obtain Output C;

[0066] Weightedly sum Output A and Output B to obtain Output D, and weightedly sum Output C and Output D to obtain Output E;

[0067] After flattening Output E, pass it through a position encoder, a frequency encoder, and the FAE to obtain Output F;

[0068] Pass Output F through a classifier to obtain the classification result of the trained model;

[0069] Step 4: Use the trained data collaborative classification system to jointly process the HSI and LiDAR-DSM images of the area to be classified, and obtain the classification result.

[0070] Before inputting into the classification network, the HSI and LiDAR-DSM images of the area to be classified need to be preprocessed, that is, the HSI needs to be dimension-reduced and sliced, and the LiDAR-DSM image needs to be sliced.

[0071] Specific Embodiment 2: Combine Figure 2 Describe this embodiment. The difference between this embodiment and Specific Embodiment 1 is that the working process of the hyperspectral-spatial dual-fusion graph convolution module is as follows:

[0072] Use the low-frequency features of the HSI after 3D-DWT as the input of the sixth convolutional layer with a convolutional kernel size of 3×3×3, and then use the output of the sixth convolutional layer as the input of the multi-scale two-dimensional convolution with convolutional kernel sizes of 1×1, 3×3, and 5×5 respectively;

[0073] Use the output of the multi-scale two-dimensional convolution as the input of the dynamic attention mechanism;

[0074] Use the low-frequency features of the HSI after 3D-DWT as the input of the GCN;

[0075] Combine the output of the dynamic attention mechanism and the output of the GCN to obtain Output a;

[0076] Other steps and parameters are the same as those in Specific Embodiment 1.

[0077] The hyperspectral-spatial dual-fusion graph convolution module of the present invention fuses the features of multi-modal data through 3D-2D hybrid convolution, and can effectively improve the classification accuracy by connecting the GCN in parallel.

[0078] Embodiment 3: Figure 3 This embodiment will be described. The difference between this embodiment and Embodiment 1 or 2 is that the working process of the asymmetric residual convolution block is as follows:

[0079] Use the low-frequency features of the LiDAR-DSM image after 2D-DWT as the input of the seventh convolutional layer with a kernel size of 1×1;

[0080] Use the output of the seventh convolutional layer as the input of the horizontal convolution with a kernel size of 3×1, and then use the output of the horizontal convolution as the input of the vertical convolution with a kernel of 1×3;

[0081] Use the output of the vertical convolution as the input of the eighth convolutional layer with a kernel size of 3×3;

[0082] Merge the output of the seventh convolutional layer and the output of the eighth convolutional layer through residual connection to obtain output b;

[0083] Other steps and parameters are the same as those in Embodiment 1 or 2.

[0084] Through direction decomposition and residual enhancement, the asymmetric residual convolution block achieves a balance between efficient calculation and strong representation ability.

[0085] Embodiment 4: Figure 4 This embodiment will be described. The difference between this embodiment and any one of Embodiments 1 to 3 is that the working process of the dynamic multi-scale fusion module is as follows:

[0086] Use output b as the input and process it through parallel dilated convolutions. Expand the receptive field by inserting spaces between the kernel elements. The formula is

[0087] s=(k - 1)×d + 1

[0088] where s is the equivalent kernel size, k is the original kernel size, and d is the dilation rate;

[0089] Sum the outputs of the convolutional layers with receptive fields of 1×1, 3×3, and 5×5 respectively to obtain output c;

[0090] Other steps and parameters are the same as those in any one of Embodiments 1 to 3.

[0091] In the present invention, through low-cost dilated convolutions and learnable weight assignment, the dynamic multi-scale fusion module achieves efficient cooperation of multi-scale features, while improving the model accuracy and maintaining the computational efficiency.

[0092] Embodiment 5: The difference between this embodiment and any one of Embodiments 1 to 4 is that the working process of the feature enhancement module is as follows:

[0093] Take the output c as the input of the ninth convolutional layer with a convolutional kernel size of 3×3, and then take the output of the ninth convolutional layer as the input of the depth convolution with a convolutional kernel size of 3×3;

[0094] Take the output after depth convolution as the input of the pointwise convolution with a convolutional kernel size of 1×1, and then take the output of the pointwise convolution as the input of the cross-attention mechanism;

[0095] Take the output of the cross-attention as the input of the large receptive field spatial modulator, and merge the output of the large receptive field spatial modulator with the output after passing through the channel attention mechanism as the output C;

[0096] Other steps and parameters are the same as those in any one of the specific embodiments one to four.

[0097] In the present invention, the feature enhancement module realizes the multi-dimensional adaptive enhancement of terrain-related features through the cascaded design of channel expansion and elevation attention, while improving the model accuracy and maintaining the computational efficiency.

[0098] Embodiment

[0099] This embodiment proposes a collaborative classification system for HSI and LiDAR data that combines frequency-domain feature learning with CNN. The method implementation process is shown in Table 1:

[0100] Table 1 Algorithm process of the fusion classification network structure based on feature confidence

[0101]

[0102]

[0103] The specific implementation steps are as follows:

[0104] Step 1: Obtain HSI data and LiDAR-DSM data (using publicly available data).

[0105] Step 2: Preprocess the LiDAR-DSM data and HSI data. According to the number of labeled pixel samples in the HSI and LiDAR-DSM images, specify the number of training sets and test sets, and then perform slicing processing on the LiDAR-DSM data and HSI data respectively.

[0106] Step 3: Training and classification of the data collaborative classification system that combines frequency-domain feature learning with CNN

[0107] Step 3.1: Construct a data collaborative classification system that combines frequency-domain feature learning with CNN. The training set and test set images used in the present invention both adopt a resolution of 9×9 pixels. As Figure 1As shown, H×W represents the spatial dimension, and B represents the third-dimensional band, obtaining a series of hyperspectral images with a size of 9×9×C and LiDAR-DSM data with a size of 9×9.

[0108] Step 3.2: The feature extraction structure of the data collaborative classification system for frequency-domain feature learning is mainly divided into three parts, namely 3D-DWT and 2D-DWT for HSI data, and 2D-DWT for LiDAR-DSM data.

[0109] Step 3.3: The high-frequency features extracted from the HSI data branch of 3D-DWT enter 3D-CNN, and the low-frequency features enter the SSDFGM module. The output after 3D-CNN and the output after the SSDFGM module are merged.

[0110] Step 3.4: The high-frequency features extracted from the HSI data branch of 2D-DWT enter 2D-CNN. The low-frequency features go through cascaded convolutions, and the output after cascaded convolutions is used as the input of the spatial attention mechanism. The output after 2D-CNN and the output after the spatial attention mechanism are merged.

[0111] Step 3.5: The output of Step 3.3 and the output of Step 3.4 are merged and output as the extracted HSI features.

[0112] Step 3.6: The high-frequency features extracted from the LiDAR-DSM data branch of 2D-DWT enter 2D-CNN, and the output after 2D-CNN is used as the input of the channel attention mechanism. The low-frequency features go through the ARCB module and the DMSF module in sequence, and the output after ARCB and the output after DMSF are connected in a residual manner and output. The output goes through the FEM module, and the output after the channel attention mechanism and the output after the FEM module are merged and output as the extracted LiDAR-DSM features.

[0113] Step 3.7: The features obtained in Step 3.5 and Step 3.6 are merged. After flattening the output, it is output through the position encoder, frequency encoder, and FAE module, and the output goes through the classifier to obtain the classification result of the training model.

[0114] Step 4: Analysis of classification results.

[0115] The network model training and classification result verification experiments of the present invention are completed on the following platform:

[0116] The hardware configuration includes an Intel i7-13700H processor, 32GB of memory, an NVIDIA GeForce RTX 4060 graphics card, a storage system consisting of a 1TB solid-state drive (SSD), and the operating system is Windows 11 Pro. The experiment used three well-known multi-modal remote sensing datasets, specifically including the Trento dataset, the dataset of the University of Houston and its surrounding areas in Texas, USA (Houston 2013), and the Augsburg dataset. The spatial distribution and ground object category information of each dataset are respectively as Figure 5 , Figure 6 and Figure 7 shown, where different colors represent different ground object categories. The detailed statistical information and feature descriptions of each dataset are respectively as Figure 8 , Figure 9 and Figure 10 shown.

[0117] 1. The capture location of the Trento dataset is in the rural area around the city of Trento in Italy. The dataset contains hyperspectral images and LiDAR images. Among them, the hyperspectral image has a size of 600×166 pixels, contains 63 bands, the band range covers the spectral band from 420.89 to 989.09 nanometers, the spectral resolution is 9.2 nanometers, and the spatial resolution is 1 meter. The LiDAR image is a single-channel image, containing the altitude corresponding to the ground position, and the image size is the same as that of the hyperspectral image. There are a total of six ground object categories in the annotation information of the dataset.

[0118] 2. The Houston 2013 dataset was specially provided by the IEEE Geoscience and Remote Sensing Society (IEEE GRSS) for the 2013 data fusion competition, covering the University of Houston and its surrounding areas in Texas, USA. The dataset contains 15 ground object categories, with a spatial resolution of 2.5 meters per pixel, including 340×1905 pixels of HSI and LiDAR-DSM data. Among them, the HSI has 144 bands, and the wavelength range is from 0.38 to 1.05 micrometers.

[0119] 3. The Augsburg dataset was captured over the city of Augsburg in Germany. The HSI data was obtained by the DAS-EOC HySpex sensor, and the LiDAR-DSM data was collected by the DLR-3K system. The spatial resolution of the two images was uniformly downsampled to 30 meters to fully manage multi-modal fusion. In this dataset, the HSI data consists of 180 bands, with a range of 0.4 to 2.5 micrometers, while the LiDAR-DSM data has only one raster. The size of this dataset is 332×485 pixels. This dataset depicts seven different land cover categories.

[0120] For the evaluation metrics of the model, three of the most widely used objective evaluation metrics in the industry are selected: Overall Accuracy (OA), Average Accuracy (AA), and Kappa Coefficient (K). These three evaluation metrics are all calculated based on the Confusion Matrix.

[0121] The following will introduce these evaluation metrics separately:

[0122] (1) Confusion Matrix: The confusion matrix is an n×n error matrix, which represents the standard form for calculating accuracy, and the detailed content that makes up the matrix is obtained from Table 2. The confusion matrix is a specific matrix used to visualize the performance of an algorithm. Each column of the confusion matrix represents the predicted value, and each row represents the true class. The name of this matrix is because its composition form can clearly show whether there is confusion between multiple classes, that is, whether one class is predicted as another class.

[0123] Table 2 Composition form of the confusion matrix

[0124]

[0125] (2) Overall Classification Accuracy (OA): OA is a basic evaluation metric, which represents the proportion of the number of correctly classified samples to the total number of samples, that is, the overall classification accuracy. The overall classification accuracy represents an indicator of the classification accuracy of an algorithm as a whole, which is the ratio of the number of correctly classified class pixels to the total number of class pixels:

[0126]

[0127] (3) Average Classification Accuracy (AA): AA represents an evaluation metric of classification accuracy in a classification algorithm. First, the classification accuracies of each type of ground object are summed, and then the average value is calculated as the classification accuracy. The average classification accuracy is different from the overall classification accuracy, and it focuses on evaluating the classification accuracy of the algorithm for each type of ground object:

[0128]

[0129] (4) Kappa Coefficient (K): The Kappa coefficient represents the proportion of the reduction in the completely random classification error. It is also an evaluation metric of classification accuracy. As a supplement to the error matrix, it can clearly reflect the effect of classification accuracy:

[0130]

[0131] Among them, n represents the number of rows and columns of the classification matrix, m ijRepresents the value of the i-th row and j-th column in the confusion matrix, m i+ is the row synthesis of the confusion matrix, and m +i is the column synthesis of the confusion matrix, and N represents all elements included in the confusion matrix.

[0132] The objective classification results of the network model used in the present invention on three data sets are shown in Table 3. From left to right, the first column is the data set; the second column is the OA data of the classification result; the third column is the AA data of the classification result; the fourth column is the K×100 data of the classification result; the fifth column is the training duration of the classification; the sixth column is the test duration of the classification. The subjective classification results are as Figure 11 shown, (a) represents the classification effect of the Trento data set, (b) represents the classification effect of the Houston 2013 data set, (c) represents the classification effect of the Augsburg data set, and the higher the accuracy, the less salt-and-pepper noise in the picture.

[0133] Table 3

[0134]

[0135] The above examples of the present invention are only to illustrate in detail the calculation model and calculation process of the present invention, rather than to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made on the basis of the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. A hyperspectral and LiDAR data collaborative classification system that combines frequency-domain feature learning with CNN, characterized in that, The method specifically includes the following steps: Step 1: Obtain Light Detection and Ranging-Digital Surface Model (LiDAR-DSM) image data and Hyperspectral Imagery (HSI) from the dataset; Step 2: Slice the LiDAR-DSM image data and the hyperspectral image data respectively, and then divide the sliced LiDAR-DSM image data and the hyperspectral image data into two parts: a training set and a test set; Step 3: Construct a collaborative classification system that combines frequency-domain feature learning with a Convolutional Neural Network (CNN). Use the training set to train the constructed fusion classification network until the maximum number of training times is reached or the classification accuracy of the classification network on the test set no longer improves, and then stop training to obtain a trained collaborative classification system; The collaborative classification system that combines frequency-domain feature learning with a CNN includes an HSI data processing branch and a LiDAR-DSM data processing branch. The working process of the collaborative classification system that combines frequency-domain feature learning with a CNN is as follows: Take the HSI data as the input of the HSI data processing branch. In the HSI data processing branch, the input HSI first undergoes a 3D-Discrete Wavelet Transform (3D-DWT); Take the high-frequency output after 3D-DWT as the input of a convolutional layer with a kernel size of 1×1×1; Take the low-frequency output after 3D-DWT as the input of the Spectral-Spatial Dual Fusion GraphConv Module (SSDFGM); Take the HSI data as the input of the HSI data processing branch. In the HSI data processing branch, the input HSI first undergoes a 2D-Discrete Wavelet Transform (2D-DWT); Take the high-frequency output after 2D-DWT as the input of a second convolutional layer with a kernel size of 1×1; Take the low-frequency output after 2D-DWT as the input of a third convolutional layer with a kernel size of 3×3; then take the output of the third convolutional layer as the input of a fourth convolutional layer with a kernel size of 3×3, and then take the output of the fourth convolutional layer as the input of the spatial attention mechanism; Take the LiDAR-DSM image data as the input of the LiDAR-DSM data processing branch. In the LiDAR-DSM data processing branch, the input LiDAR-DSM image first undergoes a 2D-DWT; Take the high-frequency output after 2D-DWT as the input of a fifth convolutional layer with a kernel size of 1×1, and then take the output of the fifth convolutional layer as the input of the channel attention mechanism; The low-frequency output after 2D-DWT is used as the input of the AsymResConv Block (ARCB), and the output of the AsymResConv Block is used as the input of the Dynamic Multi-Scale Fusion Module (DMSF). The output after ARCB and the output after DMSF are used as the input of the Feature Enhancement Module (FEM) through residual connection. The output after the first convolutional layer and the output after the SSDFGM module are merged into Output A. The output after the second convolutional layer and the output after the spatial attention mechanism are merged into Output B. The output after the channel attention mechanism and the output after the feature enhancement module are merged into Output C. Output A and Output B are weighted and summed to obtain Output D, and Output C and Output D are weighted and summed to obtain Output E. After Output E is flattened, Output F is obtained through the position encoder, frequency encoder, and Frequency-domain Attention-based Encoder (FAE). Output F passes through the classifier to obtain the classification result of the trained model. Step 4: Use the trained data collaborative classification system to jointly process the HSI and LiDAR-DSM images of the area to be classified to obtain the classification result.

2. The collaborative classification system of HSI and LiDAR data by combining frequency domain feature learning with CNN according to claim 1, wherein, The working process of the spatio-spectral dual-fusion graph convolutional module is as follows: The low-frequency features of the HSI after 3D-DWT are used as the input of the sixth convolutional layer with a convolutional kernel size of 3×3×3, and the output of the sixth convolutional layer is used as the input of the multi-scale two-dimensional convolution with convolutional kernel sizes of 1×1, 3×3, and 5×5 respectively. The output after the multi-scale two-dimensional convolution is used as the input of the dynamic attention mechanism. The low-frequency features of the HSI after 3D-DWT are used as the input of the Graph Convolutional Network (GCN). The output after the dynamic attention mechanism and the output after GCN are merged into Output a.

3. The HSI and LiDAR data collaborative classification system combining frequency domain feature learning with CNN according to claim 2, characterized in that, The working process of the AsymResConv Block is as follows: The low-frequency features of the LiDAR-DSM image after 2D-DWT are used as the input of the seventh convolutional layer with a convolutional kernel size of 1×1. The output of the seventh convolutional layer is used as the input of the horizontal convolution with a convolutional kernel size of 3×1, and the output of the horizontal convolution is used as the input of the vertical convolution with a convolutional kernel of 1×3. The output of the vertical convolution is used as the input of the eighth convolutional layer with a convolutional kernel size of 3×3. The output after the seventh convolutional layer and the output after the eighth convolutional layer are merged through residual connection into Output b.

4. The HSI and LiDAR data collaborative classification system combining frequency domain feature learning with CNN according to claim 3, wherein The working process of the dynamic multi-scale fusion module is as follows: Output b is used as the input and processed through parallel dilated convolution to expand the receptive field by inserting spaces between the convolutional kernel elements. The formula is s=(k - 1)×d + 1 where s is the equivalent kernel size, k is the original kernel size, and d is the dilation rate. The outputs of convolutional layers with receptive fields of 1×1, 3×3, and 5×5 are weighted and summed to obtain the output c.

5. The HSI and LiDAR data collaborative classification system for frequency domain feature learning combined with CNN according to claim 4, characterized in that The working process of the feature enhancement module is as follows: The output c is used as the input of a convolutional layer with a 3×3 convolutional kernel as the ninth layer, and then the output of the ninth convolutional layer is used as the input of a depthwise convolution with a 3×3 convolutional kernel; The output after depthwise convolution is used as the input of a pointwise convolution with a 1×1 convolutional kernel, and then the output of the pointwise convolution is used as the input of the cross-attention mechanism; The output of the cross-attention is used as the input of the large receptive field spatial modulator, and the output of the large receptive field spatial modulator is merged with the output after passing through the channel attention mechanism to obtain the output C.

Citation Information

Cited By

  • HSI and ranging data domain classification method and system based on domain expansion and feature decoupling

    CN121788922A