A collaborative classification method for hyperspectral and LiDAR data based on a dual-branch architecture

Through the dual-branch network and feature pruning technology, combined with hollow convolution and capsule network, the problem of insufficient utilization of spatial information in the coordinated classification of hyperspectral and LiDAR data is solved, and higher classification accuracy is achieved.

CN114429564BActive Publication Date: 2025-07-11HARBIN QIAOSHENG TECH DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210018148.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-08
Publication Date
2025-07-11
Estimated Expiration
2042-01-08

AI Technical Summary

Technical Problem

The traditional hyperspectral and LiDAR data collaborative classification methods have limited performance in feature learning and classification, resulting in insufficient utilization of spatial information, and the large number of features extracted by different sensors are prone to dimensional disasters.

Method used

The hyperspectral and LiDAR data were extracted respectively by a dual-branch network, and the hyperspectral image band was selected using the pruning method, combining hollow convolution and capsule network to extract features, and strengthening the complementary data advantages through softmax classification.

Benefits of technology

The accuracy of land object classification is improved, the problems of insufficient utilization of space information and dimensional disasters in traditional methods are overcome, and the accuracy of land object classification is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114429564B_ABST
    Figure CN114429564B_ABST
Patent Text Reader

Abstract

A collaborative classification method for hyperspectral and LiADR data based on a dual-branch belongs to the technical field of image classification. The method sequentially performs the following steps: input the registered.tif data of hyperspectral and LiDAR, and input the data into a dual-branch network; use a pruning method to select bands for the hyperspectral image; extract features for space and spectrum respectively; use dilated convolution to extract features for the LiDAR branch; splice the features extracted from the hyperspectral image branch and the LiDAR data branch; finally, use softmax to classify the spliced features to obtain sample classification labels. The collaborative classification method for hyperspectral and LiADR data based on a dual-branch of the present invention utilizes the respective characteristics of hyperspectral and LiDAR data, complements each other's advantages, and improves the classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A collaborative classification method for hyperspectral and LiDAR data based on a dual-branch in the present invention belongs to the technical field of image classification. Background Art

[0002] Hyperspectral HSI and Light Detection and Ranging LiDAR remote sensing technologies are important ways to obtain surface information. With the rapid development of earth observation technology and aerospace technology, people can obtain multi-source data sets of the same scene simultaneously, which provides a better platform for the collaborative classification of hyperspectral and LiDAR data. By integrating the individual heterogeneity and data diversity of multi-source remote sensing, improving the accuracy of scene observation and classification has become a new trend in current research. Using collaborative classification can effectively combine and also provide new solutions and effective means for solving the problem of ground object classification. Hyperspectral remote sensing and LiDAR are two common remote sensing means, each with different characteristics. Hyperspectral images can well characterize the spectral information of ground objects and reflect their characteristics such as materials and textures, while LiDAR can efficiently and accurately obtain the elevation data of the ground. Although the above data can describe the characteristics of ground object categories from different angles, the spatial resolution of hyperspectral images is relatively low, and the spectral-spatial information of LiDAR is relatively scarce. Therefore, to achieve and finally implement reliable ground object classification, on the one hand, it is necessary to develop the full-spectrum ground object feature expression of hyperspectral data, accurately extract the ground object spatial information, and efficiently perform feature fusion and classification, so as to improve the accuracy of single-source hyperspectral image information interpretation and classification as much as possible; on the other hand, actively exploring other effective information, improving the classification reliability of hyperspectral data, and constructing multi-source remote sensing collaborative expression are effective means to enhance the earth observation efficiency.

[0003] Traditional collaborative methods first separately perform traditional feature extraction on hyperspectral images and LiDAR images, such as morphological features, wavelet features, texture features, etc., and then use traditional classifiers such as Support Vector Machine (SVM) and Random Forest (RF) to classify the features extracted from the two images. The performance of traditional classification methods in feature learning and classification is limited. Traditional features may lead to insufficient utilization of spatial information. In addition, the more features extracted by different sensors, although the information of the two images can be characterized in more detail, it will simultaneously cause serious dimensionality disasters.

[0004] With the rise of deep learning, a large number of algorithms based on deep learning have been proposed. The convolutional neural network CNN simulates the concept of "local vision" in the human visual system, converts the fully connected layer into local connections, and uses local connections to process spatial dependencies, significantly reducing the number of parameters to be trained and the computational cost. In addition, the convolutional neural network has the ability to learn rich hierarchical representations and autonomous learning, and can adaptively extract appropriate features according to different data sources. Therefore, it is suitable for the fusion classification of hyperspectral images and LiDAR images. Summary of the Invention

[0005] The present invention provides a collaborative classification method for hyperspectral and LiDAR data based on a dual-branch structure, using two branches to extract features from hyperspectral and LiDAR data respectively. First, a pruning method is used to select bands for the hyperspectral image. Then, for the hyperspectral spatial branch, dilated convolution and 2D capsule network are used to extract features, and for the spectral branch, dilated convolution and 1D capsule network are used to extract features. In the LiDAR branch, a cascade module and dilated convolution are used to extract features. Finally, the features of the dual-branches of hyperspectral and LiDAR are concatenated and then softmax classified to strengthen the complementary advantages between data and improve the classification accuracy of ground object classification.

[0006] The object of the present invention is achieved as follows:

[0007] A collaborative classification method for hyperspectral and LiDAR data based on a dual-branch structure, comprising the following steps:

[0008] Step a, input the registered.tif data of hyperspectral and LiDAR into the dual-branch network;

[0009] Step b, use a pruning method to select bands for the hyperspectral image;

[0010] Specifically: for the hyperspectral branch, the entire HSI band is used as input to train the original network parameters; for each band, all parameters in the branch network are integrated to measure the band importance; under the constraint of the new band significance factor, the convolutional neural network is pruned, and some representative weights are retained, and the small sub-network is retrained to finally solve the hyperspectral band selection problem.

[0011] Input the full-band HSI slices and the network structure, which are divided into a training stage and a band selection stage. Training stage: First, randomly initialize the CNN model parameters. Then, train the CNN model with all bands until the accuracy tends to be stable. Band selection stage: First, for each band, calculate the band significance factor. Then calculate the pruning band with the minimum band significance factor and its corresponding kernel matrix. Then retrain the pruned sub-network to restore the accuracy and refine the network by removing the pruned parameters.

[0012] Calculate the band significance factor. Let B represent the number of HSI bands, H / W represent the height / width of the input slice, and N represent the number of output feature maps. The first convolutional layer converts the input HSI slice into a feature map, which is achieved by adopting a 2D kernel of N×B On the HSI slice, it can be expressed as:

[0013]

[0014] where * represents the convolution operation, X j represents the j-th band image, and Y i is the feature map generated by the first layer. All kernels form a kernel matrix Since the bands with smaller kernel weights tend to produce feature maps with weak activation compared to other bands, the band significance factor of the j-th band in HSI is defined as the sum of the absolute values of its corresponding kernel matrix, and the formula is as follows:

[0015]

[0016] The band significance factor reflects the influence of the input frequency band on the output feature map. To achieve band selection, the band with the smallest band effective factor and its corresponding kernel matrix are set to zero and marked as pruning parameters. These pruned parameters will not be updated in the backpropagation step. After pruning the bands and kernels, it is necessary to retrain the network to compensate for the performance degradation. Finally, the network can be refined by removing these pruning parameters to obtain the selected bands and a compact model.

[0017] Step c: Extract features from space and spectrum respectively;

[0018] Specifically:

[0019] Step c1: For the spatial branch of hyperspectral, use a 2D capsule network and dilated convolution to extract features from the spatial information of hyperspectral;

[0020] The spatial branch of hyperspectral contains eight layers, and the structure is as follows: dilated convolution layer 1 → BN layer → Swish activation function → dilated convolution layer 2 → BN layer → Swish activation function → 2D primary capsule layer → digital capsule layer; among them, the convolution kernel size of the first dilated convolution is 3×3, the filter size is 256, and the dilation rate is 2. The convolution kernel size of the second dilated convolution is 3×3, the filter size is 512, and the dilation rate is 3; the capsule network mainly includes a primary capsule layer and a digital capsule layer. The convolution kernel size of the primary capsule layer is 3×3, the stride is 2×2, and the number of routings is set to 3. The number of capsules in the output layer of the digital capsule layer is the same as the number of classifications, the number of capsules is 15, and the capsule dimension is 7;

[0021] Step c2: For the spectral branch of hyperspectral images, use a 1D capsule network and dilated convolution to extract features from the spatial information of hyperspectral images;

[0022] The spectral branch of hyperspectral images consists of eight layers, and the structure is as follows: dilated convolution layer 1 → BN layer → Swish activation function → dilated convolution layer 2 → BN layer → Swish activation function → 1D primary capsule layer → digital capsule layer; among them, the convolution kernel size of the first dilated convolution is 11, the filter size is 64, and the dilation rate is 1. The convolution kernel size of the second dilated convolution is 3, the filter size is 128, and the dilation rate is 2; the convolution kernel size of the primary capsule layer is 3, the stride is 2, and the number of routing is set to 3. The number of capsules in the digital capsule layer is 15, and the vector dimension is 7.

[0023] Step d: Use dilated convolution to extract features from the LiDAR branch;

[0024] Specifically:

[0025] The LiDAR branch consists of six layers, and the structure is as follows: dilated convolution layer → Swish activation function → cascade block → max pooling layer → Swish activation function → cascade block; among them, the convolution kernel size of the first dilated convolution is 3×3, the filter size is 64, and the dilation rate is 2. Then it enters the cascade block, and the size of the max pooling layer is 2×2; the cascade block consists of twelve layers, and the structure is as follows: dilated convolution layer 1 → BN layer → LeakyReLU activation function → dilated convolution layer 2 → BN layer → LeakyReLU activation function → dilated convolution layer 1 → BN layer → LeakyReLU activation function → dilated convolution layer 2 → BN layer → LeakyReLU activation function. Add the same convolution operations mathematically, add the dilated convolutions 1 mathematically, and add the dilated convolutions 2 mathematically. The convolution kernel size of dilated convolution 1 is 3×3, the filter size is 128, and the dilation rate is 2. The convolution kernel size of dilated convolution 2 is 1×1, the filter size is 64, and the dilation rate is 3. The input filter size of the first cascade block is 64, and the input filter size of the second cascade block is 128.

[0026] Step e: Concatenate the features extracted from the hyperspectral image branch and the LiDAR data branch;

[0027] Step f: Finally, use softmax to classify the concatenated features to obtain the sample classification labels;

[0028] Specifically:

[0029] The data information extracted by the hyperspectral branch and the LiDAR branch is concatenated, and then operations are performed through a BN layer, a dropout layer, a Swish activation function, etc. After that, softmax is used to classify the features of the fully connected layer to obtain the sample classification labels.

[0030] Beneficial effects:

[0031] The present invention adopts the collaborative classification of hyperspectral and LiDAR data, utilizes the respective characteristics of hyperspectral and LiDAR data, complements each other's advantages, and improves the classification accuracy. Aiming at the fact that scalar neurons in traditional convolutional neural networks cannot express feature position information, and the spatial resolution decreases and detailed information is lost after continuous pooling and downsampling of images, the present invention uses dilated convolution and capsule network to replace the traditional CNN to extract features from space and spectrum respectively. The characteristic of the capsule network is that it takes vectors as the input of the network. Compared with the scalars in traditional convolutional neural networks, vectors can represent image features more completely, achieving the purpose of improving the recognition rate; all the important information of the feature states in capsule detection will be encapsulated by the capsules in the form of vectors. Compared with pooling in CNN, the capsule network has translational and rotational equivariance for images. Therefore, this method also makes progress in spatial information extraction. In a convolutional neural network, as the size of the convolutional kernel increases, the corresponding receptive field will become larger. In addition, the number of learning parameters will also increase, resulting in overfitting of data during the training process. Without introducing more parameters and without causing information loss, dilated convolution can obtain a larger receptive field than conventional convolution while keeping the size of the convolutional kernel unchanged, which is beneficial to enhancing the accuracy of pixel classification in image segmentation. Dilated convolution effectively increases the receptive field by injecting holes into the convolutional kernel without introducing new parameters. Description of the drawings

[0032] Figure 5 It is the flow chart of the collaborative classification of hyperspectral and LiDAR data in the method of the present invention.

[0033] Figure 1 It is the schematic diagram of the principle of the dual-branch network based on the collaborative classification of hyperspectral and LiDAR in the method of the present invention.

[0034] Figure 6 It is the schematic diagram of the principle of band selection by the pruning method in the method of the present invention.

[0035] Figure 2 It is the schematic diagram of the principle of the hyperspectral branch in the method of the present invention.

[0036] Figure 3 It is the schematic diagram of the principle of the LiDAR branch in the method of the present invention.

[0037] Figure 4It is a schematic diagram of the cascade block principle in the method of the present invention.

[0038] Figure 7 It is the result of dilated convolution with a continuous dilation rate of 2 in the method of the present invention.

[0039] Figure 8 It is the result of dilated convolution with dilation rates of 1 / 2 / 3 in the method of the present invention.

[0040] Figure 9 It is the ground truth map of the Houston dataset in the method of the present invention.

[0041] Figure 10 It is the HSI classification result map in the Houston dataset in the method of the present invention.

[0042] Figure 11 It is the LiDAR classification result map in the Houston dataset in the method of the present invention.

[0043] Figure 12 It is the classification result map of the Bi-CNN method in the Houston dataset in the method of the present invention.

[0044] Figure 13 It is the classification result map of the method adopted in the Houston dataset in the method of the present invention. Detailed implementation manners

[0045] The following further describes the detailed implementation manners of the present invention with reference to the accompanying drawings.

[0046] A hyperspectral and LiDAR data collaborative classification method based on a double-branch under this detailed implementation manner, the flowchart is as Figure 1 shown, and includes the following steps:

[0047] Step a, input the registered hyperspectral and LiDAR.tif data, and input the data into a double-branch network, as Figure 2 shown;

[0048] In this detailed implementation manner, the publicly available Houston dataset is adopted. The Houston dataset consists of HSI and LiDAR data introduced by the 2013 GRSS Data Fusion Contest. The data size is 349×1905 pixels, the spatial resolution is 2.5m, the HSI scene consists of 144 spectral bands, the wavelength range is 0.38μm to 1.05μm, and it includes 15 classes.

[0049] Step b, use a pruning method to perform band selection on the hyperspectral image;

[0050] For the hyperspectral branch, the entire HSI band is used as the input to train the original network parameters; for each band, all the parameters in the branch network are integrated to measure the band importance; under the constraint of the new band significance factor, the convolutional neural network is pruned to retain some representative weights, and the small sub-network is retrained to finally solve the hyperspectral band selection problem; specifically:

[0051] For the pruning method of the hyperspectral branch, it measures the importance of each band in a simple and effective way by calculating the band significance factor. The distribution of the band effective factor before and after pruning shows that the remaining band information is more effectively utilized. Using the entire spectral band HSI as the input, unimportant spectral bands and their corresponding kernels are automatically identified and pruned during the training process to obtain the selected spectral bands and a compact model.

[0052] Input the full-band HSI slices and the network structure, which are divided into a training stage and a band selection stage. Training stage: First, randomly initialize the CNN model parameters. Then, train the CNN model with all bands until the accuracy tends to be stable. Band selection stage: First, for each band, calculate the band significance factor. Then calculate the pruning band with the minimum band significance factor and its corresponding kernel matrix. Then retrain the pruned sub-network to restore the accuracy and refine the network by removing the pruned parameters.

[0053] Calculate the band significance factor. Let B represent the number of HSI bands, H / W represent the height / width of the input slice, and N represent the number of output feature maps. The first convolutional layer converts the input HSI slice into a feature map, which is achieved by adopting a 2D kernel of N×B and can be expressed on the HSI slice as:

[0054]

[0055] where * represents the convolution operation, X j represents the image of the j-th band, and Y i is the feature map generated by the first layer. All the kernels form the kernel matrix Because the bands with smaller kernel weights tend to produce feature maps with weak activations compared to other bands, the band significance factor of the j-th band in the HSI is defined as the sum of the absolute values of its corresponding kernel matrix, and the formula is as follows:

[0056]

[0057] The band significance factor reflects the influence of the input band on the output feature map. As Figure 3As shown, dark blocks represent convolution kernels with larger significance factors, and white blocks represent zero. To achieve band selection, the bands with the smallest significance factors and their corresponding kernel matrices are set to zero and marked as pruned parameters. These pruned parameters will not be updated in the back-propagation step. After pruning the bands and kernels, the performance degradation needs to be compensated by retraining the network. Finally, the network can be refined by removing these pruned parameters to obtain the selected bands and a compact model.

[0058] Step c: extracting features from the space and spectrum respectively; specifically:

[0059] Step c1: for the spatial branch of the hyperspectral spectrum, a 2D capsule network and a dilated convolution are used to extract the features of the spatial information of the hyperspectral spectrum;

[0060] The spatial branch of the hyperspectral image contains eight layers, and the structure is as follows: dilated convolution layer 1 → BN layer → Swish activation function → dilated convolution layer 2 → BN layer → Swish activation function → 2D primary capsule layer → digital capsule layer, as shown in Figure 4 As shown in the figure, the convolution kernel size of the first layer of dilated convolution is 3×3, the filter size is 256, and the dilation rate is 2. The convolution kernel size of the second layer of dilated convolution is 3×3, the filter size is 512, and the dilation rate is 3. The capsule network mainly includes the main capsule layer and the digital capsule layer. The convolution kernel size of the main capsule layer is 3×3, the step size is 2×2, and the number of routes is set to 3. The number of capsules in the output layer of the digital capsule layer is the same as the number of categories, the number of capsules is 15, and the capsule dimension is 7.

[0061] The spatial channel will be in pixels p ij The slice with the center and radius r is used as the input data. The purpose of being sent to the dilated convolution layer is to capture the features of the input data and output a feature map. The 2D capsule network is mainly composed of the main capsule layer and the digital capsule layer. The main capsule layer performs convolution processing on the feature map output by the convolution layer and converts the generated scalar into a vector. At this point, the capsule network completes the conversion from scalar to vector, that is, the construction of the capsule is completed. The last layer is the digital capsule layer. Similar to the fully connected layer, the size of the output vector modulus is used to measure the probability of a certain category. The vector with the largest modulus is the output category.

[0062] Step c2: for the spectral branch of the hyperspectral spectrum, a 1D capsule network and a dilated convolution are used to extract the spatial information of the hyperspectral spectrum;

[0063] The spectral branch of hyperspectral contains eight layers, and the structure is as follows: Dilated Convolution Layer 1 → BN Layer → Swish Activation Function → Dilated Convolution Layer 2 → BN Layer → Swish Activation Function → 1D Primary Capsule Layer → Digital Capsule Layer; among them, the convolution kernel size of the first dilated convolution is 11, the filter size is 64, and the dilation rate is 1. The convolution kernel size of the second dilated convolution is 3, the filter size is 128, and the dilation rate is 2; the convolution kernel size of the primary capsule layer is 3, the stride is 2, and the number of routings is set to 3. The number of capsules in the digital capsule layer is 15, and the vector dimension is 7.

[0064] Different from 2D-ConvCaps, 1D-ConvCaps uses capsules as the convolution unit and performs 1D convolution in the spectral direction to extract features between channels. The spectral channels are concentrated at the central pixel at position p ij It includes dilated convolution layers, batch normalization layers, Swish activation functions, and 1D capsule networks. The batch normalization involved allows for a higher learning rate to accelerate convergence and normalizes the data for each training mini-batch. After the convolution and max-pooling layers, the output spectral features are flattened

[0065] Then the spatial-spectral features are concatenated together and then input into the fully connected layer. The output of the fully connected layer can be expressed as where || represents the operation of concatenating the spatial and spectral feature vectors, and respectively represent the weights and biases of the fully connected layer. Then, the joint spatial-spectral features represented as F hsi are input into the softmax classification layer. In the designed dual-tunnel CNN branch, when forward calculation is required, the two tunnels move forward simultaneously, and the weight update based on backpropagation will follow the chain rule. Among them, different from CNN using neurons as data processing units, the capsule network combines all neurons that recognize the same object to form a capsule. These neurons contain all the information of various attributes in a specific pattern and can recognize a class of patterns in different directions and positions. Its input and output are vectors containing the feature information of each characteristic.

[0066] Step d: Use dilated convolution to extract features from the LiDAR branch;

[0067] ​The LiDAR branch contains six layers, and the structure is as follows: dilated convolutional layer → Swish activation function → cascading block → max pooling layer → Swish activation function → cascading block; among them, the convolutional kernel size of the first dilated convolution is 3×3, the filter size is 64, the dilation rate is 2, and it enters the cascading block. The max pooling layer size is 2×2; the cascading block contains twelve layers, and the structure is as follows: dilated convolutional layer 1 → BN layer → LeakyReLU activation function → dilated convolutional layer 2 → BN layer → LeakyReLU activation function → dilated convolutional layer 1 → BN layer → LeakyReLU activation function → dilated convolutional layer 2 → BN layer → LeakyReLU activation function. The same convolutional operations are added mathematically, the dilated convolutions 1 are added mathematically, and the dilated convolutions 2 are added mathematically. The convolutional kernel size of the dilated convolution 1 is 3×3, the filter size is 128, the dilation rate is 2, the convolutional kernel size of the dilated convolution 2 is 1×1, the filter size is 64, the dilation rate is 3, and the input filter size in the first cascading block is 64, and the input size of the second filter is 128.

[0068] The LiDAR branch uses dilated convolution to extract features from LiDAR elevation information. The LiDAR branch consists of a dilated convolutional layer, a Swish activation function layer, a cascading block, a max pooling layer, and a Flatten layer. Among them, the cascading design is used to combine different-level features from different layers to achieve feature reuse and propagation.

[0069] Figure 5 The flowchart of the LiDAR branch is shown, which is used to extract features in LiDAR, including a dilated convolutional layer, a cascading block, and a max pooling layer. In this branch, first, the LiDAR slice data, the central pixel L at position p ij is the input data, and the normalized data is first input into the network after the convolutional operation of the spatial kernel (3×3 kernel). After performing multi-scale feature extraction once through the first cascading module, a max pooling operation is performed, and the result after the max pooling operation is subjected to second multi-scale feature extraction through the second cascading module, thereby outputting multi-scale elevation information. ij , For preference, the cascading block is defined as

[0070] y

[0071] = g m (x1, {W m}) + x1 i y = g

[0072] s (x s , {W j}) + x s s

[0073] Among them, g m (x1, {W i}) and g s (x s , {W j}) represent the function mapping operations between two corresponding shortcut paths. Here, x and y are the cascaded input and output vectors respectively. Note that x1 and x s are the outputs of the first convolution and the Swish activation function respectively, and y m has the same dimension as x1.

[0074] The detailed architecture of the cascading block is as Figure 6 , for image classification, the cascading design is used to combine different-level features from different layers to achieve feature reuse and propagation. The cascading block consists of seven layers of operations, including a dilated convolutional layer, batch normalization, and the Swish activation function. These two paths bridge the first and intermediate convolutions and two activation operations. The paths pass the previous features to the subsequent layers through simple mathematical addition. In addition, the mathematical addition is performed channel-wise on two feature maps with the same shape. Then, the fused features are propagated to the next layer in the forward phase. In the backward propagation weight update phase, the chain rule is adopted. Among them, dilated convolution introduces the concept of dilation rate (r) on the basis of ordinary convolution and has the function of expanding the receptive field of the convolution kernel. By setting different dilation rates, multi-scale feature information of LiDAR images can be captured. Dilated convolution can expand the action range of the convolution kernel.

[0075] The calculation formulas for the dilated convolution kernel and the receptive field are:

[0076] n = k + (k - 1) × (r - 1)

[0077]

[0078] In the formula, k represents the size of the original convolution kernel; n represents the size of the dilated convolution kernel; l m-1 represents the receptive field size of the (m - 1)th layer; f m represents the size of the convolution kernel of the mth layer; l m represents the receptive field size of the mth layer after dilated convolution; S i represents the stride size of the ith layer.

[0079] For the design of the dilation rate in the model, if only the convolution kernels with the same dilation rate are stacked multiple times, a grid effect will be generated. The grid effect will cause discontinuity between pixels, there will be some holes, some pixels will be missed, resulting in local information loss and damaging the continuity of the information. These problems will reduce the training and testing accuracy of the classification model.

[0080] The dilation rate design must satisfy:

[0081] r max,i = max[r max,i+1 -2r i , r max,i+1 -2(r max,i+1 -r i ), r i

[0082] Where r i is the dilation rate of the i-th layer, and r max,i refers to the maximum dilation rate of the i-th layer. The hybrid dilated convolution requires that the dilation rates of the stacked convolutions cannot have a common divisor greater than 1. In this paper, the convolution kernel is dilated by the method of odd-even hybrid dilation rates, and the dilation rates are set to a cyclic structure of [1, 2, 3], which can cover each pixel point on the image and avoid information loss.

[0083] However, if non-coprime dilation coefficients are set in consecutive dilated convolution layers, a problem of discontinuous sampling of the feature map will occur, that is, the grid effect, resulting in the loss of a large amount of feature information. A structure called HDC (Hybrid Dilated Convolution) is designed to solve the problem of discontinuous convolution kernels, and it has the following characteristics:

[0084] First, the dilation rates of the stacked convolutions cannot have a common divisor greater than 1.

[0085] Second, the dilation rates are designed into a zigzag structure, such as [1, 2, 3, 1, 2, 3]. Figure 7 is the result of continuously performing dilated convolution with a dilation rate of 2, Figure 8 is the result of separately performing dilated convolutions with dilation rates of 1 / 2 / 3. The advantage of the latter is that a complete and continuous 3×3 region is retained from the beginning, and the subsequent dilation rate designs ensure the coherence of the receptive field, even if there is overlap, it is airtight.

[0086] Step e: Concatenate the features extracted from the hyperspectral image branch and the LiDAR data branch;

[0087] Step f: Finally, use softmax to classify the concatenated features to obtain the sample classification labels;

[0088] Concatenate the data information extracted from the hyperspectral branch and the LiDAR branch, and then perform operations through the BN layer, dropout layer, Swish activation function, etc. After that, use softmax to classify the fully connected features to obtain the sample classification labels.

[0089] Based on the comparison results in Table 1 and Table 2, it can be seen that by making full use of the advantages between different sensor data, the present invention effectively improves the fusion quality and classification accuracy of multi-sensor remote sensing images.

[0090] ​Table 1 Comparison of Classification Accuracies of Different Classification Methods for Houston Dataset (%)

[0091]

[0092]

[0093] Table 2 Comparison of Classification Accuracies of Different Classification Methods for Houston Dataset OA, AA and Kappa (%)

[0094] LiDAR HSI Bi-CNN Proposed Overall Accuracy OA 68.02±0.32 72.68±0.71 93.02±1.37 96.38±0.45 Average Accuracy AA 69.55±0.73 75.63±0.75 94.18±3.00 96.50±0.44 Kappa Coefficient 65.74±0.33 70.73±0.76 92.52±1.47 96.32±0.55

[0095] To subjectively evaluate the classification effect, Figure 9 、 Figure 10 、 Figure 11 、 Figure 12 and Figure 13 respectively show the true-color map of the Houston dataset and the pseudo-color maps of the classification results of each method. It can be seen that compared with the HSI single-branch, LiDAR single-branch and dual-branch CNN (Bi-CNN), the method in this paper is closer to the true ground object distribution, and the misclassified area is greatly reduced.

Claims

1. A collaborative classification method for hyperspectral and LiDAR data based on a dual-branch structure, characterized in that It includes the following steps: Step a: Input the registered.tif data of hyperspectral and LiDAR, and input the data into a dual-branch network; Step b: Use a pruning method to perform band selection on the hyperspectral image; Specifically: For the hyperspectral branch, use the entire HSI band as input to train the original network parameters; for each band, synthesize all the parameters in this branch network to measure the band importance; under the constraint of the new band significance factor, prune the convolutional neural network, retain some representative weights, and retrain the small sub-network to finally solve the hyperspectral band selection problem. Input the full-band HSI slices and network structure, which are divided into a training stage and a band selection stage. Training stage: First, randomly initialize the CNN model parameters. Then, train the CNN model with all bands until the accuracy tends to be stable. Band selection stage: First, for each band, calculate the band significance factor. Then calculate the pruning band with the minimum band significance factor and its corresponding kernel matrix. Then retrain the pruned sub-network to restore the accuracy and refine the network by removing the pruned parameters. Calculate the band significance factor. Let B denote the number of HSI bands, H / W denote the height / width of the input slice, and N denote the number of output feature maps. The first convolutional layer converts the input HSI slice into a feature map, which is achieved by adopting a 2D kernel of N×B on the HSI slice, and it can be expressed as: Among them, * represents the convolution operation, and X j represents the image of the j-th band, and Y i is the feature map generated by the first layer. All kernels form a kernel matrix Because the bands with smaller kernel weights tend to produce feature maps with weak activation compared to other bands, the band significance factor of the j-th band in HSI is defined as the sum of the absolute values of its corresponding kernel matrix, and the formula is as follows: The band significance factor reflects the influence of the input band on the output feature map. To achieve band selection, set the band with the minimum band effective factor and its corresponding kernel matrix to zero and mark them as pruned parameters. These pruned parameters will not be updated in the backpropagation step. After pruning the bands and kernels, it is necessary to retrain the network to compensate for the performance degradation. Finally, the network can be refined by removing these pruned parameters to obtain the selected bands and a compact model. Step c: Extract features from the space and spectrum respectively; Specifically: Step c1: For the spatial branch of the hyperspectral, use a 2D capsule network and dilated convolution to extract features from the spatial information of the hyperspectral; Step c2: For the spectral branch of the hyperspectral, use a 1D capsule network and dilated convolution to extract features from the spatial information of the hyperspectral; Step d: Use dilated convolution to extract features from the LiDAR branch; Step e: Concatenate the features extracted from the hyperspectral image branch and the LiDAR data branch; Step f: Finally, use softmax to classify the concatenated features to obtain the sample classification labels.

2. A method for collaborative classification of hyperspectral and LiDAR data based on a dual-branch according to claim 1, characterized in that The spatial branch of the hyperspectral in step c1 includes eight layers, and the structure is successively: dilated convolution layer 1 → BN layer → Swish activation function → dilated convolution layer 2 → BN layer → Swish activation function → 2D primary capsule layer → digital capsule layer; among them, the convolution kernel size of the first dilated convolution is 3×3, the filter size is 256, and the dilation rate is 2; the convolution kernel size of the second dilated convolution is 3×3, the filter size is 512, and the dilation rate is 3; the capsule network mainly includes a primary capsule layer and a digital capsule layer. The convolution kernel size of the primary capsule layer is 3×3, the stride is 2×2, the number of routings is set to 3, the number of capsules in the output layer of the digital capsule layer is the same as the number of classifications, the number of capsules is 15, and the capsule dimension is 7; The spectral branch of the hyperspectral in step c2 contains eight layers, and the structure is as follows: Dilated Convolution Layer 1 → BN Layer → Swish Activation Function → Dilated Convolution Layer 2 → BN Layer → Swish Activation Function → 1D Primary Capsule Layer → Digital Capsule Layer; among them, the convolution kernel size of the first dilated convolution is 11, the filter size is 64, and the dilation rate is 1. The convolution kernel size of the second dilated convolution is 3, the filter size is 128, and the dilation rate is 2; the convolution kernel size of the primary capsule layer is 3, the stride is 2, and the number of routing is set to 3. The number of capsules in the digital capsule layer is 15, and the vector dimension is 7.

3. A collaborative classification method for hyperspectral and LiDAR data based on a dual-branch, characterized in that, Step d is specifically as follows: The LiDAR branch contains six layers, and the structure is as follows: Dilated Convolution Layer → Swish Activation Function → Cascade Block → Max Pooling Layer → Swish Activation Function → Cascade Block; among them, the convolution kernel size of the first dilated convolution is 3×3, the filter size is 64, and the dilation rate is 2, and it enters the cascade block. The size of the max pooling layer is 2×2; the cascade block contains twelve layers, and the structure is as follows: Dilated Convolution Layer 1 → BN Layer → LeakyReLU Activation Function → Dilated Convolution Layer 2 → BN Layer → LeakyReLU Activation Function → Dilated Convolution Layer 1 → BN Layer → LeakyReLU Activation Function → Dilated Convolution Layer 2 → BN Layer → LeakyReLU Activation Function. The same convolution operations are added mathematically, the dilated convolutions 1 are added mathematically, and the dilated convolutions 2 are added mathematically. The convolution kernel size of the dilated convolution 1 is 3×3, the filter size is 128, and the dilation rate is 2. The convolution kernel size of the dilated convolution 2 is 1×1, the filter size is 64, and the dilation rate is 3. The input filter size of the first cascade block is 64, and the input size of the second filter is 128.

4. A collaborative classification method for hyperspectral and LiDAR data based on a dual-branch, characterized in that, Step f is specifically as follows: The data information extracted from the hyperspectral branch and the LiDAR branch is concatenated, and then operations are performed through the BN layer, dropout layer, Swish activation function, etc. After that, softmax is used to classify the fully connected features to obtain the sample classification labels.

Citation Information

Patent Citations

  • Classification method based on fusion of hyperspectral image and DSM data

    CN110210420A