Hyperspectral and laser radar combined classification method

Features are extracted through two-path neural networks and two-dimensional residual convolutional neural networks, combined with weighted fraction Fourier transform and Pareto soft optimization strategies, the problems of information magnitude imbalance and modal learning conflicts in the joint classification of hyperspectral and lidar are solved, and the classification accuracy is improved.

CN120298773APending Publication Date: 2025-07-11HARBIN ENG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510358333.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

There are serious information magnitude imbalance and modal learning conflicts in the existing joint classification methods of hyperspectral and lidar, resulting in insufficient classification accuracy.

Method used

The features of high-spectral and lidar images are extracted using two-path neural networks and two-dimensional residual convolutional neural networks, and feature fusion and model training are combined with weighted fraction Fourier transform and Pareto soft optimization strategies. The spatial spectral integration module and elevation information enhancement module are designed. Through the weighted fraction Fourier enhancement fusion module and Pareto soft optimization strategy, the feature contribution weight is adaptively adjusted to alleviate information order imbalance and modal learning conflicts.

Benefits of technology

The accuracy of the joint classification of hyperspectral and lidar has been improved, more accurate identification of geographic categories has been achieved, and the accuracy of classification results has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298773A_ABST
    Figure CN120298773A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral and laser radar combined classification method, and belongs to the technical field of hyperspectral and laser radar combined classification. The method aims at solving the problems that an existing fusion strategy is serious in information magnitude unbalance and modal learning conflicts exist in a model. According to the method, spatial features and spectral features of a hyperspectral image are extracted by adopting a dual-path neural network pair, and feature fusion is performed on the spatial features and the spectral features to obtain hyperspectral image features; performing feature extraction on the laser radar image by using a two-dimensional residual convolutional neural network to obtain a feature map; and performing feature fusion on the hyperspectral image and the laser radar image by using a weighted score Fourier enhancement fusion module, and obtaining a final classification result through a classifier. In a model training process, a Pareto soft optimization strategy is used to carry out iterative training on the model so as to obtain an optimal model parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hyperspectral and lidar joint classification, and particularly relates to a hyperspectral and lidar joint classification method. Background Art

[0002] With the advantage of extremely high spectral resolution of hyperspectral imaging technology, the obtained hyperspectral images can accurately reflect the physical characteristics of the observed targets. Such images include not only one-dimensional features representing spectral responses but also two-dimensional geometric information describing the spatial distribution of targets. Therefore, hyperspectral images are widely used in multiple fields, including pollution monitoring, urban planning, geological exploration, and precision agriculture, etc. Among them, hyperspectral image classification, as a key technology for interpreting hyperspectral images, utilizes a multi-dimensional feature space and assigns corresponding class labels to each pixel. Due to the phenomenon of "same spectrum but different objects" commonly existing in hyperspectral data, the classification accuracy is restricted. To break through this bottleneck, lidar data with ground elevation information is introduced to construct a multi-source remote sensing collaborative classification framework. In recent years, although significant progress has been made in the hyperspectral and lidar joint classification method, there are two key limitations in the existing research:

[0003] Firstly, the existing fusion strategies are difficult to effectively balance the information density differences of heterogeneous data, and there is a relatively serious problem of information magnitude imbalance;

[0004] Secondly, multi-modal features are prone to generate modal conflicts during the joint learning process, resulting in the model converging to a sub-optimal solution.

[0005] Therefore, the current effect of hyperspectral and lidar joint classification needs to be further improved. Summary of the Invention

[0006] The present invention aims to solve the problems of serious information magnitude imbalance existing in the existing fusion strategies and modal learning conflicts existing in the model.

[0007] A hyperspectral and lidar joint classification method uses a joint classification model to perform hyperspectral and lidar joint classification. The classification process of the joint classification model includes:

[0008] S1. For the hyperspectral image, a dual-path neural network is used to extract spatial features and spectral features. One path of the dual-path neural network extracts spectral features based on one-dimensional convolution B is the dimension of the latent space; the other path uses a Transformer block to capture spatial features n is the number of tokens; the spatial features and spectral features are fused to obtain hyperspectral image features, denoted as integrated features

[0009] For lidar images, a two-dimensional residual convolutional neural network is used to extract features to obtain a feature map, denoted as elevation features.

[0010] S2. Based on Obtain where Ψ is the kernel function of the ξ-order weighted fractional Fourier transform, x(n) is the value corresponding to the channel dimension with sequence index n in, and u represents the fractional domain transformation plane;

[0011] Combine together, and then process through a linear layer and a Sigmoid activation function to obtain the dynamic weight β; then, use the dynamic weight β to re-weight, and after a non-linear mapping, obtain the preliminary fusion features where W1 represents the weight, GN represents group normalization, represents matrix multiplication;

[0012] Obtain the fusion information through weighted summation where ε is a weighting coefficient;

[0013] Subsequently, obtain the correlation features by multiplying the two fusion features where W2 represents the parameter of the linear layer; then apply a residual connection to obtain the final fusion features where W3 represents the parameter of the linear layer;

[0014] Finally, input into the classifier Classifier3 to obtain the class prediction probability, and the ground object classification result can be obtained according to the maximum probability value index.

[0015] Furthermore, for hyperspectral images, the process of using a dual-path neural network to extract spatial and spectral features includes:

[0016] For the hyperspectral image Take the central pixel x i ′ as the input, which is first input into a 1D residual network for preliminary spectral feature extraction of the HSI; the 1D residual network includes a 1D convolutional layer, a batch normalization layer, an activation layer, and a 1D residual unit; subsequently, an average pooling layer and a linear layer are introduced, and the feature obtained by the residual network is denoted as The subscript spec represents spectral features;

[0017] Meanwhile, combine Divide into sub-patches through patch embedding, and input the output of patch embedding into multiple Transformer modules, where the Transformer module includes a normalization layer, a multi-head attention layer, and a multi-layer perceptron layer; finally, obtain the output The subscript spac represents spatial features.

[0018] Furthermore, the 1D residual unit includes a 1D convolutional layer, a batch normalization layer, and an activation layer.

[0019] Furthermore, the process of feature fusion of spatial features and spectral features to obtain hyperspectral image features includes:[[]]

[0020] Denote the spectral feature and the spatial feature as Use a vector expansion operation to expand the spectral feature To the shape of the spatial feature And apply a convolutional layer to generate V spec And Q, apply a convolutional layer to the spatial feature To generate V spa And K;

[0021] Apply the Softmax layer to generate the spatial attention map A = Softmax(exp(QK T ), where exp(·) represents the exponential function;

[0022] Multiply the spatial attention map by V spec And V spa Respectively through matrix multiplication to obtain the spatially reconstructed feature map:

[0023]

[0024] Among them, Trans spec And Trans spa Represents a transformation function composed of convolutional layers;

[0025] Based on Obtain the integrated feature

[0026]

[0027] Among them, Pool is the pooling layer processing, MLP represents the MLP layer processing, Sigmoid represents the Sigmoid function layer; ⊙ represents the Hadamard product operation, and σ is the balance hyperparameter.

[0028] Furthermore, the process of feature extraction using a two-dimensional residual convolutional neural network includes:

[0029] Duplicate the lidar image according to the number of hyperspectral image channels to obtain Will Input into the 2D residual network Φ 1 Among them, the 2D residual network mainly includes a 2D convolutional layer, a batch normalization layer, a ReLU activation function, and a 2D residual unit; the 2D residual network outputs features At the same time, extract the central pixel from the sample and expand it to the same size as by replication, denoted as x i ″; Input x i ″ into the 2D residual network 1 with the same structure as Φ and sharing parameters for feature extraction to obtain elevation features Then calculate the feature similarity:

[0030]

[0031] where x a and z b represent the input features with spatial position indices a and b respectively, φ(·) represents the 2D convolutional layer, ||·|| represents the L2 norm, and T represents the transpose operation;

[0032] Calculate the affinity between the pixel with index value a and other pixels through inner product, and suppress negative values through the ReLU activation function, and then calculate the weighted representation of z b and the affinity: where S(x a ) represents

[0033] Connect the output and the input feature with residual connection to obtain the feature Then send it into a 2D residual network Φ with GAPooling 2 to obtain the elevation feature of LiDAR

[0034] Furthermore, the 2D residual unit includes a 2D convolutional layer, a batch normalization layer, and an activation layer.

[0035] Furthermore, for the hyperspectral image, the number of tokens n corresponding to the spatial feature =(p / 2) 2 +1, where p is the patch size of the hyperspectral image

[0036] Furthermore, the kernel function Ψ of the ξ-order weighted fractional Fourier transform is as follows:

[0037] ​

[0038] where ω c (ξ) is the weighting coefficient of the weighted fractional Fourier transform, c ∈ (0, 1, 2, 3); j represents the imaginary unit.

[0039] Furthermore, the weighting coefficient of the weighted fractional Fourier transform where ξ is an adaptive hyperparameter.

[0040] Furthermore, the joint classification model is iteratively trained using the Pareto soft optimization strategy, and the specific process includes:

[0041] Set two additional classifiers, and the additional classifiers are respectively loaded on the HSI features and LiDAR features after that;

[0042] Based on the two additional classifiers and the classifier Classifier3, calculate the cross-entropy loss of the classifiers respectively, and obtain the loss L calculated by the cross-entropy function H 、L L and L M ;

[0043] In SGD optimization, for any loss L ε , ε ∈ (m, u), where m represents M, u represents H or L; in the t-th mini-batch S, the gradient of the parameter θ k is expressed as:

[0044]

[0045] where ▽ represents the differential operator; |S| represents the capacity of the mini-batch S;

[0046] Express the gradient of modality k as k = h, l represent HSI or LiDAR respectively; according to the Pareto principle, when there are many target tasks, at each model iteration, the gradients for different targets will be re-weighted and integrated; finally, the model will converge to a compromise state along the gradient direction optimized by Pareto, meeting Pareto optimality; the formula for specifically calculating the Pareto optimal solution is:

[0047]

[0048] s.t. η m , η u ≥0, η m + η u = 1,

[0049] where ||·|| represents the L2 norm, η m 、ηu respectively represent for L m 、L u the calculated gradients the corresponding weighting coefficients;

[0050] Obtain the Pareto optimal solution of the loss through calculation:

[0051]

[0052] Here represents the total gradient for mode k, and the superscript Pareto represents Pareto;

[0053] Based on the Pareto optimal solution Improve and process according to two different scenarios:

[0054] Conflict-free situation: First, calculate the gradient angle β between When cosβ > 0, there is no need to adopt the Pareto optimization strategy, and the gradient at this time is the unified gradient

[0055] Conflict situation: When cosβ < 0, based on the Pareto optimal solution calculate the Pareto coefficient to solve and Obtain the final gradient based on the enhancement factor μ > 1:

[0056] Beneficial effects:

[0057] In the hyperspectral and lidar joint classification algorithm proposed by the present invention, in the hyperspectral and lidar feature extraction stages, a spatial-spectral integration module and an elevation information enhancement module are designed, effectively utilizing the patch features sharing the same land cover label with the central pixel; in the information fusion stage, a weighted fractional Fourier enhanced fusion module is designed, which applies the weighted fractional Fourier transform to the hyperspectral features, thereby enhancing their representation and combining them with the lidar features to generate balanced fusion features; the fused features are sent to the classifier to obtain the classification result map. In addition, a Pareto soft optimization strategy tailored for hyperspectral and lidar joint classification is introduced in the training stage, effectively preventing optimization conflicts between the two modes. The present invention can adaptively adjust the feature contribution weights of hyperspectral and lidar data, effectively alleviating the problem of information magnitude imbalance; secondly, it can effectively achieve the dynamic balance of collaborative learning and competition inhibition between modes, reducing the modal learning conflicts in the model and enabling more accurate identification of ground object categories. Through experimental analysis, the method proposed by the present invention can obtain OA values of 0.9232, 0.9366, and 0.9165 on three datasets respectively. Description of the Drawings

[0058] Figure 1 It is a flowchart of a hyperspectral and lidar joint classification method based on Pareto optimization and fractional Fourier enhanced fusion network;

[0059] Figure 2 It is a schematic diagram of a residual network. Among them, (a) is a one-dimensional residual network, and (b) is a two-dimensional residual network;

[0060] Figure 3 It is a schematic diagram of a spatial-spectral integration module;

[0061] Figure 4 It is a schematic diagram of a weighted fractional Fourier enhanced fusion module;

[0062] Figure 5 They are the hyperspectral false-color image, lidar grayscale image, ground truth map and classification result map of dataset I. (a) is the hyperspectral false-color image, (b) is the lidar grayscale image, (c) is the ground truth map, and (d) is the classification result map;

[0063] Figure 6 They are the hyperspectral false-color image, lidar grayscale image, ground truth map and classification result map of dataset II. (a) is the hyperspectral false-color image, (b) is the lidar grayscale image, (c) is the ground truth map, and (d) is the classification result map;

[0064] Figure 7 They are the hyperspectral false-color image, lidar grayscale image, ground truth map and classification result map of dataset III. (a) is the hyperspectral false-color image, (b) is the lidar grayscale image, (c) is the ground truth map, and (d) is the classification result map. Detailed implementation mode

[0065] Detailed implementation mode one: Combined with Figure 1 Illustrate this implementation mode,

[0066] This implementation mode is a hyperspectral and lidar joint classification method, which includes the following steps:

[0067] S1. Extract features from the hyperspectral image and the lidar image respectively. For the hyperspectral image, a dual-path neural network is used to extract spatial features and spectral features, and after feature fusion, hyperspectral image features are obtained; for the lidar image, a two-dimensional residual convolutional neural network is used for feature extraction.

[0068] S2. Use the weighted fractional Fourier enhanced fusion module to fuse the features of the hyperspectral and lidar images, and obtain the final classification result through a classifier.

[0069] S3. The present invention uses the Pareto soft optimization strategy to iteratively train the model to obtain optimal model parameters.

[0070] It can be seen that the present invention is essentially a hyperspectral and lidar joint classification method based on the fusion network of Pareto optimization and fractional Fourier transform enhancement. The specific processing process is as follows:

[0071] S1. Extract features from the hyperspectral image and the lidar image respectively:

[0072] In order to extract rich spectral and spatial information in the hyperspectral image (HSI), a dual-path network is adopted to achieve spatial-spectral feature extraction. One path has excellent learning ability for high-dimensional sequences by means of one-dimensional convolution, and uses one-dimensional residual blocks (1D ResBlocks) to extract spectral features; at the same time, the other path uses Transformer blocks to capture spatial features, so as to effectively capture global context information. The lidar (LiDAR) data provides spatial geometric information and elevation details of the surface coverage. Since the two-dimensional convolutional neural network (2D CNN) has excellent performance in extracting 2D spatial features, this method uses 2D ResBlocks to extract LiDAR features. The specific process includes:

[0073] Step 1.1. Hyperspectral image feature extraction:

[0074] Represent the HSI sample set as The corresponding label is denoted as where p is any sample in the HSI sample set The patch size (a small-sized image of size p×p centered on a certain pixel cut from a hyperspectral image), d is the spectral dimension; the superscript h represents HSI, i represents the sample index value, N is the total number of samples, and C is the number of classes.

[0075] To extract spectral features, take The central pixel of as the input. As shown in Figure 2 (a), x i ′ is first input into the 1D residual network for preliminary spectral feature extraction of HSI.

[0076] The 1D residual network mainly includes a 1D convolutional layer, a batch normalization (BN) layer, an activation layer (LeakyReLU), and a residual unit (1D ResBlock). The residual unit includes a 1D convolutional layer, a batch normalization (BN) layer, and an activation layer (Leaky ReLU). Subsequently, an average pooling (AvgPool1D) layer and a linear layer are introduced to further map the features obtained by the residual network to a deeper latent space. The output of the above entire network can be expressed as:

[0077]

[0078] where B is the dimension of the latent space, represents the entire above network, and represents the spectral characteristics of the HSI, with the subscript spec representing the spectrum.

[0079] In terms of spatial feature extraction, the Vision Transformer (ViT) is used to extract spatial domain features. As Figure 1 shown, first, is divided into smaller sub-patches through patch embedding, and learnable class tokens and position embeddings are added. Subsequently, the output of the patch embedding is input into the Transformer module, which mainly includes a normalization layer, a multi-head attention layer (4 attention heads), and a multi-layer perceptron (MLP) layer. To extract deeper features, 4 Transformer blocks are stacked in this embodiment to generate the final output:

[0080]

[0081] where, represents all Transformer blocks, the number of tokens n = (p / 2) 2 +1, and represents the spatial features of the HSI, with the subscript spac representing the space.

[0082] To minimize the interference from irrelevant pixel features, as Figure 3 shown, a Spatial-Spectral Integration Module (SSIM) is designed to integrate the spectral and spatial features of the HSI. First, the spectral feature is expanded to the shape using a vector expansion operation, and a convolutional layer is applied to generate V spec and Q, where, Similarly, a convolutional layer is applied to the spatial feature to generate V spa and K, where V spa , Next, the matrix multiplication of Q and K is performed to obtain the feature map, and a Softmax layer is applied to generate the spatial attention map, denoted as:

[0083] A = Softmax(exp(QK T )) (3)

[0084] where exp(·) represents the exponential function.

[0085] Next, multiply the spatial attention map with V spec and V spa through matrix multiplication to obtain the feature map after spatial reconstruction:

[0086]

[0087] where, Trans spec and Trans spa represent transformation functions composed of convolutional layers.

[0088] To further enhance the feature representation, a channel attention module is applied to the features. First, use the global pooling (GAPooling) layer to aggregate the features into a single channel respectively; then, an MLP layer learns the interdependencies between channels and uses the Sigmoid layer to generate a channel weighting vector, and then uses this vector to re-weight the feature map; finally, combine the output with the features through a residual connection and adaptively weight the spatial and spectral features to obtain the final integrated features The specific calculation process is as follows:

[0089]

[0090] where, ⊙ represents the Hadamard product operation, σ is a balancing hyperparameter, and

[0091] Step 1.2, LiDAR image feature extraction:

[0092] First, copy the LiDAR image to match the number of channels of the HSI to generate a sample set where, p is any sample in the LiDAR sample set The patch size, d is the channel dimension; the superscript l represents LiDAR, i represents the sample index value, and N is the total number of samples.

[0093] Subsequently, input into a 2D residual network to extract high-level semantic features. As Figure 2 (b) shows, the 2D residual network mainly includes a 2D convolutional layer, a batch normalization (BN) layer, a ReLU activation function, and a residual unit (2DResBlock). The residual unit includes a 2D convolutional layer, a batch normalization (BN) layer, and an activation layer (ReLU). Then the output of the entire 2D residual network can be expressed as:

[0094]

[0095] where, Φ 1 represents the entire 2D residual network, Represents the shallow features of LiDAR.

[0096] To further extract elevation information, first, extract the central pixel from the sample and expand it to size p×p by replication, thus obtaining Subsequently, input x i ″ into a 2D residual network with the same structure as Φ 1 and sharing parameters for feature extraction, obtaining the elevation feature representation:

[0097]

[0098] where represents the 2D residual network sharing parameters with Φ 1 , represents the elevation feature of LiDAR.

[0099] Similarly, to suppress the interference of irrelevant pixels in the LiDAR features while retaining important geometric information, as Figure 1 shown, this method uses an Elevation Information Enhancement Module (EIEM) to integrate pixel-level features. Specifically, use cosine distance to calculate the feature similarity between the pure elevation information and the patch-level feature pixels:

[0100]

[0101] where x a and z b represent the input features with their spatial position indices being a and b, φ(·) represents a 2D convolutional layer, ||·|| represents the L2 norm, and T represents the transpose operation.

[0102] Then, in the normalized feature space, calculate the affinity between the pixel with index a and other pixels through inner product, and suppress negative values through the ReLU activation function. The resulting LiDAR feature is the weighted representation of z b and the affinity:

[0103]

[0104] where S(x a ) represents

[0105] To retain the original spatial geometric features of LiDAR, connect the output and the input feature with a residual connection to obtain the feature

[0106] Then, through the 2D residual network with the same structure as in Figure 2 (b), deeper features are extracted, and then the final extracted features are obtained through spatial dimension GAPooling, which are represented as:

[0107]

[0108] where, Φ 2 represents the 2D residual network with GAPooling, represents the high-level elevation features of LiDAR.

[0109] S2. Feature fusion of hyperspectral and LiDAR images:

[0110] As Figure 4 shown, based on the Weighted Fractional Fourier Enhanced Fusion Module (WFrFEF) to aggregate HSI and LiDAR features, that is As a variant of the FrFT, the Weighted Fractional Fourier Transform (WFrFT) shows a more uniform energy distribution and can be more widely extended in the time-frequency plane. Utilizing this property, the WFrFT is applied to the HSI features to effectively reduce redundant information, thereby effectively alleviating the problem of information magnitude imbalance in modal feature fusion. Specifically, the WFrFT of the HSI features is expressed as:

[0111]

[0112] where, Ψ is the kernel function of the weighted fractional Fourier transform of order ξ, x(n) is the value corresponding to the channel dimension with sequence index n, and u represents the fractional domain transformation plane. The definition of the kernel function is as follows:

[0113]

[0114] where, ω c (ξ) is the weighting coefficient of the weighted fractional Fourier transform, c ∈ (0, 1, 2, 3); j represents the imaginary number;

[0115] ω c (ξ) is specifically as follows:

[0116]

[0117] In this method, ξ is an adaptive and learnable hyperparameter.

[0118] Obtained after WFrFT Initially Concatenated together, and then processed through a linear layer and a Sigmoid activation function to obtain the dynamic weight β. Then, use this dynamic weight to re-weight these two features and pass through a non-linear mapping layer to obtain the preliminary fused feature:

[0119]

[0120] Among them, W1 represents the learnable weight, GN represents group normalization, represents matrix multiplication.

[0121] To retain the original information of HSI features and LiDAR features and enhance the stability of feature propagation, the fused information is obtained through a weighted summation process Denoted as Among them, ε is a learnable weighting coefficient, contains the original high-level information.

[0122] Subsequently, by multiplying the two fused features, the correlation feature is obtained:

[0123]

[0124] Among them, W2 represents the parameter of the linear layer.

[0125] Next, apply a residual connection to obtain the final fused feature:

[0126]

[0127] Among them, W3 represents the parameter of the linear layer, which is used to adjust the dimension through a linear transformation.

[0128] Subsequently Input into the classifier (Classifier3) to obtain the class prediction probability, and the ground object classification result can be obtained according to the index of the maximum probability value.

[0129] S3. Iteratively train the model using the Pareto soft optimization strategy:

[0130] Differences in the feature fitting capabilities of hyperspectral images and lidar in the network may lead to an unbalanced learning process between them, which may result in underutilization of a certain modality. To solve this problem, set two additional classifiers, and the additional classifiers are respectively loaded on the HSI features and LiDAR features After that, the learning ability of the single modality is enhanced. However, during the backpropagation process, the optimization for the shared objective may lead to optimization conflicts. Therefore, the present invention proposes a Pareto-based soft optimization strategy, called HLPareto, for HSI and LiDAR feature learning. Specifically, a classifier is designed and applied to calculate the classification loss, L Cls denotes the cross-entropy loss of the classifier, and its calculation formula is:

[0131]

[0132] where, denotes the probability output of the classifier; is the label, s is the category, and C is the number of categories.

[0133] First, three classifiers C H 、C L and C M are designed for the HSI branch, LiDAR branch, and fused features respectively, and the cross-entropy function is calculated by formula (17) to calculate the losses L H 、L L and L M .

[0134] In SGD optimization, for any loss L ε , ε ∈ (m, u), where m represents M, and u represents H or L. In the t-th mini-batch S, the gradient of the parameter θ k is expressed as:

[0135]

[0136] where, denotes the differential operator; |S| represents the capacity of the mini-batch S.

[0137] To simplify the calculation, the gradient of modality k is expressed as k = h, l represent HSI or LiDAR respectively. According to the Pareto principle, when there are many target tasks, during each model iteration, the gradients for different targets will be re-weighted and integrated. Finally, the model will converge to a compromise state along the gradient direction optimized by Pareto, conforming to Pareto optimality. The formula for specifically calculating the Pareto optimal solution is:

[0138]

[0139] where, ||·|| represents the L2 norm, η m 、η u represent the gradients m 、L u calculated for L The corresponding weighting factor.

[0140] The Pareto optimal solution of the loss can be obtained through calculation:

[0141]

[0142] Here represents the total gradient for mode k, and the superscript Pareto represents Pareto. Due to the implicit influence of the noise intensity of SGD, directly applying the Pareto algorithm to the multi-modal model fails to achieve the expected performance. Therefore, the present invention proposes a more flexible Pareto optimization method called HLPareto. Specifically, HLPareto aims to handle two different scenarios:

[0143] Conflict-free situation: First, calculate the gradient angle β between . When there is no conflict between the gradients, that is, then there is no need to adopt the Pareto optimization strategy, and the gradient at this time is the unified gradient

[0144]

[0145] Conflict situation: When a gradient conflict occurs, that is, cosβ < 0, calculate the Pareto coefficient based on equation (20) to solve To achieve a smoother adjustment rather than a rigid direction change, the original gradient is combined into the Pareto-based solution. In addition, to enhance the generalization ability, an enhancement factor μ > 1 is introduced during the gradient optimization process to obtain the final gradient:

[0146]

[0147] Generally speaking, the complete algorithm of the present invention is shown in Table 1, and its schematic diagram is as shown in Figure 1 shown.

[0148] It should be noted that: for the conflict situation, actually η is first calculated and determined according to formula (2) m and η u , and then it can be calculated according to formula (22).

[0149] Table 1

[0150]

[0151]

[0152] After the above training process, the overall trained network model is obtained. During the actual hyperspectral and lidar joint classification process, the joint classification result is obtained by using Classifier3.

[0153] Embodiment

[0154] In this part, the effectiveness of the hyperspectral and lidar joint classification method based on Pareto optimization and fractional Fourier enhanced fusion network proposed by the present invention is illustrated using three commonly used datasets. The detailed information of the three datasets used is listed in Table 1. The experimental results adopt the overall accuracy (OA), specific class accuracy, average accuracy (AA), and Kappa coefficient as evaluation indicators. The higher the value of all evaluation indicators, the better the classification effect.

[0155] Table 1 Details of the hyperspectral and lidar images used

[0156]

[0157] For different data, the optimal parameter settings of the method of the present invention are shown in Table 2. The parameter μ is the enhancement factor, and p is the image patch size. The invented algorithm is implemented using the PyTorch framework in Python 3.7; all experiments are carried out on the same hardware platform: GTX-3090 GPU, Intel 4210R CPU, 30GB of memory. In the network architecture, the spatial size of the tokens in the Transformer block is set to 2×2, and the embedding dimension is 192. In addition, during training, the batch size is set to 32, the number of epochs is set to 100, SGD is used as the optimizer, the momentum is 0.9, and the weight decay is set to 1e -4 , and the learning rate is set to 5e -3 .

[0158] Table 2 Optimal parameters and evaluation index values on three groups of experimental data

[0159]

[0160] The classification results of Dataset I to Dataset III are as Figures 5 - 7 shown, where Figure 5 are the hyperspectral false color map, lidar grayscale map, ground truth map, and classification result map of Dataset I. (a) is the hyperspectral false color map, (b) is the lidar grayscale map, (c) is the ground truth map, and (d) is the classification result map. Figure 6 are the hyperspectral false color map, lidar grayscale map, ground truth map, and classification result map of Dataset II. (a) is the hyperspectral false color map, (b) is the lidar grayscale map, (c) is the ground truth map, and (d) is the classification result map. Figure 7 are the hyperspectral false color map, lidar grayscale map, ground truth map, and classification result map of Dataset III. (a) is the hyperspectral false color map, (b) is the lidar grayscale map, (c) is the ground truth map, and (d) is the classification result map.

[0161] The above numerical examples of the present invention are only used to illustrate in detail the calculation model and calculation process of the present invention, rather than to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or modifications derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. A hyperspectral and lidar joint classification method, characterized in that, Hyperspectral and lidar joint classification is performed using a joint classification model, and the classification process of the joint classification model includes: S1. For hyperspectral images, a dual-path neural network is used to extract spatial features and spectral features. One path of the dual-path neural network extracts spectral features based on one-dimensional convolution B is the dimension of the latent space; the other path uses Transformer blocks to capture spatial features n is the number of tokens; the spatial features and spectral features are fused to obtain hyperspectral image features, denoted as integrated features For lidar images, a two-dimensional residual convolutional neural network is used to extract features to obtain a feature map, denoted as elevation features S2. Based on obtain where Ψ is the kernel function of the ξ-order weighted fractional Fourier transform, x(n) is the value corresponding to the channel dimension with sequence index n in and u represents the fractional domain transform plane; Concatenate with , then process it through a linear layer and a Sigmoid activation function to obtain the dynamic weight β; then, use the dynamic weight β to re-weight and , and through a layer of non-linear mapping, obtain the preliminary fusion feature where W1 represents the weight, GN represents group normalization, represents matrix multiplication; The fused information is obtained by weighted summation where ε is a weighting coefficient; Subsequently, the correlation feature is obtained by multiplying the two fusion features where W2 represents the parameter of the linear layer; then the residual connection is applied to obtain the final fusion feature where W3 represents the parameter of the linear layer; Finally, it is input into the classifier Classifier3 to obtain the class prediction probability, and the ground object classification result can be obtained according to the maximum probability value index.

2. The hyperspectral and lidar joint classification method according to claim 1, wherein For hyperspectral images, the process of using a dual-path neural network to extract spatial and spectral features includes: For hyperspectral images Take The central pixel x i ′ as the input, which is first input into a 1D residual network for preliminary spectral feature extraction of the HSI; the 1D residual network includes a 1D convolutional layer, a batch normalization layer, an activation layer, and a 1D residual unit; subsequently, an average pooling layer and a linear layer are introduced, and the feature representation obtained by the residual network is The subscript spec represents spectral features; Meanwhile, The patch embedding is divided into sub-patches through patch embedding, and the output of the patch embedding is input into multiple Transformer modules, where the Transformer module includes a normalization layer, a multi-head attention layer, and a multi-layer perceptron layer; finally, the output is obtained. The subscript spac represents the spatial feature.

3. A hyperspectral and lidar joint classification method according to claim 2, characterized in that, The 1D residual unit includes a 1D convolutional layer, a batch normalization layer, and an activation layer.

4. A hyperspectral and lidar joint classification method according to claim 1, characterized in that The process of performing feature fusion on spatial and spectral features to obtain hyperspectral image features includes: Denote the spectral feature and the spatial feature as and Use the vector expansion operation to expand the spectral feature into the shape of the spatial feature and apply a convolutional layer to generate V spec and Q. Apply a convolutional layer to the spatial feature to generate V spa and K; Apply the Softmax layer to generate the spatial attention map A = Softmax(exp(QK T ))), where exp(·) represents the exponential function; Multiply the spatial attention map with V and V respectively through matrix multiplication to obtain the feature map after spatial reconstruction: spec and V spa ​ Among them, Trans spec and Trans spa represent transformation functions composed of convolutional layers; Based on Integrated features are obtained Among them, Pool represents pooling layer processing, MLP represents MLP layer processing, Sigmoid represents the Sigmoid function layer; ⊙ represents Hadamard product operation, and σ is a balance hyperparameter.

5. A hyperspectral and lidar joint classification method according to claim 1, characterized in that For lidar images, the process of using a two-dimensional residual convolutional neural network to extract features includes: The lidar image is replicated according to the number of hyperspectral image channels to obtain Input into the 2D residual network Φ 1 . The 2D residual network mainly includes a 2D convolutional layer, a batch normalization layer, a ReLU activation function, and a 2D residual unit; the 2D residual network outputs features At the same time, extract the central pixel from the sample and expand it to a feature with the same size as by replication, denoted as x″ i ; Input x″ i into a 2D residual network 1 with the same structure as Φ and sharing parameters for feature extraction to obtain elevation features Then calculate the feature similarity: where x a and z b represent the input features and whose spatial position indices are a and b, φ(·) represents a 2D convolutional layer, ||·|| represents the L2 norm, and T represents the transpose operation; Calculate the affinity between the pixel with index value a and other pixels through the inner product, and suppress negative values through the ReLU activation function to calculate z b Weighted representation of the affinity: Among them, S(x a ) represents The output and the input features are connected by a residual connection to obtain features which are then fed into a 2D residual network Φ with GAPooling 2 to obtain the elevation features of LiDAR 6. A hyperspectral and lidar joint classification method according to claim 5, characterized in that The 2D residual unit includes a 2D convolutional layer, a batch normalization layer, and an activation layer.

7. A hyperspectral and lidar joint classification method according to claim 1, characterized in that, For hyperspectral images, the spatial features The corresponding number of tokens n = (p / 2) 2 + 1, where p is the patch size of the hyperspectral image of the hyperspectral image.

8. A hyperspectral and lidar joint classification method according to claim 1, characterized in that The kernel function Ψ of the ξ-order weighted fractional Fourier transform is as follows: Among them, ω c (ξ) is the weighting coefficient of the weighted fractional Fourier transform, c ∈ (0, 1, 2, 3); j represents the imaginary number.

9. A hyperspectral and lidar joint classification method according to claim 8, characterized in that Weighting Coefficient of Weighted Fractional Fourier Transform where ξ is an adaptive hyperparameter.

10. A hyperspectral and lidar joint classification method according to any one of claims 1 to 9, characterized in that, The joint classification model is iteratively trained using the Pareto soft optimization strategy, and the specific process includes: Set two additional classifiers, and the additional classifiers are respectively loaded after the HSI features and the LiDAR features afterwards; Based on two additional classifiers and classifier Classifier3, calculate the cross-entropy loss of the classifiers respectively to obtain the loss L calculated by the cross-entropy function H , L L and L M ; In SGD optimization, for any loss L ε , ε ∈ (m, u), where m represents M, and u represents H or L; in the t-th mini-batch S, the gradient of the parameter θ k is expressed as: Among them, represents a differential operator; |S| represents the capacity of the mini-batch S; The gradient of modality k is expressed as respectively represent HSI or LiDAR; According to the Pareto principle, when there are many target tasks, during each model iteration, the gradients for different targets will be reweighted and integrated; Eventually, the model will converge to a compromise state along the gradient direction optimized by Pareto, which conforms to Pareto optimality; The formula for specifically calculating the Pareto optimal solution is: s.t. η m , η u ≥0, η m + η u =1, where, ||·|| represents the L2 norm, η m , η u respectively represent the gradients m , L u calculated for the corresponding weighting coefficients; The Pareto optimal solution of the loss is obtained by calculation: Here represents the total gradient for mode k, where the superscript Pareto represents Pareto; Based on the Pareto optimal solution Improve and process according to two different scenarios: Conflict-free situation: First, calculate and the gradient angle β therebetween, and calculate When cosβ > 0, the Pareto optimization strategy does not need to be adopted, and the gradient at this time is the unified gradient Conflict situation: when cosβ < 0, based on the Pareto optimal solution calculate the Pareto coefficient to solve and obtain the final gradient based on the enhancement factor μ > 1:

Citation Information

Cited By

  • Joint classification method for hyperspectral image and laser radar data

    CN121305244A