Local and global multi-resolution LiDAR-hyperspectral image collaborative classification method
By adopting local-global directed MLP structure and multi-resolution multi-scale feature extraction strategy in the hyperspectral-LiDAR collaborative classification method, combining the local and global features of hyperspectral images, and through the feature enhancement module based on LiDAR information, the problem that existing methods are difficult to effectively combine local and global information is solved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202510265456.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing hyperspectral-LiDAR collaborative classification method is difficult to effectively combine local and global information, and ignores feature information at different resolutions and scales, resulting in insufficient classification accuracy.
The local-global directed MLP structure and multi-resolution multi-scale feature extraction strategy are adopted, and the local and global features of hyperspectral images are combined, and the elevation information is introduced to improve feature representation through a feature enhancement module based on LiDAR information. Use cross-attention mechanisms to narrow the semantic gap between hyperspectral and LiDAR data.
The classification accuracy and effect of hyperspectral-LiDAR image collaborative classification is improved. By effectively combining local and global information, the land cover characteristics at different scales are accurately obtained, which alleviates the problem of insufficient hyperspectral elevation information.
Smart Images

Figure CN119942238A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning and relates to a local-global multi-resolution LiDAR-hyperspectral image collaborative classification method. Background Art
[0002] Hyperspectral images provide more information from hundreds of spectral bands and have been widely used in many fields, including urban construction, smart agriculture, and environmental monitoring. Among these applications, hyperspectral image classification is one of the most important and basic technologies. However, in hyperspectral images, since it only contains one-dimensional spectral and two-dimensional spatial attributes of land cover, its classification performance on land cover with high differences is limited. In terms of the application of three-dimensional information, laser radar (LiDAR) data with elevation information as the core has been widely used and can effectively guide HSI classification. With the help of LiDAR data, face recognition can achieve higher classification accuracy. Therefore, the information fusion and interaction between them has become an important research topic.
[0003] In recent years, in order to solve the problems of poor generalization ability and classification accuracy of traditional methods, deep learning has been introduced in the field of multimodal joint classification. Methods based on convolutional neural networks (CNNs) have been widely used. These methods can achieve reconstruction by learning a mapping model from two-dimensional measurements to three-dimensional HSI. Research based on convolutional neural networks has explored a variety of methods to capture deep features for the collaborative classification of hyperspectral-LiDAR data, and has made a series of progress. However, due to the limitation of the receptive field, convolutional neural networks cannot effectively combine global information. Therefore, researchers have tried to introduce multi-layer perceptrons (MLPs) to solve the problem of integrating global and local information. Unlike the local receptive field of CNN, the linear layer of MLP can perform global perception, and thanks to the absence of attention matrix multiplication, its computational overhead is less than that of the Transformer model. The introduction of MLP effectively improves the classification effect, but there are still some problems to be solved. First, existing MLP-based models tend to prioritize spatial information rather than spectral features, and only focus on local window areas, while ignoring global areas that help enhance feature representation. Second, the same surface cover exhibits different characteristics at different resolutions, which focuses on the detailed texture and edge information of different surface covers. At the same time, different surface covers have different scales, but the existing models ignore the characteristic information at different resolutions and scales. In the field of hyperspectral-LiDAR collaborative classification, how to effectively integrate local-global information, reasonably use information at different scales, and effectively use LiDAR data information to assist in enhancing classification results is the key to improving classification accuracy. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a local-global multi-resolution LiDAR-hyperspectral image collaborative classification method. The algorithm invents a multi-resolution and multi-scale feature extraction strategy to accurately obtain detailed information and features of different land covers at different scales. By designing a local-global oriented MLP structure, the local and global features of the hyperspectral image are effectively combined. By designing a feature enhancement module based on LiDAR images, the elevation information of the laser radar is introduced to improve the feature representation of the hyperspectral image. The semantic gap between hyperspectral and LiDAR data is narrowed by adopting a cross-attention mechanism.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A local-global multi-resolution LiDAR-hyperspectral image collaborative classification method comprises the following steps: S1: preprocessing of hyperspectral data and LiDAR data, and normalizing the hyperspectral data and LiDAR data by maximum-minimum threshold; S2: data set division; randomly selecting pixel points as training sets according to a certain ratio, and the remaining pixel points as test sets; S3: multi-resolution and multi-scale feature extraction; S4: local-global directed MLP feature learning; S5: feature fusion guided by LiDAR information; S6: loss calculation and model parameter update; S7: determining sample classification results.
[0007] Further, in step S3, multi-resolution and multi-scale feature extraction is performed on the hyperspectral data, including the following steps:
[0008] First, the hyperspectral image H1 h,w,c Perform two downsampling operations to obtain H2 and H3, where h is the length of the hyperspectral image, w is the width of the hyperspectral image, and c is the number of hyperspectral image channels. Specifically, downsampling uses a convolution layer with a convolution kernel size of 3*3 and a step size of 2 to downsample the hyperspectral image. The formula is as follows:
[0009]
[0010] After two downsamplings, convolutional layers, batch normalization layers, and activation layers are used to embed features on H1, H2, and H3, respectively, so that their dimensions become H h,w,c,j , j is the embedding dimension. Specifically, a 3D convolution layer with a convolution kernel size of (3,3,1) is used for convolution operation, and the activation layer uses the ReLu function. The formula is as follows:
[0011]
[0012] Among them, H ex Indicates H xAfter embedding, ReLu represents the activation function, and BatchNormed() represents the batch normalization function. Represents a 3D convolutional layer with a kernel size of (3,3,1).
[0013] Perform window-based multi-scale feature extraction. By selecting non-overlapping windows of size 3*3, 5*5, and 7*7, H e1 ,H e2 ,H e3 Specifically, H is divided by Rearrange operation. h,w,c Expand to H h / ws,w / ws,ws,ws,c In the form of, ws is a different window size, and then the local-global directed MLP is used to extract features and then the Concat operation is used to combine them. The formula is as follows:
[0014]
[0015] Among them, Reshape() represents the Reshape deformation operation, Concat() represents the splicing operation, and DMLP() represents the local-global directed MLP feature extraction module. Corresponding to H ex The features after y*y window division, O x Indicates the corresponding H x The output result is:
[0016] After extracting multi-scale features from H, H1, and H2 respectively, O2 and O3 are restored to the same scale as O1 through upsampling, and then spliced to obtain the output features of the hyperspectral information. Specifically, upsampling uses a transposed convolution with a convolution kernel size of 3*3, and the formula is as follows:
[0017]
[0018] Among them, H out Represents the output characteristics of hyperspectral data, Concat() represents the concatenation operation, Represents a transposed convolution with a convolution kernel size of 3*3.
[0019] Further, in step S4, feature learning is performed using a local-global directed MLP. First of all Perform shuffle operation internally to obtain the shuffled
[0020] Each H and Three data forms are obtained by expanding according to the horizontal, vertical and feature dimensions (H 1 ∈Rw*j*c,h ,H 2 ∈R h*j*c,w ,H 3 ∈R w*j*h,c ), then input it into the multi-layer perceptron MLP for feature learning, and then restore it to its original scale through the Reshape operation and add the output feature representation The formula is as follows:
[0021]
[0022] Among them, Reshape() represents the reshape deformation operation, and MLP() represents the multi-layer perceptron.
[0023] Further, in step S5, data feature fusion is performed based on the LiDAR information guidance, including the following steps:
[0024] S51, LiDAR data L h,w,d Perform preliminary feature extraction, where h, w of L are the same as H, and d represents the dimensional information of LiDAR data. First, embed L into K dimensions to transform it into L h,w,d,k Specifically, a 2D convolution layer with a convolution kernel size of (3,3), a batch normalization layer, and a ReLu activation layer are used for feature embedding. The formula is as follows:
[0025]
[0026] S52, for L embed Perform window-based multi-scale feature extraction. Select non-overlapping windows of size 3*3, 5*5, 7*7 for division. Then in Perform shuffle operation internally to obtain the shuffled After expanding the window in three dimensions, namely horizontal, vertical and feature, three data forms are obtained (L 1 ∈R w*k*d,h ,L 2 ∈R h*k*d,w ,L 3 ∈R w*k*h,d ), using a 2D convolution-based attention layer (ConvAttention, CA). Perform feature learning, then restore it to its original scale through the Reshape operation and add the output feature representation The formula is as follows:
[0027]
[0028] Among them, CA first converts the input features into query values, key values, and attribute values through three 2D convolution layers with a convolution kernel size of 1*1. The formula is as follows:
[0029]
[0030] Among them, L represents the input feature map, query, key, and value represent the obtained query value, key value, and attribute value. Then the query value, key value, and attribute value are combined according to the attention formula:
[0031]
[0032] Among them, softmax() represents the activation function that maps the value to the interval [0,1], Q, K, V represent the query value, key value, and attribute value respectively, and d represents the feature dimension;
[0033] S53, and They are mapped to the same feature dimension through linear layers, and feature refinement is performed using a parameter-sharing dilated convolution layer. Specifically, feature alignment is performed using a dilated convolution layer with a convolution kernel size of 5*5 and a dilated rate of d=1, 3, 5, a ReLu activation function layer, and a random loss layer with a Dropout rate of 0.1, and then concatenated using the Concat operation. The formula is as follows;
[0034]
[0035] Among them, DCB y () represents a dilated convolutional layer with a dilation rate of d = y. It represents the hyperspectral result after the hole convolution layer. It represents the LiDAR result after the hole convolution layer, and DropOut() represents the random inactivation layer.
[0036] S54, yes Apply three linear layers to generate corresponding key values, query values, and attribute values respectively. The formula is as follows:
[0037]
[0038] Among them, Q, K, V represent the query value, key value, and attribute value obtained, and W q ,W k ,W v ,W Q ,W K ,W V Represent the projection weights of query value, key value, and attribute value respectively.
[0039] Then during the attention operation, The corresponding key value and The corresponding key values are swapped to achieve feature alignment and fusion. The specific formula is as follows:
[0040]
[0041] Final h and final l Perform 2D maximum pooling to enhance the important features, then perform concatenation, and adjust the dimensions through a linear layer to get the final output. The specific formula is as follows;
[0042] final out =Linear(Concat(Maxpool2d(final h )+final h ,Maxpool2d(final l )+final l ))
[0043] Further, in step S6, the loss is calculated according to the loss function and back-propagated to update the model parameters. The loss function adopted by this model is a hybrid function based on the cross entropy function. and final out After projecting into the same feature space, the cross entropy loss function is applied respectively and added together. The specific formula is as follows:
[0044]
[0045] After the training is completed, a trained hyperspectral image classifier is obtained; the input sample is discriminated by the trained classifier to obtain the category of the sample.
[0046] The beneficial effects of the present invention are:
[0047] The local-global directed MLP structure proposed in the present invention can effectively combine local features and global features. The multi-resolution and multi-scale feature extraction strategy adopted can accurately obtain detailed information and features of different land covers at different scales. Secondly, a multi-resolution and multi-scale feature extraction strategy was invented to accurately obtain detailed information and features of different land covers at different scales. Subsequently, the problem of insufficient hyperspectral elevation information was alleviated by applying a feature enhancement module guided by LiDAR information. The semantic gap between hyperspectral information and LiDAR information was eliminated through multi-scale void convolution and cross-attention mechanism. The classification accuracy and effect of hyperspectral-LiDAR image collaborative classification are improved. The experimental results on three hyperspectral-LiDAR image datasets show that the performance of the proposed network is better than the current advanced hyperspectral-LiDAR image collaborative classification method.
[0048] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0050] Figure 1 is a flow chart of the method of the present invention;
[0051] Figure 2 Schematic diagram of the local-global directed MLP network (MMD-MLP) for hyperspectral-LiDAR collaborative classification;
[0052] Figure 3 It is a structural diagram of the local global directed MLP feature extraction module (LGDM) of the present invention;
[0053] Figure 4 It is a structural diagram of the LiDAR feature extraction module (LiFR) based on convolutional attention of the present invention;
[0054] Figure 5 is a structural diagram of a cross-attention-based feature enhancement module (CAFE) of the present invention;
[0055] Figure 6 is the Muufl dataset, where (a) is a pseudo-color image, (b) is a binary LiDAR image, and (c) is a true value image.
[0056] Figure 7Visualization results of different methods on the Muufl dataset, including (a) CoupledCNN, (b) EndNet, (c) FusAtNet, (d) Fusion-HCT, (e) MDL, (f) MTNet, (g) SENet (h) True value map (i) Result map of this method. DETAILED DESCRIPTION
[0057] The technical solution of the present invention is described in detail below with reference to the accompanying drawings.
[0058] Figure 1 It is a flow chart of the method of the present invention, such as Figure 1 As shown, a local-global multi-resolution LiDAR-hyperspectral image collaborative classification method includes the following steps:
[0059] 1. Preprocessing of hyperspectral data and LiDAR data, and normalization of maximum-minimum threshold of hyperspectral data and LiDAR data;
[0060] 2. Data set division: randomly select pixels as training set according to a certain ratio, and the remaining pixels as test set;
[0061] 3. Perform multi-resolution and multi-scale feature extraction on hyperspectral data. First, extract the hyperspectral image H1 h,w,c Perform two downsampling operations to obtain H2 and H3, where h is the length of the hyperspectral image, w is the width of the hyperspectral image, and c is the number of hyperspectral image channels. Specifically, downsampling uses a convolution layer with a convolution kernel size of 3*3 and a step size of 2 to downsample the hyperspectral image. The formula is as follows:
[0062]
[0063] After two downsamplings, convolutional layers, batch normalization layers, and activation layers are used to embed features on H1, H2, and H3, respectively, so that their dimensions become H h,w,c,j , j is the embedding dimension. Specifically, a 3D convolution layer with a convolution kernel size of (3,3,1) is used for convolution operation, and the activation layer uses the ReLu function. The formula is as follows:
[0064]
[0065] Among them, H ex Represents the embedded features, ReLu represents the activation function, and BatchNormed() represents the batch normalization function. Represents a 3D convolutional layer with a kernel size of (3,3,1).
[0066] Perform window-based multi-scale feature extraction. By selecting non-overlapping windows of size 3*3, 5*5, and 7*7, H e1 ,H e2 ,H e3 Specifically, H is divided by Rearrange operation. h,w,c Expand to H h / ws,w / ws,ws,ws,c In the form of, ws is a different window size, and then the local-global directed MLP is used to extract features and then the Concat operation is used to combine them. The formula is as follows:
[0067]
[0068] Among them, Reshape() represents the Reshape deformation operation, Concat() represents the splicing operation, and DMLP() represents the local-global directed MLP feature extraction module. Corresponding to H ex The features after y*y window division, O x Indicates the corresponding H x The output result of
[0069] After extracting multi-scale features from H, H1, and H2 respectively, O2 and O3 are restored to the same scale as O1 through upsampling, and then spliced to obtain the output features of the hyperspectral information. Specifically, upsampling uses a transposed convolution with a convolution kernel size of 3*3, and the formula is as follows:
[0070]
[0071] Among them, H out Represents the output characteristics of hyperspectral data, Concat() represents the concatenation operation, Represents a transposed convolution with a convolution kernel size of 3*3.
[0072] 3. Use local-global directed MLP for feature learning. For the multi-resolution and multi-scale feature information of the input First of all Perform shuffle operation internally to obtain the shuffled
[0073] Each H and Three data forms are obtained by expanding according to the horizontal, vertical and feature dimensions (H 1 ∈R w*j*c,h ,H 2 ∈R h*j*c,w ,H 3 ∈R w*j*h,c), then input it into the multi-layer perceptron MLP for feature learning, and then restore it to its original scale through the Reshape operation and add the output feature representation The formula is as follows:
[0074]
[0075] Among them, Reshape() represents the reshape deformation operation, and MLP() represents the multi-layer perceptron.
[0076] 5. Perform data feature fusion based on LiDAR information guidance, including the following steps: S51, h,w,d Perform preliminary feature extraction, where h, w of L are the same as H, and d represents the dimensional information of LiDAR data. First, embed L into K dimensions to transform it into L h,w,d,k Specifically, a 2D convolution layer with a convolution kernel size of (3,3), a batch normalization layer, and a ReLu activation layer are used for feature embedding. The formula is as follows:
[0077]
[0078] S52, for L embed Perform window-based multi-scale feature extraction. Select non-overlapping windows of size 3*3, 5*5, 7*7 for division. Then in Perform shuffle operation internally to obtain the shuffled After expanding the window in three dimensions, namely horizontal, vertical and feature, three data forms are obtained (L 1 ∈R w*k*d,h ,L 2 ∈R h*k*d,w ,L 3 ∈R w*k*h,d ), using a 2D convolution-based attention layer (ConvAttention, CA). Perform feature learning, then restore it to its original scale through the Reshape operation and add the output feature representation The formula is as follows:
[0079]
[0080] Among them, CA first converts the input features into query values, key values, and attribute values through three 2D convolution layers with a convolution kernel size of 1*1. The formula is as follows:
[0081]
[0082] Among them, L represents the input feature map, query, key, and value represent the obtained query value, key value, and attribute value. Then the query value, key value, and attribute value are combined according to the attention formula:
[0083]
[0084] Among them, softmax() represents the activation function that maps the value to the interval [0,1], Q, K, V represent the query value, key value, and attribute value respectively, and d represents the feature dimension;
[0085] S53, and They are mapped to the same feature dimension through linear layers, and feature refinement is performed using parameter-sharing dilated convolution layers. Specifically, feature alignment is performed using dilated convolution layers with dilated rates d=1, 3, 5 and a kernel size of 5*5, ReLu activation function layers, and random layers with a Dropout rate of 0.1, and then concatenated using the Concat operation. The formula is as follows;
[0086]
[0087] Among them, DCB y () represents a dilated convolutional layer with a dilation rate of d = y. It represents the hyperspectral result after the hole convolution layer. It represents the LiDAR result after the hole convolution layer, and DropOut() represents the random inactivation layer.
[0088] S54, yes Apply three linear layers to generate corresponding key values, query values, and attribute values respectively. The formula is as follows:
[0089]
[0090] Among them, Q, K, V represent the query value, key value, and attribute value obtained, and W q ,W k ,W v ,W Q ,W K ,W V Represent the projection weights of query value, key value, and attribute value respectively.
[0091] Then during the attention operation, The corresponding key value and The corresponding key values are swapped to achieve feature alignment and fusion. The specific formula is as follows:
[0092]
[0093] Finalh and final l Perform 2D maximum pooling to enhance the important features, then perform concatenation, and adjust the dimensions through a linear layer to get the final output. The specific formula is as follows;
[0094] final out =Linear(Concat(Maxpool2d(final h )+final h ,Maxpool2d(final l )+final l ))
[0095] 6. Calculate the loss according to the loss function and back propagate to update the model parameters. The loss function adopted by this model is a hybrid function based on the cross entropy function. and final out After projecting into the same feature space, the cross entropy loss function is applied respectively and added together. The specific formula is as follows:
[0096]
[0097] 7. After the training is completed, a trained hyperspectral image classifier is obtained; the input sample is discriminated by the trained classifier to obtain the category of the sample.
[0098] like Figure 6 This is the experimental result of the method of the present invention on an open source hyperspectral-LiDAR multimodal dataset Muufl. It can be seen that the types of surface objects are well identified. The classification effect of the present invention can be further illustrated by comparative experiments. The method of the present invention and other existing methods CoupledCNN, EndNet, FusAtNet, Fusion-HCT, MDL, MTNet, and S2ENet are compared on the Muufl dataset to calculate the overall accuracy OverallAccuracy, the average accuracy Avage Accuracy, and the Kappa coefficient Kappa. The higher the value, the better the effect. Table 1 shows the accuracy comparison of different methods:
[0099] Table 1 Comparison of various methods on the Muufl dataset
[0100]
[0101] Figure 7 Several deep learning-based methods are given to visualize the detection results on a scene. It can be seen that the performance of the method described in the present invention is better than other hyperspectral-LiDAR collaborative classification methods.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A local-global multi-resolution LiDAR-hyperspectral image collaborative classification method, characterized in that: The following steps are involved: S1: Hyperspectral data and LiDAR data preprocessing; S2: Dataset division; S3: multi-resolution and multi-scale feature extraction; S4: Local-global MLP feature learning; S5: Feature fusion guided by LiDAR information; S6: loss calculation and model parameter update; S7: Determine the sample classification result.
2. The local-global multi-resolution LiDAR-hyperspectral image collaborative classification method according to claim 1, characterized in that: In step S1, the hyperspectral data and LiDAR data are normalized by maximum-minimum threshold.
3. The local-global multi-resolution LiDAR-hyperspectral image collaborative classification method according to claim 2, characterized in that: In step S2, pixels are randomly selected according to a certain ratio as a training set, and the remaining pixels are used as a test set.
4. The dual-stage multi-scale hyperspectral snapshot compression imaging image reconstruction method according to claim 3 is characterized in that: In step S3, multi-resolution and multi-scale feature extraction is performed on the hyperspectral data, including the following steps: First, the hyperspectral image H1 h,w,c Perform two downsampling operations to obtain H2 and H3, where h is the length of the hyperspectral image, w is the width of the hyperspectral image, and c is the number of hyperspectral image channels; specifically, downsampling uses a convolution layer with a convolution kernel size of 3*3 and a step size of 2 to downsample the hyperspectral image. The formula is as follows: After two downsamplings, convolutional layers, batch normalization layers, and activation layers are used to embed features on H1, H2, and H3, respectively, so that their dimensions become H h , w,c,j , j is the embedding dimension. Specifically, a 3D convolution layer with a convolution kernel size of (3,3,1) is used for convolution operation, and the activation layer uses the ReLu function. The formula is as follows: Among them, H ex Represents the embedded features, ReLu represents the activation function, and BatchNormed() represents the batch normalization function. Represents a 3D convolutional layer with a kernel size of (3,3,1). Perform window-based multi-scale feature extraction by selecting non-overlapping windows of size 3*3, 5*5, and 7*7 for H e1 ,H e2 ,H e3 Specifically, H is divided by Rearrange operation. h,w,c Expand to H h / ws,w / ws,ws,ws,c In the form of ws, the window size is different. The local-global MLP is used to extract features and then the Concat operation is used to combine them. The formula is as follows: Among them, Reshape() represents the Reshape deformation operation, Concat() represents the splicing operation, and DMLP() represents the local-global directed MLP feature extraction. Corresponding to H ex The features after y*y window division, O x Indicates the corresponding H x The output result; after multi-scale feature extraction of H, H1, and H2, O2 and O3 are restored to the same scale of O1 through upsampling operation, and then spliced to obtain the output features of hyperspectral information. Specifically, upsampling adopts transposed convolution with a convolution kernel size of 3*3, and the formula is as follows: Among them, H out Represents the output characteristics of hyperspectral data, Concat() represents the concatenation operation, Represents a transposed convolution with a convolution kernel size of 3*3.
5. The local-global multi-resolution LiDAR-hyperspectral image collaborative classification method according to claim 4, characterized in that: In step S4, a local-global directed MLP is used for feature learning. First of all Perform shuffle operation internally to obtain the shuffled Each H and Three data forms are obtained by expanding according to the horizontal, vertical and feature dimensions (H 1 ∈R w *j*c,h ,H 2 ∈R h*j*c,w ,H 3 ∈R w*j*h,c ), then input it into the multi-layer perceptron MLP for feature learning, and then restore it to its original scale through the Reshape operation and add the output feature representation The formula is as follows: Among them, Reshape() represents the reshape deformation operation, and MLP() represents the multi-layer perceptron.
6. The local-global multi-resolution LiDAR-hyperspectral image collaborative classification method according to claim 4, characterized in that: In step S5, data feature fusion is performed based on the LiDAR information guidance, including the following steps: S51, LiDAR data L h,w,d Perform preliminary feature extraction, where h, w of L are the same as H, and d represents the dimensional information of LiDAR data. First, embed L into K dimensions to transform it into L h,w,d,k Specifically, a 2D convolution layer with a convolution kernel size of (3,3), a batch normalization layer, and a ReLu activation layer are used for feature embedding. The formula is as follows: S52, for L embed Perform window-based multi-scale feature extraction by selecting non-overlapping windows of size 3*3, 5*5, and 7*7 for division Then in Perform shuffle operation internally to obtain the shuffled After expanding the window in three dimensions, namely horizontal, vertical and feature, three data forms are obtained (L 1 ∈R w *k*d,h ,L 2 ∈R h*k*d,w ,L 3 ∈R w*k*h,d ), using a 2D convolution-based attention layer (ConvAttention, CA) for L and Perform feature learning, then restore it to its original scale through the Reshape operation and add the output feature representation The formula is as follows: Among them, CA first converts the input features into query values, key values, and attribute values through three 2D convolution layers with a convolution kernel size of 1*1. The formula is as follows: Among them, L represents the input feature map, query, key, and value represent the obtained query value, key value, and attribute value, and then the query value, key value, and attribute value are combined according to the attention formula: Among them, softmax() represents the activation function that maps the value to the interval [0,1], Q, K, V represent the query value, key value, and attribute value respectively, and d represents the feature dimension; S53, and They are mapped to the same feature dimension through linear layers respectively, and feature refinement is performed using parameter-sharing dilated convolution layers. Specifically, feature alignment is performed using dilated convolution layers with convolution kernels of 5*5 and dilated rates d=1, 3, 5, ReLu activation function layers, and random layers with a Dropout rate of 0.1, and then concatenated using the Concat operation. The formula is shown below. Among them, DCB y () represents a dilated convolutional layer with a dilation rate of d = y. It represents the hyperspectral result after the hole convolution layer. Represents the LiDAR result after the hole convolution layer, and DropOut() represents the random dropout layer; S54, yes Apply three linear layers to generate corresponding key values, query values, and attribute values respectively. The formula is as follows: Among them, Q, K, V represent the query value, key value, and attribute value obtained, and W q ,W k ,W v ,W Q ,W K ,W V Represent the projection weights of query value, key value, and attribute value respectively. Then during the attention operation, The corresponding key value and The corresponding key values are swapped to achieve feature alignment and fusion. The specific formula is as follows: Final h and final l Perform a 2D maximum pooling operation to enhance the important features, then perform a splicing operation, and adjust the dimensions through a linear layer to obtain the final output result. The specific formula is as follows; final out =Linear(Concat(Maxpool2d(final h )+final h ,Maxpool2d(final l )+fina l ))。 7. The local-global multi-resolution LiDAR-hyperspectral image collaborative classification method according to claim 6, characterized in that: In step S6, the loss is calculated and back-propagated according to the loss function to update the model parameters. The loss function adopted by this model is a hybrid function based on the cross entropy function. and final out After projecting into the same feature space, the cross entropy loss function is applied respectively and added together. The specific formula is as follows:
8. The local-global multi-resolution LiDAR-hyperspectral image collaborative classification method according to claim 7, characterized in that: After the training is completed, a trained hyperspectral image classifier is obtained; the input sample is discriminated by the trained classifier to obtain the category of the sample.
Citation Information
Patent Citations
Hyperspectral and laser radar classification method for loop generation learning based on modal attention
CN117893827A
Hyperspectral and laser radar multilayer fusion classification method based on adversarial learning
CN117934978A
Two-stage multi-scale hyperspectral snapshot compression imaging image reconstruction method
CN117974909A