A hyperspectral and laser radar multi-modal remote sensing data classification method and system

By combining the Cross Feature Reconstruction (CFR) module with CNN, the problem of insufficient feature representation capability in the classification of hyperspectral and lidar multimodal remote sensing data is solved, and more accurate object recognition and classification results are achieved.

CN114387505BActive Publication Date: 2025-11-25SHANDONG NORMAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111507974.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-11-25
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

Existing classification methods for hyperspectral and lidar multimodal remote sensing data have limited feature representation capabilities when processing heterogeneous data, and deep network training is prone to gradient vanishing problems.

Method used

By employing a cross-feature reconstruction module (CFR) combined with a convolutional neural network (CNN), and through cross-channel prediction and an encoder-decoder architecture, the hyperspectral and lidar image features are tightly integrated. The cross-feature reconstruction module is used to exchange and fuse information at the feature level.

Benefits of technology

It improves the classification accuracy of multimodal remote sensing data, generates clearer and more detailed feature maps, and achieves more accurate object recognition and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387505B_ABST
    Figure CN114387505B_ABST
Patent Text Reader

Abstract

The application belongs to the field of image processing, and provides a hyperspectral and laser radar multi-modal remote sensing data classification method and system. The method comprises the following steps: acquiring a hyperspectral image and a laser radar image; extracting a feature map of the hyperspectral image and a feature map of the laser radar image by using a CNN; obtaining a fused multi-modal feature by using a cross feature reconstruction module based on the feature map of the hyperspectral image and the feature map of the laser radar image; and obtaining a classification result by using a classifier based on the multi-modal feature.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of image processing, and particularly relates to a hyperspectral and laser radar multi-modal remote sensing data classification method and system. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] Hyperspectral remote sensing image refers to an image obtained by a hyperspectral imager, which has very rich spatial information and spectral information. In addition, the hyperspectral image also has a larger number of wavebands and extremely high resolution, so that detailed ground feature characteristics can be obtained by analyzing the spectral characteristics and spatial characteristics. Laser radar uses a laser as a light source and is an active remote sensing device using photoelectric detection technology. Laser radar is an advanced detection method combining laser technology and modern photoelectric detection technology. Compared with ordinary microwave radar, laser radar uses a laser beam, and the working frequency is much higher than that of microwave, thus bringing many advantages, mainly including: (1) high resolution (2) good concealment and strong anti-active interference capability (3) good low-altitude detection performance (4) small size and light weight, and has great advantages in mapping the 3D contour of an object. At present, hyperspectral and laser radar imaging technology has been widely applied, including precision agriculture, atmospheric monitoring, ocean detection and other fields. With the wide application of hyperspectral and laser radar imaging technology in more and more fields, how to more quickly and accurately identify objects and classify substances has become a problem to be solved.

[0004] For the multi-modal remote sensing data classification task of hyperspectral image and laser radar image, morphological profile and subspace learning based method are two main methods for multi-modal remote sensing data feature extraction and classification, although these traditional shallow models have achieved satisfactory performance, but due to the large gap between different modes, the feature representation ability is still limited. Inspired by the recent success of various deep networks in extracting more discriminative features from data, such as convolutional neural network (CNN), recurrent neural network (RNN), graph convolution network (GCN) and deep similarity network, etc., some preliminary multi-modal networks have been proposed and used for remote sensing images, and better classification results have been obtained than using single mode. However, the fusion strategy is the key factor to determine the performance of multi-modal network, which can be roughly divided into two groups: cascade-based fusion and alignment-based fusion. Generally speaking, the former directly stacks the images or features at the early, middle or late stage, resulting in early fusion, middle fusion and late fusion, respectively, while the latter, as the name implies, aligns different modes by similarity measurement or constraint, and effectively realizes the fusion process. These newly developed methods have been proven to be effective methods for fusing multiple remote sensing data sources, but the ability of these existing methods in dealing with multi-modal data, especially heterogeneous data (such as HS and SAR data, HS and laser radar), is still limited. This may be because there is a lack of more advanced fusion strategies to better eliminate the gap between different modes and obtain features.

[0005] In the process of implementing the present disclosure, the inventors found that the prior art has the following technical problems

[0006] There is a large gap between different data modes, and the feature representation using morphological profile and subspace learning based method is still limited.

[0007] (2) The commonly used direct stacking of images or features at the early, middle or late stage, these existing methods still have limited ability in dealing with multi-modal data, especially heterogeneous data (such as HS and SAR data, HS and laser radar).

[0008] (3) Due to the limited training samples of hyperspectral data set, with the increase of network depth, the phenomenon of gradient disappearance may occur. SUMMARY

[0009] In order to solve the technical problems existing in the background art, the present application provides a hyperspectral and laser radar multi-modal remote sensing data classification method and system, which obtains more accurate image classification results, which is helpful for subsequent land research, forestry monitoring, etc., and has certain practicality.

[0010] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows:

[0011] The first aspect of the present application provides a hyperspectral and lidar multi-modal remote sensing data classification method.

[0012] The hyperspectral and lidar multi-modal remote sensing data classification method comprises the following steps.

[0013] Obtaining a hyperspectral image and a lidar image;

[0014] Extracting a feature map of the hyperspectral image and a feature map of the lidar image respectively by using a CNN;

[0015] Based on the feature map of the hyperspectral image and the feature map of the lidar image, a cross-feature reconstruction module is used to obtain a fused multi-modal feature;

[0016] Based on the multi-modal feature, a classifier is used to obtain a classification result.

[0017] The second aspect of the present application provides a hyperspectral and lidar multi-modal remote sensing data classification system.

[0018] The hyperspectral and lidar multi-modal remote sensing data classification system comprises the following steps.

[0019] An obtaining module configured to obtain a hyperspectral image and a lidar image;

[0020] A feature extraction module configured to extract a feature map of the hyperspectral image and a feature map of the lidar image respectively by using a CNN;

[0021] A fusion module configured to obtain a fused multi-modal feature based on the feature map of the hyperspectral image and the feature map of the lidar image by using a cross-feature reconstruction module;

[0022] A classification module configured to obtain a classification result based on the multi-modal feature by using a classifier.

[0023] The third aspect of the present application provides a computer readable storage medium.

[0024] A computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the steps of the hyperspectral and lidar multi-modal remote sensing data classification method according to the first aspect.

[0025] The fourth aspect of the present application provides a computer device.

[0026] A computer device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, the processor executing the program to implement the steps of the hyperspectral and lidar multi-modal remote sensing data classification method according to the first aspect.

[0027] Compared with the prior art, the present application has the beneficial effects that:

[0028] The present application adopts convolutional neural network for feature extraction, which can accurately extract feature images.

[0029] The present application adopts the image fusion mode of hyperspectral and laser radar, which can fully combine the advantages of both, can more accurately identify objects, and thus more accurately realize classification.

[0030] The present application adopts a new encoder-decoder network, namely cross feature reconstruction model (CFR), which can more closely fuse features than traditional early fusion, mid-fusion and late fusion. The cross reconstruction mode in the CFR module can realize effective information exchange and more compact fusion at the feature level, thereby improving classification accuracy and making the classification result more accurate and clear. BRIEF DESCRIPTION OF DRAWINGS

[0031] The drawings constituting a part of the specification of the present application are used to provide further understanding of the present application, and the illustrative embodiments of the present application and the description thereof are used to explain the present application, and do not constitute undue limitation on the present application.

[0032] Figure 1 is the network architecture diagram of the hyperspectral and laser radar multi-modal remote sensing data classification method shown by the present application;

[0033] Figure 2 is the network framework diagram of the image feature map extraction by the CNN shown by the present application;

[0034] Figure 3 demonstrates the difference comparison diagram between early fusion, mid-fusion, late fusion and CFR fusion;

[0035] Figure 4 is the structure diagram of the CFR module shown by the present application;

[0036] Figure 5(a) is a visualization visual of feature map when the CFR module is not used according to the present application Figure 1 ;

[0037] Figure 5(b) is a visualization visual of feature map when the CFR module is used according to the present application Figure 1 ;

[0038] Figure 5(c) is a visualization visual of feature map when the CFR module is not used according to the present application Figure 2 ;

[0039] Figure 5(d) is a visualization visual of feature map when the CFR module is used according to the present application Figure 2 ;

[0040] Figure 6 is the ablation experiment result comparison chart of the fusion strategy shown in the present application;

[0041] Figure 7 is the effect comparison chart of the evaluation effect shown in the present application and the current advanced classification model;

[0042] Figure 8 is the flow chart of the hyperspectral and laser radar multi-modal remote sensing data classification method shown in the present application. DETAILED DESCRIPTION

[0043] The present application will be further described below in conjunction with the drawings and embodiments.

[0044] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as would be commonly understood by one of ordinary skill in the art to which the present application belongs.

[0045] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and it should also be understood that when the terms "comprise" and / or "include" are used in the specification, they indicate the presence of the features, steps, operations, devices, components and / or combinations thereof.

[0046] It should be noted that the flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the method and system according to various embodiments of the present disclosure. It should be noted that each block in the flowchart or block diagram can represent a module, a program segment or a part of code, which can include one or more executable instructions for implementing the logical functions defined in various embodiments. It should also be noted that in some alternative implementations, the functions indicated in the blocks can also occur in an order different from that indicated in the drawings. For example, two blocks indicated in succession can actually be executed substantially in parallel, or they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.

[0047] Embodiment one

[0048] As Figure 8As shown, the embodiment provides a hyperspectral and laser radar multi-modal remote sensing data classification method. The method is applied to a server for illustration. It can be understood that the method can also be applied to a terminal and can be applied to a system including a terminal and a server and is realized through interaction of the terminal and the server. The server can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communication, middleware services, domain name services, security services CDN, and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, and the like, but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application. In the embodiment, the method includes the following steps:

[0049] Obtaining a hyperspectral image and a laser radar image;

[0050] Extracting a feature map of the hyperspectral image and a feature map of the laser radar image by using a CNN;

[0051] Based on the feature map of the hyperspectral image and the feature map of the laser radar image, a cross-feature reconstruction module is used to obtain a fused multi-modal feature;

[0052] Based on the multi-modal feature, a classifier is used to obtain a classification result.

[0053] The overall architecture of the hyperspectral and laser radar multi-modal remote sensing data classification method based on the convolutional neural network provided in the embodiment is as shown in Figure 1 The cross-feature reconstruction (CFR) module is as shown in Figure 4 Figure 7 The evaluation effect of the embodiment and the comparison effect diagram of the current advanced classification model are shown. The specific steps are as follows:

[0054] As shown in Figure 1 First, the hyperspectral image and the laser radar image are simultaneously sent into the CNN for image feature map extraction. The double-path CNN provides a feature extraction module for each modality, as shown in Figure 2 ​As shown; because the input layer of the convolutional neural network can process multi-dimensional data, and the network is widely used in the field of computer vision, many studies assume three-dimensional input data when introducing the structure of the network, that is, two-dimensional pixels and RGB channels on a plane. In recent years, CNN has been proven to be an effective tool in extracting discriminative feature representations of remote sensing data. In this embodiment, the extractor consists of two CNN blocks in each modality stream, each block having a 3x3 convolutional layer, a batch normalization (BN) layer, a 2x2 max pooling layer, and a ReLU activation layer.

[0055] A Cross-Feature Reconstruction (CFR) module is constructed. Previous multi-modal feature fusion methods usually follow a cascade-based fusion strategy, such as early fusion, mid-fusion, late fusion, etc. The encoder-decoder architecture is a relatively new fusion method, and a more compact feature representation can be generated through the encoder-decoder architecture. In addition, the Cross-Feature Reconstruction (CFR) module is constructed by cross-channel prediction following the encoder architecture, so that the features extracted from different modalities are fused in a more sufficient manner. More specifically, this cross-reconstruction pattern in the CFR module can achieve effective information exchange and more compact fusion at the feature level. Figure 3 The differences between early fusion, mid-fusion, late fusion, and the cross-fusion strategy proposed in this embodiment are shown. The fusion process in the CFR module can be expressed as

[0056]

[0057] wherein represents the pixel-level fused feature in the l-th layer. The gen function is defined as an encoder network in the CFR, represents the output of the p-th layer of the CNN, represents the weights of the p-th layer CNN in the encoder and the weights of the l-th layer in the CNN, represents the offset of the l-th layer in the CNN; l represents the layer number of the CNN; p represents the number of layers of the CNN; m represents an infinite quantity, which essentially represents a number.

[0058] And the decoder part can be expressed as:

[0059]

[0060] wherein represents the pixel-level reconstructed feature in the l-th layer, V i (m) represents the pixel-level fused feature in the m-th layer; This represents the weights of the l-th layer in the CNN within the decoder; The offset in the decoder is represented by l; l represents the CNN layer number; m represents an infinite quantity, essentially referring to a number; in order to perform the cross-feature reconstruction module, features are learned from the input features. To output reconstruct features Network mapping. CFR modules, such as... Figure 4 As shown.

[0061] The network parameters in the Cross-Feature Reconstruction (CFR) module mentioned in this embodiment can be updated by jointly optimizing the following overall loss function:

[0062] L = L l +αL rec

[0063] In our example, the parameter α was determined to be 1 through experimentation and experience, resulting in relatively stable performance. l It is the fusion feature V i (m) And a label Y i The cross-entropy loss between them can be expressed as:

[0064]

[0065] Where L rec The loss function for CFR is:

[0066]

[0067] in L represents the output reconstructed features. rec Reconstructed features are calculated using L2-norm. and cross-channel characteristics The loss regularization norm between them.

[0068] In addition to studying the quantitative classification accuracy, the embodiment also verifies the effectiveness of the CFR module in fusing multi-modal features from the visual perspective of visualization. The feature maps with and without the CFR module in the region of interest of the hyperspectral image are visualized as shown in Figs. 5(a)-5(d). The edges and structures of the feature maps generated by the CFR module are clearer than those without the CFR module, resulting in more detailed and fine fusion feature maps. In addition, we also conducted an ablation experiment on the fusion strategy. Early fusion, mid-fusion and late fusion are three main fusion strategies in multi-modal fusion networks, and the CFR model is a derivative model of mid-fusion. In order to verify the effectiveness of the CFR module, early fusion, mid-fusion, late fusion and CFR fusion are quantitatively compared on the HS LiDAR Houston 2013 dataset, and the results are shown in Fig. 6. The performance of early fusion is poor because multi-modal remote sensing data coupling occurs in the initial stage, while the late fusion strategy focuses on the decision layer, to some extent, ignoring the information exchange at the feature layer. As a trade-off, mid-fusion tends to achieve better fusion results. Figure 6

[0069] The embodiment conducts a large number of experiments on the HS LiDAR Houston 2013 dataset. The results obtained by the model are compared with the Ground Truth of the dataset, and the classification results are evaluated by three evaluation indexes (overall accuracy), (mean accuracy), coefficient, and the evaluation effect is compared with that of the current advanced classification model, as shown in Fig. 7, which shows that the classification accuracy of the embodiment is higher and the classification effect is better, and has a certain practicality. Figure 7

[0070] Embodiment Two

[0071] The embodiment provides a hyperspectral and laser radar multi-modal remote sensing data classification system.

[0072] A hyperspectral and laser radar multi-modal remote sensing data classification system comprises:

[0073] An acquisition module configured to acquire a hyperspectral image and a laser radar image;

[0074] A feature extraction module configured to extract a feature map of the hyperspectral image and a feature map of the laser radar image respectively by using a cross-feature reconstruction module;

[0075] A fusion module configured to obtain fused multi-modal features based on the feature map of the hyperspectral image and the feature map of the laser radar image by using the cross-feature reconstruction module;

[0076] ​​a classification module configured to: based on the multi-modal feature, employ a classifier to obtain a classification result.

[0077] It should be noted that the above-mentioned acquisition module, feature extraction module, fusion module and classification module have the same examples and application scenarios as the steps in Embodiment One, but are not limited to the content disclosed in Embodiment One. It should be noted that the above-mentioned modules as part of the system can be executed in a computer system such as a set of computer executable instructions.

[0078] Embodiment Three

[0079] The embodiment provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the steps in the hyperspectral and laser radar multi-modal remote sensing data classification method according to the above-mentioned Embodiment One.

[0080] Embodiment Four

[0081] The embodiment provides a computer device, which includes a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps in the hyperspectral and laser radar multi-modal remote sensing data classification method according to the above-mentioned Embodiment One when executing the program.

[0082] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can be in the form of a hardware embodiment, a software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer usable program code.

[0083] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general purpose computer, a special purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that carries out the functions specified in one or more flows and / or blocks.

[0084] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 of the flow or flows and / or blocks Figure 1 of the block or blocks specified in the flow.

[0085] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 of the flow or flows and / or blocks Figure 1 of the block or blocks specified in the flow.

[0086] Those of ordinary skill in the art can understand that all or part of the flow of the above-mentioned embodiment method can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the flow of the above-mentioned embodiment method. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like.

[0087] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A hyperspectral and lidar multimodal remote sensing data classification method, characterized in that, The method comprises the following steps: obtaining a hyperspectral image and a lidar image; extracting a feature map of the hyperspectral image and a feature map of the lidar image respectively by using a CNN; obtaining a fused multi-modal feature based on the feature map of the hyperspectral image and the feature map of the lidar image by using a cross-feature reconstruction module; obtaining a classification result based on the multi-modal feature by using a classifier; optimizing network parameters of the cross-feature reconstruction module by using a loss function; the loss function is: wherein represents the output reconstructed feature, L rec is a reconstruction feature computed by L2-norm and cross-channel feature between the loss regularized norm.

2. The hyperspectral and lidar multimodal remote sensing data classification method according to claim 1, characterized in that, features are extracted by using a double-path CNN, and each path of the CNN comprises a convolution layer, a batch normalization layer, a pooling layer and a ReLU activation layer.

3. The hyperspectral and lidar multimodal remote sensing data classification method of claim 1, wherein, The cross-feature reconstruction module comprises an encoder module and a decoder module.

4. The hyperspectral and lidar multimodal remote sensing data classification method according to claim 3, characterized in that, The encoder module is configured to: wherein, represents the pixel-level fusion feature in the lth layer, the gen function is defined as an encoder network in the CFR, represents the output of the pth layer of the CNN, represents the weights of the pth layer CNN in the encoder and the weights of the lth layer in the CNN, represents the offset of the lth layer in the CNN in the encoder; l represents the layer number of the CNN; p represents the number of layers of the CNN; m represents an infinite quantity, which essentially refers to a number.

5. The hyperspectral and lidar multimodal remote sensing data classification method according to claim 4, characterized in that, The decoder module is configured to: wherein represents the pixel-level reconstructed feature in the lth layer, V i (m) represents the pixel-level fused feature in the mth layer; represents the weight of the lth layer in the CNN in the decoder; represents the offset in the decoder; l represents the layer number of the CNN; m represents an infinite quantity, which essentially represents a number; in order to perform the cross-feature reconstruction module, it is necessary to learn the mapping from the input feature to the output reconstructed feature network.

6. A hyperspectral and lidar multimodal remote sensing data classification system, characterized in that, The method comprises the following steps: an acquisition module configured to obtain a hyperspectral image and a lidar image; a feature extraction module configured to extract a feature map of the hyperspectral image and a feature map of the lidar image respectively by using a CNN; a fusion module configured to obtain a fused multi-modal feature based on the feature map of the hyperspectral image and the feature map of the lidar image by using a cross-feature reconstruction module; a classification module configured to obtain a classification result based on the multi-modal feature by using a classifier; optimizing network parameters of the cross-feature reconstruction module by using a loss function; the loss function is: wherein represents the output reconstructed feature, L rec is a reconstruction feature computed by L2-norm and cross-channel features between the loss regularized norm.

7. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by a processor to implement the steps in the hyperspectral and lidar multi-modal remote sensing data classification method according to any one of claims 1-5.

8. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the steps in the hyperspectral and lidar multi-modal remote sensing data classification method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Classification method based on fusion of hyperspectral image and DSM data

    CN110210420A

  • Coastal wetland deep learning classification method and device, equipment and storage medium

    CN111898662A

  • Multi-source image combined urban land surface coverage classification method

    CN113435253A