Hyperspectral and LiDAR 4D anti-shake fusion method and system based on deep learning
Through the hyperspectral camera and lidar fusion network based on deep learning, the external parameter relationship is calculated in real time, and the problem of degradation of the fusion accuracy of hyperspectral images and lidar 4D information in the existing technology is solved, achieving efficient and accurate 4D fusion in a jitter environment.
Patent Information
- Application Number
- CN202210394912.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-15
AI Technical Summary
The existing 4D information fusion technology of hyperspectral images and lidar relies on complicated calibration steps and hard calibration, resulting in a decrease in fusion accuracy and inefficiency in a long-term jitter environment.
Using a hyperspectral camera and lidar fusion network (HLFN) based on deep learning, the external parameters relationship between the hyperspectral camera and lidar is calculated in real time through feature extraction and fusion, and a prediction transformation matrix is generated to achieve 4D anti-shake fusion.
The 4D fusion information of hyperspectral morphology is quickly and accurately collected under the mechanical jitter conditions of the equipment, which improves the fusion accuracy and efficiency, reduces the dependence on the calibration plate, and can work stably in a long-term jitter environment.
Smart Images

Figure CN114882329B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of hyperspectral and laser radar data fusion, and in particular to a hyperspectral and laser radar 4D anti-shake fusion method and system based on deep learning. Background Art
[0002] In intelligent automation equipment, in order to overcome the information dimension perception defects of a single sensor, more and more applications require multi-sensor fusion technology.
[0003] In intelligent driving and intelligent robotics, the data fusion technology of hyperspectral cameras and lidars has attracted much attention. On the one hand, hyperspectral cameras can simultaneously collect spectral information of multiple bands of a target, and their rich spectral features lay a solid foundation for target recognition, semantic segmentation, material property analysis and classification; on the other hand, lidars can obtain accurate 3D point cloud data of the target, which is of great value in the perception of data such as the distance, shape, size and volume of the target.
[0004] However, most of the current 4D information fusion technologies for hyperspectral images and lidar are based on external parameter calibration using calibration plates, which requires complex calibration steps and a lot of manpower, limiting the efficiency of data fusion between hyperspectral cameras and lidars. At the same time, the existing technology calibrates the external parameters between hyperspectral cameras and lidars once, and the calibration parameters are no longer changed during subsequent use, which leads to a decrease in fusion accuracy due to changes in actual external parameters in a long-term jitter environment. Summary of the invention
[0005] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide a hyperspectral and laser radar 4D anti-shake fusion method and system based on deep learning. The present invention can realize the rapid and accurate acquisition of hyperspectral morphology 4D fusion information under the condition of mechanical shaking of the equipment, which is helpful for applications such as target recognition, semantic segmentation, material property analysis, and precise positioning.
[0006] A hyperspectral and lidar 4D anti-shake fusion method based on deep learning, the steps are as follows:
[0007] 1) The point cloud collected by the LiDAR is first transformed from 3D to 2D space using known internal parameters to generate the corresponding depth map;
[0008] 2) The depth map and the hyperspectral image are input into the HLFN (Hyperspectral Camera and LiDAR Fusion Network) together. The HLFN extracts features from the hyperspectral and depth maps respectively and fuses the features of the two. The rotation vector r and translation vector t are calculated based on the fused features.
[0009] 3) The rotation vector r is transformed into a rotation matrix R through the Rodriguez formula (SE(3)), and finally a predicted transformation matrix T is generated, which indicates the external parameter relationship between the hyperspectral camera and the lidar;
[0010] 4) The original point cloud is then transformed into a 3D-3D space through the predicted transformation matrix T to generate a corrected point cloud;
[0011] 5) The corrected point cloud and the hyperspectral image are in the same coordinate system, and the two are fused to generate the final 4D fusion data.
[0012] The HLFN training method comprises the following steps:
[0013] 2.1) Use the real transformation matrix to perform 3D-3D spatial transformation on the point cloud and fuse it with the hyperspectral image to obtain the reference point cloud, reference depth map and reference spectrum map;
[0014] 2.2) Calculate the prediction transformation matrix of the point cloud and the hyperspectral image according to steps 1) to 3);
[0015] 2.3) Use the predicted transformation matrix to perform 3D-3D spatial transformation on the point cloud and fuse it with the hyperspectral image to obtain the corrected point cloud, corrected depth map and corrected spectrum map;
[0016] 2.4) A geometric loss constraint is formed between the reference point cloud and the corrected point cloud, a photometric loss constraint is formed between the reference depth map and the corrected depth map, and a spectral loss constraint is formed between the reference spectrum map and the corrected spectrum map;
[0017] 2.5) The final loss is the sum of the three losses described in 2.4), and they are balanced by three hyperparameters;
[0018] 2.6) Through back propagation, the HLFN network parameters are continuously iteratively trained to minimize the loss function.
[0019] In step 2.1), the true transformation matrix is obtained by the hyperspectral camera and lidar using the traditional extrinsic calibration method.
[0020] 2.6) are the loss functions described, namely, spectral loss, photometric loss and geometric loss.
[0021] The photometric loss is calculated by calculating the sum of the differences between the pixel values at the same positions of the reference depth map and the corrected depth map.
[0022] The spectral loss is calculated by calculating the sum of the distances between all spectral vectors at the same positions between the reference spectral graph and the correction spectral graph.
[0023] The geometric loss: The geometric loss calculates the spatial distance between the reference point cloud and the correction point cloud.
[0024] The HLFN network structure: the depth map and the hyperspectral image first pass through two branches composed of convolutional layers for feature extraction, and then their features are connected in series, and then through 7 convolutional layers for deep feature fusion. Finally, the fused features are divided into two branches, each of which passes through a 1×1 convolution and a fully connected layer to obtain the rotation vector r and the translation vector t respectively.
[0025] A device using the method described above comprises a hyperspectral camera module, a laser radar module, a power module, a wireless communication module, a host computer module, a display module, an information processing and control module, and a hyperspectral image and point cloud anti-shake fusion module;
[0026] The information processing and control module is respectively connected to the hyperspectral camera module, the laser radar module, the power module, the hyperspectral image and point cloud anti-shake fusion module, and the information processing and control module is connected to the host computer module through the wireless communication module; the information processing and control module is used for information processing and control of the connected modules;
[0027] The hyperspectral camera module is used to obtain hyperspectral information of the target object;
[0028] The laser radar module is used to obtain point cloud data of the target object;
[0029] The hyperspectral image and point cloud anti-shake fusion module is used for 4D anti-shake fusion of hyperspectral image and point cloud data;
[0030] The power module is used for supplying power;
[0031] The wireless communication module is used to transmit data between the host computer module and the information processing and control module.
[0032] The host computer module is provided with a hyperspectral camera and a laser radar calibration module, which is called by the host computer module and is used for external parameter calibration between the hyperspectral camera and the laser radar.
[0033] The hyperspectral camera module comprises an electro-optical tunable filter and a high-speed CCD camera.
[0034] The hyperspectral camera module further includes an RGB camera, which is combined with the high-speed CCD camera in a conjugate mode.
[0035] Beneficial effects of the present invention:
[0036] In the existing technology, the calibration of external parameters between hyperspectral cameras and lidars mainly relies on calibration plates, which requires complicated steps and a lot of manpower, which limits the efficiency of data fusion between hyperspectral cameras and lidars. At the same time, the existing technology hard-calibrates the external parameters between hyperspectral cameras and lidars, that is, once calibrated, the calibration parameters will not change during subsequent use, which leads to the problem of decreased fusion accuracy due to changes in actual external parameters in a long-term jitter environment. In the past, methods either failed to obtain accurate fusion data in a jitter environment, or required multiple recalibrations, which was very time-consuming and labor-intensive.
[0037] The present invention trains a hyperspectral camera and laser radar fusion network (HLPN), which can achieve the following in actual use: (1) the transformation matrix between the hyperspectral camera and the laser radar can be quickly and accurately obtained without the need for a calibration plate, thereby achieving accurate 4D information fusion; (2) the jitter interference of the environment can be suppressed. In a jitter environment, the HLPN will calculate the external parameters between the hyperspectral camera and the laser radar in real time, and can stably perform accurate 4D fusion without being disturbed by jitter. The present invention has significant advantages in applications in automated production lines or robots with long-term jitter environments. On the one hand, hyperspectral images can help with target recognition, semantic segmentation, material property analysis, and other applications. On the other hand, point clouds provide important information such as the size and distance of the target. The present invention lays a foundation for the application of robust, accurate fusion and stable operation between hyperspectral cameras and laser radars. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of the modules of the present invention.
[0039] Figure 2 It is a schematic diagram of the hyperspectral image and point cloud anti-shake fusion module of the present invention.
[0040] Figure 3 Schematic diagram for training HLPN in the present invention.
[0041] Figure 4 This is the network structure diagram of the HLPN in the present invention. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] like Figure 1As shown, a hyperspectral and lidar 4D anti-shake fusion acquisition system based on deep learning is provided, which comprises a hyperspectral camera module, a lidar module, a power module, a wireless communication module, a host computer module, a display module, an information processing and control module, a hyperspectral camera and lidar calibration module, and a hyperspectral image and point cloud anti-shake fusion module. The hyperspectral image and point cloud anti-shake fusion module is implemented by a neural network based on deep learning, and the features of the spectral space and the geometric space of the hyperspectral image and the point cloud are respectively extracted, and the real-time 6DoF pose between the two is jointly calculated, thereby realizing anti-shake 4D information fusion.
[0044] like Figure 2 As shown, the 4D information fusion steps are as follows:
[0045] 1) The point cloud collected by the LiDAR is first transformed from 3D to 2D space using known internal parameters to generate the corresponding depth map;
[0046] 2) The depth map and the hyperspectral image are input into HLFN together. HLFN extracts features from the hyperspectral image and the depth map respectively and fuses the features of the two. The rotation vector r and the translation vector t are calculated based on the fused features.
[0047] 3) The rotation vector r is transformed into a rotation matrix R through the Rodriguez formula (SE(3)), and finally a predicted transformation matrix T is generated, which indicates the external parameter relationship between the hyperspectral camera and the lidar;
[0048] 4) The original point cloud is then transformed into a 3D-3D space through the predicted transformation matrix T to generate a corrected point cloud;
[0049] 5) The corrected point cloud and the hyperspectral image are in the same coordinate system, and the two are fused to generate the final 4D fusion data.
[0050] In practical applications, because the external parameters between the hyperspectral camera and the lidar are quickly calculated through HLFN, the mechanical jitter between the hyperspectral camera and the lidar during actual operation can be corrected in real time to achieve the purpose of anti-shake.
[0051] The neural network in the hyperspectral image and point cloud anti-shake fusion module needs to be trained, and the training process is constrained by the external parameters obtained by the hyperspectral camera and lidar calibration module.
[0052] like Figure 3 As shown in the figure, when training HLFN, multiple sets of hyperspectral images and point clouds with real transformation matrices (which can be obtained by hyperspectral camera and lidar calibration module) are required as training data sets. The training process is as follows:
[0053] (1) Use the real transformation matrix to perform 3D-3D spatial transformation on the point cloud and fuse it with the hyperspectral image to obtain the reference point cloud, reference depth map, and reference spectrum map;
[0054] (2) Calculate the prediction transformation matrix of the point cloud and hyperspectral image according to 1) to 3) in the 4D information fusion steps;
[0055] (3) Use the predicted transformation matrix to perform 3D-3D spatial transformation on the point cloud and fuse it with the hyperspectral image to obtain the corrected point cloud, corrected depth map, and corrected spectrum map;
[0056] (4) A geometric loss constraint is formed between the reference point cloud and the corrected point cloud, a photometric loss constraint is formed between the reference depth map and the corrected depth map, and a spectral loss constraint is formed between the reference spectrum map and the corrected spectrum map;
[0057] (5) The final loss is the sum of the three losses described in (4), and they are balanced by three hyperparameters;
[0058] (6) Through back propagation, the HLFN network parameters are continuously iteratively trained to minimize the loss function.
[0059] The neural network in the hyperspectral image and point cloud anti-shake fusion module optimizes three loss functions during training, namely, spectral loss, photometric loss and geometric loss.
[0060] The definition of photometric loss is as follows: Photometric loss is calculated as the sum of the differences between the pixel values (depth values) at the same positions of the reference depth map and the corrected depth map. If the predicted transformation matrix is closer to the real transformation matrix, the difference between the pixel values (depth values) at the same positions of the reference depth map and the corrected depth map will be smaller, and the photometric loss is relatively small; conversely, if the predicted transformation matrix is significantly different from the real transformation matrix, the photometric loss is relatively large.
[0061] Spectral loss and geometric loss are also similar. Spectral loss calculates the sum of the distances between all spectral vectors at the same position between the reference spectrum and the correction spectrum, while geometric loss calculates the spatial (point cloud three-dimensional coordinates) distance between the reference point cloud and the correction point cloud.
[0062] The HLFN network structure is as follows Figure 4 As shown in the figure, the depth map and hyperspectral image first pass through two branches composed of convolutional layers for feature extraction, and then their features are connected in series, and then deep feature fusion is performed through 7 convolutional layers. Finally, the fused features are divided into two branches, each of which passes through a 1×1 convolution and a fully connected layer to obtain the rotation vector r and translation vector t respectively.
[0063] The hyperspectral camera module comprises an electro-optical tunable filter and a high-speed CCD camera, and is connected to an information processing and control module to obtain hyperspectral information of a target object.
[0064] By continuously changing the transmission wavelength of the electro-optical tunable filter and continuously collecting data with a high-speed CCD camera, a hyperspectral image can be obtained.
[0065] The hyperspectral camera module can also be combined with an RGB camera in a conjugate mode. The RGB image has a relatively high spatial resolution and can provide better display effects for 4D fusion.
[0066] The hyperspectral camera and laser radar calibration module is located in the host computer module, and is connected to the hyperspectral camera module and the laser radar module through the wireless communication module and the information processing and control module respectively, and is used for external parameter calibration between the hyperspectral camera and the laser radar.
[0067] The hyperspectral camera and lidar calibration modules require the use of a calibration plate, while the hyperspectral image and point cloud anti-shake fusion module does not require the use of a calibration plate.
[0068] The hyperspectral camera and lidar calibration module uses the traditional calibration method, whose purpose is to obtain the real transformation matrix used to train the HLPN network; once the HLPN is trained, no calibration plate is required during the inference phase, and the network can complete the fusion between the hyperspectral image and the point cloud by identifying the features of both.
[0069] In actual use, the working steps of the entire system are as follows:
[0070] 1. If HLPN has not been trained:
[0071] First, you need to collect the data set for training:
[0072] 1) Fix the relative position of the hyperspectral camera and the laser radar, aim at a certain target, and then the host computer module sends a control command to the information processing and control module through the wireless communication module. After receiving the command to collect data, the information processing and control module controls the hyperspectral camera module to collect a hyperspectral image and controls the laser radar module to collect point clouds. These data are then sent to the host computer through the wireless transmission module, and the host computer stores them.
[0073] 2) Change the target content and collect data again according to the steps in 1) until multiple targets (for example, 20) are captured.
[0074] 3) Replace the target with the calibration plate and collect data multiple times (for example, 5 times) according to the steps in 1).
[0075] 4) The hyperspectral camera and lidar calibration module in the host computer module uses the data collected in step 3) for calibration. The calibration result is the real transformation matrix, which will be used as ground truth when training the HLPN network.
[0076] 5) Pack the data collected in step 2) and the real transformation matrix obtained by calibration in step 4) to form a set of training data.
[0077] 6) Change the relative positions of the hyperspectral camera and the lidar, repeat the above steps, obtain multiple sets of training data and form a training data set.
[0078] Then, the HLPN is trained. The host computer module sends the obtained training data set to the hyperspectral image and point cloud anti-shake fusion module through the wireless communication module, information processing and control module. Figure 3 The steps shown are used for training to obtain the trained HLPN.
[0079] 2. If HLPN has been trained, in actual application, 4D fusion of newly collected data is performed as follows:
[0080] 1) The host computer module sends a control command to the information processing and control module through the wireless communication module. After receiving the command to collect data, the information processing and control module controls the hyperspectral camera module to collect a hyperspectral image and controls the lidar module to collect point clouds.
[0081] 2) Then these data are sent to the hyperspectral image and point cloud anti-shake fusion module, which follows the described Figure 2 The steps shown are used to perform 4D fusion to obtain 4D fused data.
[0082] 3) The hyperspectral image and point cloud anti-shake fusion module sends the fused 4D data to the host computer module through the information processing and control module and the wireless communication module, and finally uses the display module in the host computer module to display the fused 4D data.
[0083] The embodiments described above can be further combined or replaced, and the embodiments are only descriptions of preferred embodiments of the present invention, and do not limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various changes and improvements made by ordinary technicians in this field to the technical solution of the present invention belong to the protection scope of the present invention. The protection scope of the present invention is given by the attached claims and any equivalents thereof.
Claims
1. A hyperspectral and laser radar 4D anti-shake fusion method based on deep learning, characterized by: Here are the steps: 1) The point cloud collected by the LiDAR is first transformed from 3D to 2D space using known internal parameters to generate the corresponding depth map; 2) The depth map and the hyperspectral image are input into the hyperspectral camera and lidar fusion network HLFN together. HLFN extracts features from the hyperspectral and depth maps respectively and fuses the features of the two. The rotation vector r and translation vector t are calculated based on the fused features. 3) The rotation vector r is transformed into a rotation matrix R through the Rodriguez formula, and finally a prediction transformation matrix T is generated, which indicates the external parameter relationship between the hyperspectral camera and the lidar; 4) The original point cloud is then transformed into a 3D-3D space through the predicted transformation matrix T to generate a corrected point cloud; 5) The corrected point cloud and the hyperspectral image are in the same coordinate system, and the two are fused to generate the final 4D fusion data; The HLFN training method comprises the following steps: 2.1) Use the real transformation matrix to perform 3D-3D spatial transformation on the point cloud and fuse it with the hyperspectral image to obtain the reference point cloud, reference depth map and reference spectrum map; 2.2) Calculate the prediction transformation matrix of the point cloud and the hyperspectral image according to steps 1) to 3); 2.3) Use the predicted transformation matrix to perform 3D-3D spatial transformation on the point cloud and fuse it with the hyperspectral image to obtain the corrected point cloud, corrected depth map and corrected spectrum map; 2.4) A geometric loss constraint is formed between the reference point cloud and the corrected point cloud, a photometric loss constraint is formed between the reference depth map and the corrected depth map, and a spectral loss constraint is formed between the reference spectrum map and the corrected spectrum map; 2.5) The final loss is the sum of the three losses described in 2.4), and they are balanced by three hyperparameters; 2.6) Through back propagation, the HLFN network parameters are continuously iteratively trained to minimize the loss function.
2. The method according to claim 1, characterized in that: In step 2.1), the true transformation matrix is obtained by the hyperspectral camera and lidar using the traditional extrinsic calibration method.
3. The method according to claim 1, characterized in that: 2.6) are spectral loss, photometric loss and geometric loss.
4. The method according to claim 1 or 3, characterized in that: The photometric loss is calculated by calculating the sum of the differences between the pixel values at the same positions of the reference depth map and the corrected depth map. The spectral loss is calculated by calculating the sum of the distances between all spectral vectors at the same positions between the reference spectral graph and the correction spectral graph. The geometric loss: The geometric loss calculates the spatial distance between the reference point cloud and the correction point cloud.
5. The method according to claim 1, characterized in that: The HLFN network structure: the depth map and the hyperspectral image first pass through two branches composed of convolutional layers for feature extraction, and then their features are connected in series, and then through 7 convolutional layers for deep feature fusion. Finally, the fused features are divided into two branches, each of which passes through a 1×1 convolution and a fully connected layer to obtain the rotation vector r and the translation vector t respectively.
6. A system using the method according to claim 1, characterized in that: It includes a hyperspectral camera module, a laser radar module, a power module, a wireless communication module, a host computer module, a display module, an information processing and control module, and a hyperspectral image and point cloud anti-shake fusion module; The information processing and control module is respectively connected to the hyperspectral camera module, the laser radar module, the power module, the hyperspectral image and point cloud anti-shake fusion module, and the information processing and control module is connected to the host computer module through the wireless communication module; the information processing and control module is used for information processing and control of the connected modules; The hyperspectral camera module is used to obtain hyperspectral information of the target object; The laser radar module is used to obtain point cloud data of the target object; The hyperspectral image and point cloud anti-shake fusion module is used for 4D anti-shake fusion of hyperspectral image and point cloud data; The power module is used for supplying power; The wireless communication module is used to transmit data between the host computer module and the information processing and control module.
7. The system according to claim 6, characterized in that: The host computer module is provided with a hyperspectral camera and a laser radar calibration module, which is called by the host computer module and is used for external parameter calibration between the hyperspectral camera and the laser radar.
8. The system according to claim 6, characterized in that: The hyperspectral camera module includes an electro-optical tunable filter and a high-speed CCD camera.
9. The system according to claim 8, characterized in that: The hyperspectral camera module further includes an RGB camera, which is combined with the high-speed CCD camera in a conjugate mode.
Citation Information
Patent Citations
Camera and laser radar calibration method and system based on end-to-end and medium
CN113160330A
Multi-source sensing method and system for water surface unmanned equipment
WO2020237693A1