Target identification system and method based on fusion of laser radar and multispectral polarization imaging

By using spatiotemporal synchronization processing and cross-modal feature fusion, the problem of data acquisition deviation between lidar and multispectral imaging was solved, achieving high-precision target recognition and environmental information output, and improving the accuracy of target recognition and multi-task processing capabilities.

CN121069410APending Publication Date: 2025-12-05HUBEI HUAZHONG PHOTOELECTRIC SCI & TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511177135.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

In existing technologies, there are discrepancies in the time and spatial location of data acquisition for lidar and multispectral imaging. Image pixels and point cloud spatial locations cannot be accurately correlated. Relying on information from only a single polarization direction results in limited ability to distinguish target material features. Point cloud data is susceptible to noise interference, and it is difficult to achieve high-precision target recognition and rich environmental information output when processing multiple tasks.

Method used

The sensor configuration and data preprocessing module performs spatiotemporal synchronous processing, and the multispectral polarization camera group and lidar are used for data acquisition. Combined with the cross-modal feature extraction and fusion module, a high-dimensional semantically enhanced point cloud representation containing image semantics and point cloud geometry is generated. The multi-task output and optimization module outputs the three-dimensional target recognition results, target velocity information and pixel-level depth map.

Benefits of technology

It achieves accurate correspondence between image pixels and point cloud spatial positions, improves the ability to distinguish target materials, reduces the impact of noise interference in point cloud data, and can simultaneously output high-precision target recognition and rich environmental information to meet diverse needs in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121069410A_ABST
    Figure CN121069410A_ABST
Patent Text Reader

Abstract

The invention discloses a laser radar and multispectral polarization imaging fused target identification system and method. The system comprises a sensor configuration and data preprocessing module, a cross-modal feature extraction and fusion module and a multi-task output and optimization module. The sensor configuration and data preprocessing module performs time-space synchronization processing on the collected original optical signals and laser signals in the environment to obtain multispectral image data and laser radar point cloud data; the cross-modal feature extraction and fusion module performs feature extraction and fusion enhancement processing on the multispectral image data and the laser radar point cloud data, and outputs high-dimensional semantic enhancement point cloud representation containing image semantics and point cloud geometry; and the multi-task output and optimization module processes the high-dimensional semantic enhanced point cloud representation and outputs a three-dimensional target recognition result, target speed information and a pixel-level depth map. Through module design and data processing, the defects in the prior art are overcome, and the accuracy of target recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of target recognition technology, and more specifically, to a target recognition system and method that integrates lidar and multispectral polarization imaging. Background Technology

[0002] In today's field of intelligent sensing, target recognition and related environmental perception technologies hold an extremely important position. Whether in autonomous driving scenarios, where precise identification of vehicles, pedestrians, traffic signs, and other targets on the road is required, along with acquiring crucial information such as their speed and location to ensure driving safety; or in security monitoring, where real-time identification and tracking of suspicious individuals and objects within a monitored area is necessary; or in industrial inspection, where accurate judgment of product defects and dimensions is crucial, all rely on efficient and accurate target recognition and environmental perception technologies. The implementation of these technologies often requires the acquisition of rich environmental information from multiple sensors. Through the processing and analysis of this information, accurate target identification and comprehensive environmental perception can be achieved.

[0003] Currently, lidar and multispectral imaging are two commonly used sensor technologies. LiDAR acquires 3D point cloud data of the surrounding environment by emitting a laser beam and measuring the time it takes for the reflected light to pass through. It can provide accurate distance information and the geometry of objects, offering significant advantages for constructing 3D models of the environment and detecting the position of objects. Multispectral imaging technology, on the other hand, uses light of different wavelengths to capture the spectral characteristics of targets. Different materials have different reflectivity at different wavelengths, and by analyzing this spectral information, the material and type of the target can be effectively identified.

[0004] However, existing technologies have several drawbacks. First, the data acquisition times of LiDAR and multispectral cameras may differ, and their spatial locations may also be inconsistent. This can lead to inaccurate correspondence between image pixels and point cloud spatial locations during data fusion, severely impacting the accuracy of subsequent feature extraction and target recognition. Second, multispectral imaging often utilizes only information from a single polarization direction, failing to fully exploit the value of polarization characteristics for target recognition, resulting in limited ability to distinguish features such as target material. Third, in harsh environments such as rain and fog, LiDAR point cloud data is severely affected by noise interference. Existing point cloud quality optimization methods are ineffective, resulting in sparse and inaccurate point cloud data that struggles to effectively support subsequent target recognition tasks. Finally, existing technologies lack sufficient collaborative processing capabilities for multi-task tasks such as 3D target recognition, velocity estimation, and pixel-level depth map generation, making it difficult to simultaneously achieve high-precision target recognition and rich environmental information output. Summary of the Invention

[0005] To address at least one deficiency or improvement need in the prior art, this invention provides a target recognition system and method that fuses lidar and multispectral polarization imaging. This system solves the problems in the prior art where there are deviations in the time and spatial position of data acquisition between lidar and multispectral cameras, the inability to accurately correspond the image pixels to the spatial position of the point cloud, the use of information from only a single polarization direction, the limited ability to distinguish features such as target material, the susceptibility of point cloud data to noise interference, and the difficulty in simultaneously achieving high-precision target recognition and rich environmental information output during multi-task processing.

[0006] To achieve the above objectives, according to a first aspect of the present invention, a target recognition system that fuses lidar and multispectral polarization imaging is provided, comprising: a sensor configuration and data preprocessing module, a cross-modal feature extraction and fusion module, and a multi-task output and optimization module connected in sequence; Among them, the sensor configuration and data preprocessing module is used to perform spatiotemporal synchronous processing on the raw optical signals and laser signals in the collected environment to obtain multispectral image data and lidar point cloud data with different polarization states. The cross-modal feature extraction and fusion module is used to perform feature extraction and fusion enhancement processing on multispectral image data and lidar point cloud data, and outputs a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry. The multi-task output and optimization module is used to process the high-dimensional semantically enhanced point cloud representation and output the 3D target recognition results, target velocity information, and pixel-level depth map.

[0007] In one possible implementation, the sensor configuration and data preprocessing module also includes: a multispectral polarization camera group, a lidar, and a spatiotemporal synchronization and calibration unit; Among them, the multispectral polarization camera group is equipped with several multispectral cameras with linear polarization filters at different angles, and dynamically adjusts the exposure parameters of each frequency band based on the ambient light intensity to acquire original multispectral images. The lidar collects point clouds and improves the quality of the point clouds through adaptive algorithms to obtain lidar point cloud data; The time-space synchronization and calibration unit is used for system time synchronization, internal parameter calibration, and external parameter calibration.

[0008] In one possible implementation, the cross-modal feature extraction and fusion module further includes: a multispectral polarization feature extraction unit, a lidar feature extraction unit, and a cross-modal feature conversion and fusion unit; Among them, the multispectral polarization feature extraction unit is used to extract feature maps of each band from the original multispectral image using a four-branch feature extraction network and then perform weighted fusion to obtain multispectral image features; The lidar feature extraction unit extracts local descriptors from lidar point cloud data through a point cloud feature extraction network and encodes them to output lidar point cloud features. The cross-modal feature conversion and fusion unit projects the LiDAR point cloud onto the image plane to establish a pixel-level correspondence with the point cloud level, and maps the multispectral image features to the point cloud space through back projection, and then stitches and fuses them with the point cloud geometric features to obtain a fused feature map.

[0009] In one possible implementation, the multi-task output and optimization module also includes: a target recognition and velocity estimation unit, a point cloud generation unit, and an all-weather robustness enhancement unit. The target recognition and velocity estimation unit inputs the fused feature map into the 3D target detection network and outputs the 3D bounding box information and category label of each target. The optical flow-based temporal modeling algorithm outputs the instantaneous velocity and acceleration of each target. The point cloud generation unit processes the fused feature map based on depth map completion and image-guided point cloud interpolation techniques, and outputs a pixel-level depth map. The all-weather robustness enhancement unit is used to evaluate the input quality of each modality in real time, dynamically generate confidence weights based on preset indicators, and dynamically adjust the features of each modality according to the confidence level.

[0010] In one possible implementation, the spatiotemporal synchronization and calibration unit further includes: a time synchronization subunit, an intrinsic parameter calibration subunit, and an extrinsic parameter calibration subunit; The time synchronization subunit connects the multispectral polarization camera group and the lidar to the same control clock through an external trigger signal, and uses a unified frequency to trigger data acquisition. The intrinsic parameter calibration subunit independently calibrates a multispectral camera with several linearly polarized filters at different angles using a calibration plate, thereby determining the camera's intrinsic parameters. The extrinsic parameter calibration subunit calculates the relative pose matrix between several multispectral cameras and the pose transformation matrix between the multispectral main camera and the lidar based on the preset calibration structure and calibration tools.

[0011] According to a second aspect of the present invention, a target recognition method based on the fusion of lidar and multispectral polarization imaging is also provided, comprising a target recognition system based on the fusion of lidar and multispectral polarization imaging as described in any of the above implementations, including a multispectral polarization camera group and a lidar, characterized in that the method comprises: The original optical and laser signals collected from the environment are processed in a time-space synchronization manner to obtain multispectral image data and lidar point cloud data with different polarization states. Feature extraction and fusion enhancement processing are performed on multispectral image data and lidar point cloud data to obtain a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry; The high-dimensional semantically enhanced point cloud representation is optimized to obtain 3D target recognition results, target velocity information, and pixel-level depth maps.

[0012] In one possible implementation, the raw optical and laser signals from the acquired environment are spatiotemporally synchronized to obtain multispectral image data and lidar point cloud data with different polarization states, and the implementation also includes: Based on the dynamic adjustment of exposure parameters for each frequency band according to ambient light intensity, original multispectral images with different polarization states are acquired. The initial point cloud data is processed using an adaptive algorithm to improve its quality, thus obtaining lidar point cloud data. The system time was synchronized, and the intrinsic and extrinsic parameters of the multispectral camera were calibrated separately.

[0013] In one possible implementation, feature extraction and fusion enhancement processing are performed on multispectral image data and lidar point cloud data to obtain a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry. This also includes: A four-branch feature extraction network is used to extract feature maps of each band from the original multispectral image and then weighted and fused to obtain multispectral image features. Local descriptors are extracted from lidar point cloud data and encoded using a point cloud feature extraction network to output lidar point cloud features. The point cloud of the LiDAR is projected onto the image plane to establish a correspondence between the pixel level and the point cloud level. The multispectral image features are then back-projected onto the point cloud space and stitched together with the geometric features of the point cloud to obtain a fused feature map.

[0014] In one possible implementation, multiple pairs of high-dimensional semantically enhanced point cloud representations are optimized to obtain 3D target recognition results, target velocity information, and pixel-level depth maps, and also include: The fused feature map is input into the 3D target detection network, which outputs the 3D bounding box information and category label of each target, and outputs the instantaneous velocity and acceleration of each target based on the optical flow temporal modeling algorithm. The fused feature map is processed based on depth map completion and image-guided point cloud interpolation techniques to output a pixel-level depth map; The system evaluates the input quality of each modality in real time, dynamically generates confidence weights based on preset indicators, and dynamically adjusts the features of each modality according to the confidence level.

[0015] One possible implementation involves synchronizing the system time and performing intrinsic and extrinsic parameter calibrations on the multispectral camera, and also includes: The multispectral polarization camera group and the lidar are connected to the same control clock by an external trigger signal, and data acquisition is triggered by a unified frequency. The camera intrinsic parameters were determined by independently calibrating a multispectral camera with several linear polarization filters at different angles using a calibration plate. The relative pose matrix between several multispectral cameras and the pose transformation matrix between the multispectral main camera and the lidar are calculated based on the preset calibration structure and calibration tools.

[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: The target recognition system fusion of lidar and multispectral polarization imaging provided by this invention, through sensor configuration and data preprocessing modules, simultaneously acquires raw multispectral images and lidar point cloud data, and performs spatiotemporal synchronous processing to ensure data consistency in time and space. This avoids temporal and spatial position deviations caused by independent operation of different sensors, ensuring accurate correspondence between image pixels and point cloud spatial positions. The cross-modal feature extraction and fusion module performs feature extraction and fusion enhancement processing on multispectral image data and lidar point cloud data, utilizing the advantages of different modal data to generate a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry. This enables more comprehensive and accurate capture of target material and other features, improving the ability to distinguish different target materials. During the cross-modal feature extraction and fusion process, the fusion enhancement processing of multispectral image data and lidar point cloud data allows for the optimization and correction of point cloud data using the rich information of the multispectral image, effectively reducing the susceptibility of point cloud data to noise interference, improving the quality and reliability of point cloud data, and thus enhancing the accuracy of target recognition. The multi-task output and optimization module can comprehensively process high-dimensional semantically enhanced point cloud representations and output three-dimensional target recognition results, target velocity information, and pixel-level depth maps. It can not only achieve high-precision target recognition, but also provide rich environmental information, meeting the diverse needs for target recognition and environmental perception in complex scenarios. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic diagram of an embodiment of the target recognition system that fuses lidar and multispectral polarization imaging provided by the present invention; Figure 2A flowchart illustrating an embodiment of the target recognition method that fuses lidar and multispectral polarization imaging provided by the present invention; Figure 3 Provided by the present invention Figure 2 A flowchart illustrating an embodiment of step S201; Figure 4 Provided by the present invention Figure 2 A flowchart illustrating an embodiment of step S202; Figure 5 Provided by the present invention Figure 2 A flowchart illustrating an embodiment of step S203; Figure 6 Provided by the present invention Figure 3 A flowchart illustrating an embodiment of step S303. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0020] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0021] This invention provides a target recognition system and method that integrates lidar and multispectral polarization imaging, which will be described below.

[0022] Please see Figure 1 , Figure 1 This is a schematic diagram of an embodiment of the target recognition system 100 that fuses lidar and multispectral polarization imaging provided by the present invention. In a specific embodiment of the present invention, a target recognition system 100 that fuses lidar and multispectral polarization imaging is disclosed, including: a sensor configuration and data preprocessing module 110, a cross-modal feature extraction and fusion module 120, and a multi-task output and optimization module 130 connected in sequence. Among them, the sensor configuration and data preprocessing module 110 is used to perform spatiotemporal synchronous processing on the raw optical signals and laser signals in the collected environment to obtain multispectral image data and lidar point cloud data with different polarization states. The cross-modal feature extraction and fusion module 120 is used to perform feature extraction and fusion enhancement processing on multispectral image data and lidar point cloud data, and outputs a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry. The multi-task output and optimization module 130 is used to process the high-dimensional semantically enhanced point cloud representation and output the three-dimensional target recognition results, target velocity information and pixel-level depth map.

[0023] In the above embodiments, the sensor configuration and data preprocessing module 110 acquires raw multispectral images and lidar point cloud data and performs preliminary processing. Its input is the raw optical signal and laser signal in the environment, and its output is multispectral image data and lidar point cloud data that have been time-synchronized, spatially calibrated and preprocessed.

[0024] The cross-modal feature extraction and fusion module 120 takes as input the preprocessed multispectral image data and lidar point cloud data output by the sensor configuration and data preprocessing module 110, and outputs a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry.

[0025] The multi-task output and optimization module 130 takes as input the high-dimensional semantically enhanced point cloud representation output by the cross-modal feature extraction and fusion module 120, and outputs the three-dimensional target recognition result, target velocity information, and full-resolution pixel-level depth map (i.e., RGB-D point cloud).

[0026] Compared with existing technologies, the target recognition system 100 fusion of lidar and multispectral polarization imaging provided in this embodiment, through the sensor configuration and data preprocessing module 110, simultaneously acquires raw multispectral images and lidar point cloud data, and performs spatiotemporal synchronization processing, ensuring data consistency in time and space. This avoids temporal and spatial position deviations caused by the independent operation of different sensors, ensuring accurate correspondence between image pixels and point cloud spatial positions. The cross-modal feature extraction and fusion module 120 performs feature extraction and fusion enhancement processing on multispectral image data and lidar point cloud data, utilizing the advantages of different modal data to generate a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry. This enables more comprehensive and accurate capture of target material and other features, improving the ability to distinguish different target materials. During the cross-modal feature extraction and fusion process, the fusion enhancement processing of multispectral image data and lidar point cloud data allows for the optimization and correction of point cloud data using the rich information of multispectral images, effectively reducing the susceptibility of point cloud data to noise interference, improving the quality and reliability of point cloud data, and thus enhancing the accuracy of target recognition. The multi-task output and optimization module 130 can comprehensively process the high-dimensional semantically enhanced point cloud representation and output three-dimensional target recognition results, target velocity information and pixel-level depth maps. It can not only achieve high-precision target recognition, but also provide rich environmental information, meeting the diverse needs for target recognition and environmental perception in complex scenarios.

[0027] In some embodiments of the present invention, the sensor configuration and data preprocessing module 110 further includes: a multispectral polarization camera group, a lidar, and a spatiotemporal synchronization and calibration unit; Among them, the multispectral polarization camera group is equipped with several multispectral cameras with linear polarization filters at different angles, and dynamically adjusts the exposure parameters of each frequency band based on the ambient light intensity to acquire original multispectral images. The lidar collects point clouds and improves the quality of the point clouds through adaptive algorithms to obtain lidar point cloud data; The time-space synchronization and calibration unit is used for system time synchronization, internal parameter calibration, and external parameter calibration.

[0028] In the above embodiments, the dynamic exposure compensation of the multispectral polarization camera group is based on the light intensity value collected in real time by the ambient light sensor. The exposure time range of each band is dynamically adjusted to 1 / 1000s - 1 / 30s and the gain coefficient range is 1x - 8x through an online photometric calibration algorithm. In low-light environments (<100 lux), the exposure time of the near-infrared band is automatically extended to 1 / 100s and the gain is increased to 4x. At the same time, the gain of the blue light band is compressed to avoid overexposure. Combined with the white balance algorithm, the color consistency of each band is corrected, and finally the original multispectral image is obtained.

[0029] In rain and fog environments, lidar optimizes point cloud quality through adaptive algorithms. Specifically, this includes filtering low-intensity, temporally isolated noise points based on point cloud echo intensity and temporal continuity analysis, and supplementing the point cloud density in sparse point clouds in low-reflectivity areas using interpolation algorithms to ensure a point density ≥ 10 points / cm² in key areas. 2 We obtained point cloud data from the lidar.

[0030] In the spatiotemporal synchronization and calibration unit, intrinsic parameter calibration uses a high-contrast calibration board to acquire at least 20 sets of images at different poses in the imaging band of each camera. The principal point position, focal length, and distortion coefficient are solved using a multi-channel camera calibration tool, and the color / illumination inconsistency between channels is compensated by a photometric calibration model. Extrinsic parameter calibration uses a planar reflective calibration board or a three-dimensional structure with marked points to simultaneously acquire LiDAR point clouds and multispectral images in multi-view scenes. The extrinsic parameters are estimated using a calibration tool, and the extrinsic parameters between the two multispectral cameras and between the multispectral master camera and the LiDAR are output. The output of the extrinsic parameter estimation is in the form of a 4x4 pose transformation matrix. The error control requirements are that the rotation error of the extrinsic parameter calibration is <0.5°, the translation error is <5mm, the calibration viewing angle span is >45° covering the horizontal and vertical directions, and the system supports periodic self-checking and automatic fine-tuning mechanisms.

[0031] In some embodiments of the present invention, the cross-modal feature extraction and fusion module 120 further includes: a multispectral polarization feature extraction unit, a lidar feature extraction unit, and a cross-modal feature conversion and fusion unit; Among them, the multispectral polarization feature extraction unit is used to extract feature maps of each band from the original multispectral image using a four-branch feature extraction network and then perform weighted fusion to obtain multispectral image features; The lidar feature extraction unit extracts local descriptors from lidar point cloud data through a point cloud feature extraction network and encodes them to output lidar point cloud features. The cross-modal feature conversion and fusion unit projects the LiDAR point cloud onto the image plane to establish a pixel-level correspondence with the point cloud level, and maps the multispectral image features to the point cloud space through back projection, and then stitches and fuses them with the point cloud geometric features to obtain a fused feature map.

[0032] In the above embodiment, in the multispectral polarization feature extraction unit, each CNN branch of the four-branch feature extraction network adopts a lightweight structure, containing 3 convolutional layers with a kernel size of 3×3, a stride of 1, and channel numbers of 16 / 32 / 64 respectively. The polarization attention mechanism module calculates the polarization difference weights of the feature maps corresponding to the 45° and 90° polarization images, and performs weighted fusion of the feature maps to obtain multispectral image features.

[0033] In the lidar feature extraction unit, the point cloud is first downsampled and outliers are removed by voxel grid filtering. The voxel size is 0.1m×0.1m×0.1m. Local descriptors such as geometric center, curvature, point density, and normal vector are calculated for the point set within each voxel. These descriptors are then input into a multilayer perceptron or PointNet++ structure for encoding, and high-dimensional semantic lidar point cloud features are output.

[0034] In the cross-modal feature transformation and fusion unit, the lidar point cloud is projected onto the multispectral image plane based on the calibrated extrinsic parameter matrix to establish a one-to-one correspondence between pixel level and point level. The constructed fusion network is a network based on the Transformer architecture, which maps the multispectral image features to the point cloud space through back projection, and performs feature-level stitching or attention-weighted fusion with the point cloud geometric features to obtain the fused feature map.

[0035] In some embodiments of the present invention, the multi-task output and optimization module 130 further includes: a target recognition and velocity estimation unit, a point cloud generation unit, and an all-weather robustness enhancement unit; The target recognition and velocity estimation unit inputs the fused feature map into the 3D target detection network and outputs the 3D bounding box information and category label of each target. The optical flow-based temporal modeling algorithm outputs the instantaneous velocity and acceleration of each target. The point cloud generation unit processes the fused feature map based on depth map completion and image-guided point cloud interpolation techniques, and outputs a pixel-level depth map. The all-weather robustness enhancement unit is used to evaluate the input quality of each modality in real time, dynamically generate confidence weights based on preset indicators, and dynamically adjust the features of each modality according to the confidence level.

[0036] In the above embodiments, the target recognition and velocity estimation unit adopts a Transformer architecture for the 3D target detection network, which includes a position encoding module, a self-attention multi-head mechanism, a feedforward encoder, and a multi-scale decoder. Combining the multimodal features of the image and point cloud, it outputs the 3D bounding box information and category label of each candidate target. The target velocity estimation part adopts a temporal modeling algorithm based on optical flow. The system continuously collects multiple frames of point cloud data, inputs them into the deep learning network to model trajectory changes, and outputs the instantaneous velocity and acceleration of each target.

[0037] In the point cloud generation unit, based on the depth map completion and image-guided point cloud interpolation technology, the input is the feature map of the previous output. The depth estimation of the occluded edge and low texture area is enhanced by the multi-scale dilated convolution module, and finally a dense RGB-D point cloud (i.e., pixel-level depth map) is generated. At the same time, the multispectral image details of the near-infrared channel are used as weight map to enhance the texture structure of the point cloud boundary and complete all regions of the network's detection head [2]. During the training stage, the pairing data of rain and fog scene images and sparse point clouds are introduced to ensure the network's ability to complete the occluded area.

[0038] In the all-weather robustness enhancement unit, the multimodal confidence assessment module evaluates the input quality of each modality in real time and dynamically generates confidence weights based on indicators such as image contrast, illumination equalization, point cloud density, and effective laser beam ratio. The fusion module dynamically adjusts the features of each modality according to confidence through attention mechanism or weighted averaging. During the model training phase, synthetic extreme environment data is introduced and domain adaptation technology is used to improve the network's generalization ability in the non-training domain. The dynamic weights are updated once before each feature fusion. Adversarial training is jointly trained with the main task, and domain adversarial optimization is performed once in each iteration. The robustness enhancement path is only activated when the sensor signal degradation score exceeds the threshold.

[0039] In some embodiments of the present invention, the spatiotemporal synchronization and calibration unit further includes: a time synchronization subunit, an internal parameter calibration subunit, and an external parameter calibration subunit; The time synchronization subunit connects the multispectral polarization camera group and the lidar to the same control clock through an external trigger signal, and uses a unified frequency to trigger data acquisition. The intrinsic parameter calibration subunit independently calibrates a multispectral camera with several linearly polarized filters at different angles using a calibration plate, thereby determining the camera's intrinsic parameters. The extrinsic parameter calibration subunit calculates the relative pose matrix between several multispectral cameras and the pose transformation matrix between the multispectral main camera and the lidar based on the preset calibration structure and calibration tools.

[0040] In the above embodiments, the multispectral polarization camera group and the lidar are connected to the same clock source by an external trigger signal, such as a TTL level pulse. Data acquisition is triggered by a uniform 10Hz frequency, and the synchronization error is controlled within ±1ms. Each frame of image and point cloud data is automatically accompanied by a uniform timestamp to ensure the timing of acquisition.

[0041] Using a high-contrast checkerboard calibration board, such as a 10mm×10mm black and white grid, two multispectral cameras were calibrated independently. The principal point position, focal length, distortion coefficient and other intrinsic parameters were solved by multi-view images. The brightness difference between each band was compensated by photometric calibration to ensure the geometric distortion correction and color consistency of the image.

[0042] Based on a planar reflective calibration board or a three-dimensional structure with marked points, LiDAR point clouds and multispectral images are simultaneously acquired in multi-view scenarios. The relative pose matrix between two multispectral cameras is calculated using calibration tools to describe the rotation and translation relationship between the two cameras. Furthermore, the pose transformation matrix between the multispectral main camera and the LiDAR is calculated to achieve pixel-level spatial mapping between the LiDAR point cloud and the multispectral image.

[0043] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the target recognition method fusion of lidar and multispectral polarization imaging provided by the present invention. According to a second aspect of the present invention, a target recognition method fusion of lidar and multispectral polarization imaging is also provided, based on a target recognition system 100 fusion of lidar and multispectral polarization imaging as described in any of the above implementations, including a multispectral polarization camera group and a lidar, characterized in that the method includes: S201. Perform spatiotemporal synchronization processing on the raw optical signals and laser signals collected from the environment to obtain multispectral image data and lidar point cloud data with different polarization states. S202. Perform feature extraction and fusion enhancement processing on multispectral image data and lidar point cloud data to obtain a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry. S203. Optimize the high-dimensional semantically enhanced point cloud representation to obtain the three-dimensional target recognition result, target velocity information, and pixel-level depth map.

[0044] In the above embodiments, a time synchronization mechanism is used to ensure the time consistency between the multispectral polarization camera group and the lidar during data acquisition, eliminating registration errors caused by time differences. Simultaneously, spatial registration technology is used to precisely align the multispectral image and the lidar point cloud in space, enabling the fusion and analysis of data from two different modes within the same coordinate system. After spatiotemporal synchronization processing, multispectral image data and lidar point cloud data are obtained.

[0045] For multispectral image data, semantic features such as texture, color, and spectral reflectance are extracted to reflect the target's material, category, and other attribute information. For LiDAR point cloud data, geometric features such as the target's shape, surface normals, and curvature are extracted to describe the target's three-dimensional structure. Through feature fusion technology, the semantic features of the multispectral image are deeply fused with the geometric features of the LiDAR point cloud to form a high-dimensional semantically enhanced point cloud representation, which retains the target's geometric information while incorporating rich semantic information.

[0046] By employing 3D object detection algorithms, target objects are searched for and located in a high-dimensional semantically enhanced point cloud, outputting the 3D bounding box information and category label for each target, achieving accurate target identification and localization. Then, through methods such as optical flow estimation or temporal difference, the motion trajectory of the target between consecutive frames is calculated, thereby deriving the target's velocity information for subsequent target tracking and behavior analysis in dynamic scenes. Finally, deep learning techniques can be used to generate pixel-level depth maps based on the high-dimensional semantically enhanced point cloud representation, accurately reflecting the distance information from each pixel in the scene to the camera.

[0047] Please see Figure 3 , Figure 3 Provided by the present invention Figure 2 A flowchart illustrating an embodiment of step S201. In some embodiments of the present invention, the raw optical signals and laser signals from the acquired environment are preprocessed to obtain multispectral image data and lidar point cloud data with different polarization states. The process also includes: S301. Based on the ambient light intensity, dynamically adjust the exposure parameters of each frequency band to acquire original multispectral images with different polarization states; S302. The initial point cloud collected is processed using an adaptive algorithm to improve the quality of the point cloud and obtain lidar point cloud data. S303. Synchronize the system time and perform internal and external parameter calibrations on the multispectral camera.

[0048] In the above embodiments, in complex and ever-changing environments, light intensity fluctuates significantly with time, weather, and scene changes. To ensure the high quality and accuracy of the acquired multispectral image data, the ambient light intensity is sensed in real time, and the exposure parameters of each frequency band are dynamically adjusted based on the light intensity data. In strong light environments, the exposure time is appropriately reduced to prevent overexposure and loss of detail; while in low light conditions, the exposure time is extended to increase the amount of light entering the image, ensuring that the image remains clear and discernible. Through this dynamic adjustment mechanism, the multispectral polarization camera group can acquire original multispectral images with rich detail, accurate color, and complete information in each frequency band under different lighting conditions.

[0049] When acquiring initial point cloud data, lidar systems are often affected by various factors, such as environmental noise, equipment errors, and target surface characteristics. This leads to problems like noisy points, outliers, and uneven density in the point cloud data, severely impacting its quality. To address these issues, statistical analysis is performed on the point cloud data to identify and locate noise and outliers. Based on the local geometric features and density distribution of the point cloud, intelligent filtering techniques are used to remove noise, while outliers are appropriately corrected or removed. Furthermore, the density distribution of the point cloud needs to be adaptively adjusted according to the complexity and importance of the target surface, retaining more point cloud data in key areas to improve the accuracy of target feature representation. Ultimately, this results in clear, accurate, and reasonably dense lidar point cloud data.

[0050] By using a high-precision time synchronization protocol, the multispectral polarization camera group and LiDAR are connected to a unified clock source, ensuring time consistency of all devices during data acquisition, eliminating time delays and errors between different devices, and ensuring strict alignment of multispectral images and LiDAR point clouds in the time dimension.

[0051] Please see Figure 4 , Figure 4 Provided by the present invention Figure 2 A flowchart illustrating an embodiment of step S202. In some embodiments of the present invention, feature extraction and fusion enhancement processing are performed on multispectral image data and lidar point cloud data to obtain a high-dimensional semantically enhanced point cloud representation that includes image semantics and point cloud geometry. The process also includes: S401. Use a four-branch feature extraction network to extract feature maps of each band from the original multispectral image and then perform weighted fusion to obtain multispectral image features. S402. Extract local descriptors from the lidar point cloud data through a point cloud feature extraction network and encode them to output lidar point cloud features. S403. Project the lidar point cloud onto the image plane to establish a correspondence between pixel level and point cloud level, and map the multispectral image features to the point cloud space through back projection, and stitch and fuse them with the point cloud geometric features to obtain a fused feature map.

[0052] In the above embodiment, the four-branch feature extraction network consists of four parallel branches, each dedicated to feature extraction for a specific band in the multispectral image. Within each branch, convolutional layers are first used to perform preliminary feature extraction on the input band image. Through the sliding operation of the convolutional kernels, local feature patterns in the image, such as edges and textures, are captured. As the number of network layers increases, each branch gradually extracts more abstract and semantic high-level features through a combination of multiple convolutional blocks and pooling layers, which can better represent the unique attributes of the target in different bands.

[0053] After feature extraction is completed in each of the four branches, a weighted fusion strategy is used to fuse the feature maps output by each branch. Specifically, different weight coefficients are assigned to the feature maps of each branch according to the importance of each band in the target recognition task. These weight coefficients can be determined through experiments and optimization algorithms, or specific values ​​can be set according to actual needs to ensure that useful information of each band is preserved to the greatest extent. Through a weighted summation operation, the feature maps of the four branches are fused into a comprehensive multispectral image feature map, which contains rich information about the target in different bands and highlights the features of key bands.

[0054] LiDAR point cloud data exists in the form of discrete three-dimensional points, describing the geometric shape and spatial location information of the target surface. A point cloud feature extraction network is used to extract meaningful features from this point cloud data. First, the input LiDAR point cloud is preprocessed, including point cloud downsampling and normalization operations, to reduce the number of points and reduce computational complexity, while preserving the main geometric structure of the point cloud. The point cloud data is transformed to a unified coordinate system and scale range to eliminate differences caused by different acquisition scenarios and equipment.

[0055] After preprocessing, a local neighborhood search algorithm is used to find the surrounding neighboring points of each point, forming a local region that reflects the geometric features of the target near that point. For the local neighborhood of each point, a multilayer perceptron (MLP) is used to extract local descriptors, converting the coordinate information of the neighboring points into discriminative feature vectors that can describe the geometric shape, curvature, and other attributes of the local region.

[0056] To further enhance the expressive power of features, the network encodes the extracted local descriptors, using an aggregation function to combine them into a global feature vector, or employing a more complex graph neural network structure to deeply fuse and encode the local descriptors. Through this encoding process, the network outputs LiDAR point cloud features to describe the target's three-dimensional geometric structure and spatial distribution information.

[0057] Based on the intrinsic and extrinsic parameters of the multispectral camera and lidar, the three-dimensional coordinates of each lidar point are converted into two-dimensional pixel coordinates on the image plane using the projection transformation formula. Each lidar point can find its corresponding pixel position in the multispectral image, thus establishing a one-to-one correspondence between the pixel level and the point cloud level.

[0058] After establishing the correspondence, the multispectral image features extracted in S401 are back-projected onto the point cloud space. Specifically, for the pixel location corresponding to each lidar point, the feature vector of that location is extracted from the multispectral image feature map and mapped to the spatial location of the lidar point. Each lidar point not only has its own geometric features but also has the multispectral image semantic features of the corresponding pixel location.

[0059] Finally, the mapped multispectral image features are concatenated and fused with the geometric features of the LiDAR point cloud extracted from S402. The fused feature vector contains both the multispectral semantic information and the three-dimensional geometric information of the target, resulting in a high-dimensional semantically enhanced point cloud representation that includes both image semantics and point cloud geometry, i.e., the fused feature map.

[0060] Please see Figure 5 , Figure 5 Provided by the present invention Figure 2 A flowchart illustrating an embodiment of step S203. In some embodiments of the present invention, multiple pairs of high-dimensional semantically enhanced point cloud representations are optimized to obtain three-dimensional target recognition results, target velocity information, and pixel-level depth maps. The method also includes: S501. Input the fused feature map into the 3D target detection network, output the 3D bounding box information and category label of each target, and output the instantaneous velocity and acceleration of each target based on the optical flow temporal modeling algorithm. S502. Based on depth map completion and image-guided point cloud interpolation techniques, the fused feature map is processed to output a pixel-level depth map; S503. Real-time evaluation of the input quality of each modality, dynamic generation of confidence weights based on preset indicators, and dynamic adjustment of each modality feature according to confidence.

[0061] In the above embodiments, the 3D object detection network adopts a deep learning architecture, combining the advantages of convolutional neural networks (CNN) and graph neural networks (GNN) to perform feature extraction and dimensionality reduction. By sliding convolutional kernels of different sizes on the feature map, it captures the feature patterns of the target at different scales, downsamples the feature map, reduces the amount of computation, and enhances the translation invariance of the features. Then, it generates a series of candidate regions containing the target on the feature map by means of a sliding window.

[0062] Each candidate region is further processed and classified. Through fully connected layers and a classifier, the network can determine whether a target exists in each candidate region and assign a category label to the target, such as vehicle, pedestrian, or building. Simultaneously, the network generates precise 3D bounding box information for each identified target, including the center coordinates, length, width, height, and rotation angle, thereby accurately locating the target's position and orientation in 3D space.

[0063] In addition to 3D bounding box information and category labels, an optical flow-based temporal modeling algorithm was employed to obtain target velocity information. This algorithm describes the motion information of pixels in the image over time. By analyzing the changes in optical flow between consecutive frames, the motion state of the target can be inferred. An optical flow estimation algorithm is used to calculate the optical flow field of the fused feature maps between adjacent frames. Each vector in the optical flow field represents the motion direction and velocity of the corresponding pixel. Then, by integrating and temporally analyzing the optical flow field, combined with the target's 3D bounding box information, the instantaneous velocity and acceleration of each target can be accurately calculated.

[0064] Although the fused feature map contains rich target information, the depth information may be incomplete or inaccurate, especially at the target edges and occluded areas. In order to obtain a more accurate and complete pixel-level depth map, depth map completion and image-guided point cloud interpolation techniques are used.

[0065] Depth map completion technology aims to repair and fill in missing or erroneous depth values ​​by utilizing existing depth information and image context information. It involves preliminary analysis of the depth information in the fused feature map to identify regions with incomplete depth. By constructing local and global context models of the depth map and combining information such as image texture, color, and edges, the missing depth values ​​are predicted and completed.

[0066] Image-guided point cloud interpolation further utilizes information from multispectral images to guide the point cloud interpolation process, thereby improving the accuracy of the depth map. It uses multispectral image information from the fused feature map as a guide, combined with the geometric information of the LiDAR point cloud, to perform fine-grained interpolation processing on the point cloud. Specifically, based on edge and texture information in the image, the direction and density of point cloud interpolation are determined, allowing the interpolated point cloud to better conform to the actual shape and surface features of the target. By projecting the interpolated point cloud onto the image plane, a pixel-level depth map is generated.

[0067] In practical applications, input data from different modalities, such as multispectral images and LiDAR point clouds, may be affected by various factors, such as lighting conditions, weather conditions, and equipment malfunctions, leading to differences in data quality. To fully utilize the advantages of each modal data and improve robustness and accuracy, it is necessary to evaluate the input quality of each modality in real time and dynamically adjust the weights of each modal feature based on the evaluation results.

[0068] For multispectral images, preset metrics include, but are not limited to, image sharpness, contrast, signal-to-noise ratio, and illumination uniformity; for lidar point clouds, preset metrics include, but are not limited to, point cloud density, integrity, noise level, and reflection intensity. By monitoring changes in these metrics in real time, the quality status of each modality of data can be understood promptly.

[0069] Then, based on the values ​​of preset indicators, a dynamic weighting algorithm is used to generate confidence weights for each modality feature. Taking into account the importance and interrelationships of each indicator, a comprehensive confidence score for each modality is calculated through weighted summation or other nonlinear combinations. The higher the confidence score, the better the data quality of that modality, and the greater the weight of its features should be in subsequent processing.

[0070] Each modal feature is multiplied by its corresponding confidence weight to obtain a weighted feature vector. This weighted feature vector is then fused and processed to give higher weight to high-quality modal features in the final result, while low-quality modal features are appropriately suppressed. This approach fully leverages the advantages of each modal data, improving the accuracy and robustness of tasks such as 3D target recognition, velocity estimation, and depth map generation, and adapting to complex environmental changes in different scenarios.

[0071] Please see Figure 6 , Figure 6 Provided by the present invention Figure 3 A flowchart illustrating one embodiment of step S303. In some embodiments of the present invention, the system time is synchronized, and the intrinsic and extrinsic parameters of the multispectral camera are calibrated separately. The method further includes: S601: Connect the multispectral polarization camera group and the lidar to the same control clock via an external trigger signal, and use a unified frequency to trigger data acquisition; S602. Independently calibrate a multispectral camera with several linearly polarizing filters at different angles using a calibration plate to determine the camera's intrinsic parameters. S603. Calculate the relative pose matrix between several multispectral cameras and the pose transformation matrix between the multispectral main camera and the lidar based on the preset calibration structure and calibration tools.

[0072] In the above embodiments, each camera in the multispectral polarization camera group and the lidar are connected to an external trigger signal distributor via a synchronization signal interface. This ensures the quality of the signal transmission line, reduces signal attenuation and interference, and guarantees that the trigger signal can accurately reach each sensor. The frequency of the external trigger signal needs to be determined based on the performance specifications of the multispectral camera and lidar, as well as the requirements of the actual application scenario. For example, if high-precision target detection and tracking are required, a higher trigger frequency is needed to obtain denser data; while in scenarios where real-time requirements are relatively low and data quality is more important, the trigger frequency can be appropriately reduced.

[0073] After completing the hardware connection and parameter settings, the waveform and timing of the trigger signal are monitored during system integration to ensure that the multispectral polarization camera group and the lidar can start data acquisition at the same trigger time and keep the acquisition frequency consistent, so that the acquired multispectral images and lidar point cloud data have the same timestamp.

[0074] Calibration is typically performed using a calibration plate with high-precision feature points, such as a checkerboard calibration plate. It's crucial to ensure the calibration plate is made of uniform material and has a smooth surface to prevent distortion during imaging. For each multispectral camera equipped with a linear polarizing filter at a different angle, an independent calibration experimental environment is set up. The calibration plate is placed within the camera's effective field of view, ensuring it is perpendicular to the camera's optical axis.

[0075] Control the camera to take multiple sets of pictures of the calibration board in different poses. During the shooting process, the pose of the calibration board should be changed slowly and smoothly to obtain a sufficient number of calibration images. The number of images in each set should be determined according to the requirements of the calibration algorithm, generally no less than 15-20 sets.

[0076] The captured calibration images are processed using professional camera calibration algorithms (such as the Zhang Zhengyou calibration method). By detecting the position of feature points on the calibration board in the image and combining the actual size information of the calibration board, a camera imaging model is established, and the camera's intrinsic parameter matrix, including focal length, principal point coordinates, and distortion coefficients, such as radial distortion coefficient and tangential distortion coefficient, is solved.

[0077] A pre-defined calibration structure should be constructed, with clearly defined feature points or markers to facilitate identification and measurement by the camera and LiDAR. For relative pose calibration between multispectral cameras, high-precision 3D measuring equipment can be used to measure the coordinates of feature points on the calibration structure in the camera coordinate system. For pose calibration between the multispectral main camera and LiDAR, joint matching can be performed using LiDAR point cloud data and camera image data.

[0078] When calibrating the relative pose between multispectral cameras, the three-dimensional coordinates of feature points on the calibration structure in each camera coordinate system are first measured using a three-dimensional measurement device. Then, the relative rotation matrix and translation vector between different cameras are solved using mathematical methods such as the least squares method, thus obtaining the relative pose matrix.

[0079] For the pose transformation matrix calibration between the multispectral main camera and the LiDAR, the calibration structure point cloud data acquired by the LiDAR is matched with the calibration structure image captured by the camera. Feature point matching algorithms (such as SIFT, SURF, etc.) can be used to extract feature points from the image and point cloud. Then, through optimization methods such as the iterative nearest-point algorithm, the pose parameters between the camera and the LiDAR are continuously adjusted to minimize the matching error between image feature points and point cloud feature points. Finally, the accurate pose transformation matrix between the multispectral main camera and the LiDAR is obtained.

[0080] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0081] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0082] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0083] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0084] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0085] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of embodiments of this disclosure upon considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

[0086] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0087] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A target recognition system of laser radar and multi-spectral polarization imaging fusion, characterized in that, Comprise: The sensor configuration and data preprocessing module, the cross-modal feature extraction and fusion module, the multi-task output and optimization module are connected in turn; Among them, the sensor configuration and data preprocessing module is used for spatio-temporal synchronous processing of the collected original optical signal and laser signal in the environment, to obtain multispectral image data of different polarization states[1] and laser radar point cloud data; The cross-modal feature extraction and fusion module is used for feature extraction and fusion enhancement processing of multispectral image data and laser radar point cloud data, and outputs high-dimensional semantic enhanced point cloud representation containing image semantics and point cloud geometry; The multi-task output and optimization module is used for processing the high-dimensional semantic enhanced point cloud representation, and outputs three-dimensional target recognition result, target speed information and pixel-level depth map.

2. The laser radar and multi-spectral polarimetric imaging fused target recognition system of claim 1, wherein, The sensor configuration and data preprocessing module further comprises a multispectral polarization camera group, a laser radar, a spatio-temporal synchronization and calibration unit; Among them, the multispectral polarization camera group is configured with several multispectral cameras with different angle linear polarization filters, which dynamically adjusts the exposure parameters of each frequency band based on the intensity of environmental light to collect original multispectral images; The laser radar collects point cloud data and improves the quality of point cloud data through adaptive algorithm processing, to obtain laser radar point cloud data; The spatio-temporal synchronization and calibration unit is used for time synchronization, internal parameter calibration and external parameter calibration of the system.

3. The laser radar and multi-spectral polarimetric imaging fused target recognition system of claim 2, wherein, The cross-modal feature extraction and fusion module further comprises a multispectral polarization feature extraction unit, a laser radar feature extraction unit and a cross-modal feature conversion and fusion unit; Among them, the multispectral polarization feature extraction unit is used for extracting feature maps of each waveband from the original multispectral image respectively by using a four-branch feature extraction network to obtain multispectral image features by weighted fusion; The laser radar feature extraction unit extracts local descriptors from the laser radar point cloud data by a point cloud feature extraction network and encodes them to output laser radar point cloud features; The cross-modal feature conversion and fusion unit projects the laser radar point cloud to the image plane to establish the corresponding relationship between pixel level and point cloud level, and maps the multispectral image features to the point cloud space by back projection, and splices and fuses the multispectral image features with the point cloud geometry features to obtain a fusion feature map.

4. The laser radar and multi-spectral polarimetric imaging fused target recognition system of claim 3, wherein, The multi-task output and optimization module further comprises a target recognition and speed estimation unit, a point cloud generation unit and a all-weather robustness enhancement unit; Among them, the target recognition and speed estimation unit inputs the fusion feature map into a three-dimensional target detection network to output three-dimensional bounding box information and class labels of each target, and outputs the instantaneous speed and acceleration of each target based on a time sequence modeling algorithm of optical flow; The point cloud generation unit processes the fusion feature map based on depth map completion and image guided point cloud interpolation technology to output a pixel-level depth map; The all-weather robustness enhancement unit is used for real-time evaluation of the input quality of each modality, and dynamically generates confidence weight according to the preset index, and dynamically adjusts each modality feature according to the confidence.

5. The laser radar and multi-spectral polarimetric imaging fused target recognition system of claim 2, wherein, The spatio-temporal synchronization and calibration unit further comprises a time synchronization subunit, an internal parameter calibration subunit and an external parameter calibration subunit; The time synchronization subunit connects the multi-spectral polarization camera group and the laser radar to the same control clock through an external trigger signal, and triggers data acquisition with a unified frequency; The intrinsic parameter calibration subunit independently calibrates the multi-spectral cameras of a plurality of linear polarization filters with different angles through a calibration board, and determines the camera intrinsic parameters; The extrinsic parameter calibration subunit calculates the relative pose matrix between a plurality of multi-spectral cameras and the pose transformation matrix between the multi-spectral main camera and the laser radar based on a preset calibration structure and a calibration tool.

6. A target recognition method of laser radar and multispectral polarimetric imaging fusion, based on the target recognition system of laser radar and multispectral polarimetric imaging fusion according to any one of claims 1-5, comprising a multispectral polarimetric camera group, a laser radar, characterized in that, The method comprises: spatially and temporally synchronizing the collected original optical signals and laser signals in the environment to obtain multi-spectral image data of different polarization states and laser radar point cloud data; performing feature extraction and fusion enhancement processing on the multi-spectral image data and the laser radar point cloud data to obtain a high-dimensional semantic enhanced point cloud representation containing image semantics and point cloud geometry; optimizing the high-dimensional semantic enhanced point cloud representation to obtain a three-dimensional target recognition result, target speed information, and a pixel-level depth map.

7. The method of claim 6, wherein the laser radar and multi-spectral polarimetric imaging fused target recognition method is characterized by, The spatial and temporal synchronization of the collected original optical signals and laser signals in the environment to obtain multi-spectral image data of different polarization states and laser radar point cloud data further comprises: adjusting the exposure parameters of each frequency band based on the intensity of the ambient light to collect original multi-spectral images of different polarization states; processing the collected initial point cloud through an adaptive algorithm to improve the quality of the point cloud and obtain laser radar point cloud data; synchronizing the system time and performing intrinsic parameter calibration and extrinsic parameter calibration on the multi-spectral cameras.

8. The laser radar and multi-spectral polarimetric imaging fused target recognition system of claim 7, wherein, The feature extraction and fusion enhancement processing on the multi-spectral image data and the laser radar point cloud data to obtain a high-dimensional semantic enhanced point cloud representation further comprises: extracting feature maps of each waveband from the original multi-spectral images using a four-branch feature extraction network to obtain a multi-spectral image feature through weighted fusion; extracting local descriptors from the laser radar point cloud data through a point cloud feature extraction network and encoding them to output laser radar point cloud features; projecting the laser radar point cloud onto the image plane to establish a correspondence between the pixel level and the point cloud level, and mapping the multi-spectral image features to the point cloud space through back projection to splice and fuse them with the point cloud geometry features to obtain a fused feature map.

9. The laser radar and multi-spectral polarimetric imaging fused target recognition system of claim 8, wherein, The optimization processing on the high-dimensional semantic enhanced point cloud representation to obtain a three-dimensional target recognition result, target speed information, and a pixel-level depth map further comprises: inputting the fused feature map into a three-dimensional target detection network to output three-dimensional bounding box information and class labels for each target, and outputting the instantaneous speed and acceleration of each target based on a time sequence modeling algorithm of optical flow; processing the fused feature map based on depth map completion and image-guided point cloud interpolation technology to output a pixel-level depth map; real-time evaluation of the input quality of each modality, dynamic generation of confidence weights based on preset indicators, and dynamic adjustment of the features of each modality according to the confidence.

10. The laser radar and multi-spectral polarimetric imaging fused target recognition system of claim 7, wherein, The synchronization of the system time and the intrinsic parameter calibration and extrinsic parameter calibration of the multi-spectral cameras further comprises: The multispectral polarization camera group is connected to the same control clock as the laser radar through an external trigger signal, and data acquisition is triggered at a unified frequency; The multispectral cameras with a plurality of linear polarization filters having different angles are independently calibrated by a calibration board to determine the camera intrinsic parameters; The relative pose matrix among the plurality of multispectral cameras and the pose transformation matrix between the multispectral main camera and the laser radar are calculated based on a preset calibration structure and a calibration tool.

Citation Information

Cited By

  • Road surface micro texture depth evaluation method and system based on phase scanning

    CN121473209A

  • Water surface floating object identification method and system fusing vision and laser radar

    CN121861368A

  • Multi-modal deep learning identification method and system for slope rock mass structural surface under complex terrain and medium

    CN122020130A

  • Slope rock mass structure plane multi-modal deep learning identification method and system under complex terrain and medium

    CN122020130B