Multi-sensor fusion device for point cloud data acquisition and calibration method thereof
Through the multi-sensor fusion device, combined with the advantages of lidar, camera and millimeter wave radar, the limitations of a single sensor in point cloud data acquisition are solved, and high-precision and all-round point cloud data acquisition and fusion are achieved, providing strong support for the fields of autonomous driving and robot navigation.
Patent Information
- Application Number
- CN202510262540.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
AI Technical Summary
A single sensor has limitations in point cloud data acquisition. For example, lidar cannot obtain object texture information, cameras have insufficient depth measurement accuracy, and millimeter-wave radars have poor detection of stationary objects, resulting in environmental perception deviations and data accuracy problems.
A multi-sensor fusion device is adopted, including lidar, camera and millimeter wave radar. Each sensor data is received and stored through the data acquisition module, and the calculation and processing module performs data fusion processing. Finally, the result generation module performs three-dimensional rendering to generate a three-dimensional visual model of the target object.
By combining the advantages of different types of sensors, we can learn from our strengths and weaknesses, and achieve high-precision and all-round point cloud data acquisition and fusion, improving data accuracy and reliability, and providing strong support for the fields of three-dimensional modeling, autonomous driving and robot navigation.
Smart Images

Figure CN120107737A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud data acquisition, and more particularly to a multi-sensor fusion device for point cloud data acquisition and a calibration method thereof. Background Art
[0002] In today's era of rapid technological development, point cloud data has become an indispensable key element in many cutting-edge fields. In the field of 3D modeling, point cloud data is like the cornerstone of building a virtual world. It can accurately depict the shape and spatial relationship of objects and scenes, making the virtual model more compatible with the real world, and providing a high-quality model foundation for application scenarios such as architectural design, cultural relics protection, and game development. In the field of autonomous driving, point cloud data is like the "eyes" of the car, helping the vehicle perceive the surrounding environment, including road conditions, other vehicles, pedestrians, and the location information of various obstacles, which is an important guarantee for achieving safe and intelligent driving. For robot navigation, point cloud data is the "compass" of the robot's actions, guiding the robot to locate and plan paths in complex environments, so as to accurately complete various tasks. Whether it is a logistics robot on an industrial production line or an intelligent robot in home services, it relies on point cloud data to achieve efficient navigation.
[0003] However, the process of collecting point cloud data faces many challenges, one of which is the limitation of a single sensor. Take LiDAR as an example. As a commonly used sensor for collecting point cloud data, it has excellent performance in measuring distance and obtaining object contours. However, LiDAR has obvious deficiencies in obtaining object texture information. The texture information of an object is of great significance for identifying the type and material of the object. However, LiDAR can only provide geometric shape information on the surface of an object, and cannot present rich texture details such as color and pattern on the surface of the object. This limits its ability to obtain comprehensive information about the object to a certain extent.
[0004] Cameras are also one of the important sensors for point cloud data collection, and they play a vital role in identifying objects and understanding scenes. However, cameras have limitations in depth measurement accuracy. Although some algorithms can be used to estimate depth information from camera images, the accuracy of this method is much lower than that of professional ranging equipment. In complex environments, cameras may cause errors in depth information due to factors such as light changes and object occlusion, which in turn affects the accuracy and reliability of point cloud data.
[0005] Millimeter-wave radar, as another sensor, has shown certain advantages in adverse weather conditions, but it is not very effective in detecting stationary objects. In application scenarios such as autonomous driving and robot navigation, accurate detection of stationary objects is crucial, and the shortcomings of millimeter-wave radar in this regard may lead to deviations in environmental perception, affecting the decision-making and safety performance of the entire system.
[0006] Therefore, how to provide a multi-sensor fusion device and a calibration method for point cloud data acquisition, combine the advantages of different types of sensors, learn from each other's strengths and weaknesses, and effectively acquire and fuse point cloud data is a problem that technical personnel in this field urgently need to solve. Summary of the invention
[0007] In view of this, the present invention provides a multi-sensor fusion device for point cloud data acquisition and a calibration method thereof, which combine the advantages of different types of sensors, learn from each other's strengths and weaknesses, effectively acquire and fuse point cloud data, and minimize errors by continuously adjusting parameters, thereby achieving precise calibration and ensuring the accuracy of data fusion, providing strong support for the further development of three-dimensional modeling, autonomous driving, robot navigation and other fields.
[0008] In order to achieve the above object, the present invention adopts the following technical solution: a multi-sensor fusion device for point cloud data acquisition, comprising:
[0009] LiDAR sensor, camera sensor, millimeter wave radar sensor, data acquisition module, calculation processing module and result generation module;
[0010] The laser radar sensor is used to obtain three-dimensional spatial information of the target object and construct point cloud data by emitting a laser beam and receiving reflected light;
[0011] The camera sensor is used to capture visual information of the target object;
[0012] The millimeter wave radar sensor is used to detect the distance and speed information of the target object;
[0013] The data acquisition module is responsible for receiving and storing data from the laser radar sensor, the camera sensor and the millimeter wave radar sensor;
[0014] The computing and processing module is used to perform fusion processing on the data received from the laser radar sensor, the camera sensor and the millimeter wave radar sensor through a fusion algorithm;
[0015] The result generation module performs three-dimensional rendering on the fused data to obtain a three-dimensional visualization model of the target object.
[0016] Preferably, the laser radar sensor, the camera sensor and the millimeter wave radar sensor are located in the same plane.
[0017] Preferably, the calculation processing module includes:
[0018] The first computing unit is used to input the image collected by the camera sensor into the knowledge distillation target recognition model of the multi-stage attention mechanism, output the target object image with the region frame in sequence, and calculate the image coordinates of the center of the target object;
[0019] The second calculation unit is used to convert the image coordinates of the center of the target object into the space coordinates P in the world coordinate system. cam , and calculate the distance d from the center of the target object to the plane where the camera sensor is located cam ;
[0020] The third computing unit is used to obtain the spatial coordinates P of the center of the target object through the point cloud image of the laser radar sensor las , and the distance d between the center of the target object and the plane where the lidar sensor is located las ;
[0021] The fourth calculation unit is used to obtain the distance d according to the millimeter wave radar sensor. mil and speed v mil , calculate the distance deviation d 1 =|d las -d cam |,d 2 =|d mil -d cam |,d 3 =|d las -d mil |;
[0022] The judgment unit is used to determine if the distance deviation d 1 d 2 d 3 Any two of them are greater than or equal to the threshold d p , then re-collect data, if the distance deviation is d 1 d 2 d 3 Any two of them are less than the threshold d p , then the center coordinates of the target object are returned.
[0023] Preferably, the knowledge distillation target recognition model of the multi-stage attention mechanism includes:
[0024] A feature extraction unit, used to detect the image collected by the camera sensor through a feature extraction model to obtain a feature image;
[0025] The visual recognition unit is used to input the feature image into the knowledge distillation target recognition model of the multi-stage attention mechanism, use the teacher network distillation to train the student network, perform fine-grained recognition based on the student network, and obtain the visual recognition result of the current target object;
[0026] The trigger unit is used to determine whether the target object has reached the specified position based on the obtained limit switch detection signal. If so, it triggers the lidar sensor to obtain point cloud data to obtain the three-dimensional spatial information of the current vehicle body, and at the same time triggers the millimeter wave radar sensor to obtain the distance and speed information of the target object.
[0027] Preferably, the teacher network is a fine-grained target recognition network with a multi-stage attention mechanism with ResNet50 as the backbone network, which extracts feature maps of the three stages Conv3_x, Conv4_x, and Conv5_x respectively, and additionally adds an attention mechanism layer, a calibration layer, and a classification layer. The calibration layer maps features of different channels and sizes to specified channels and sizes through convolution operations, and fuses the output features of the three stages. The classifier is composed of a fully connected layer, which maps the feature map into an output category vector.
[0028] Preferably, the student network is based on the MobinetV3 network as the backbone network, and the MobinetV3 network is pre-designed in stages. The output feature maps of the three stages Bneck6, Bneck11, and Bneck15 are extracted respectively, and an attention layer, a calibration layer, and a classification layer are additionally added. The output value of each stage of the student network is obtained through the feature mapping function and network parameters of the student network.
[0029] Preferably, the feature extraction model adopts the yolov8 network structure.
[0030] Preferably, during the rendering process, the visualization tool uses OpenGL, Unity3D or Unreal Engine.
[0031] Preferably, a calibration method for a multi-sensor fusion device for point cloud data acquisition is performed at a known calibration site or using a standard calibration object;
[0032] Install the LiDAR sensor, camera sensor, and millimeter-wave radar sensor in the predetermined positions and ensure that their relative positions and angles are constant;
[0033] Simultaneously start the LiDAR sensor, camera sensor, and millimeter wave radar sensor to collect point cloud data, image data, and distance and speed information including the benchmark;
[0034] The image collected by the camera sensor is input into the knowledge distillation target recognition model with a multi-stage attention mechanism, and the target object image with a region frame is output in sequence, and the image coordinates of the center of the target object are calculated;
[0035] Convert the image coordinates of the center of the target object into the spatial coordinates P in the world coordinate system cam , and calculate the distance d from the center of the target object to the plane where the camera sensor is located cam ;
[0036] The spatial coordinates P of the center of the target object are obtained through the point cloud image of the lidar sensor las , and the distance d between the center of the target object and the plane where the lidar sensor is located las ;
[0037] According to the distance d obtained by the millimeter wave radar sensor mil and speed v mil , calculate the distance deviation d 1 =|d las -d cam |,d 2 =|d mil -d cam |,d 3 =|d las -d mil |;
[0038] If the distance deviation d 1 d 2 d 3 Any two of them are greater than or equal to the threshold d p , then re-collect data, if the distance deviation is d 1 d 2 d 3 Any two of them are less than the threshold d p , then the center coordinates of the target object are returned.
[0039] Preferably, the knowledge distillation target recognition model of the multi-stage attention mechanism includes:
[0040] Detecting the image collected by the camera sensor through a feature extraction model to obtain a feature image;
[0041] The feature image is input into the knowledge distillation target recognition model of the multi-stage attention mechanism, the teacher network is used for distillation to train the student network, and fine-grained recognition is performed based on the student network to obtain the visual recognition result of the current target object;
[0042] According to the limit switch detection signal obtained, it is judged whether the target object has reached the specified position. If so, the lidar sensor is triggered to obtain point cloud data to obtain the three-dimensional spatial information of the current vehicle body, and the millimeter wave radar sensor is triggered to obtain the distance and speed information of the target object.
[0043] It can be known from the above technical solutions that compared with the prior art, the present invention discloses a multi-sensor fusion device for point cloud data acquisition and a calibration method thereof, including: a laser radar sensor, a camera sensor, a millimeter wave radar sensor, a data acquisition module, a calculation processing module and a result generation module; the laser radar sensor is used to obtain the three-dimensional spatial information of the target object with high precision, and constructs point cloud data by emitting a laser beam and receiving reflected light; the camera sensor is used to capture visual information such as color and texture of the target object; the visual understanding of the target object is enhanced based on the camera sensor; the millimeter wave radar sensor is used to detect the distance and speed information of the target object; the data acquisition module is responsible for receiving and storing the data of the laser radar sensor, the camera sensor and the millimeter wave radar sensor; the calculation processing module is used to fuse the received data of the laser radar sensor, the camera sensor and the millimeter wave radar sensor through a fusion algorithm; the result generation module performs three-dimensional rendering on the fused data to obtain a three-dimensional visualization model of the target object. It is convenient for subsequent analysis, monitoring and decision-making. The present invention combines the advantages of different types of sensors, takes advantage of their strengths and makes up for their weaknesses, effectively acquires and fuses point cloud data, and simultaneously minimizes errors by continuously adjusting parameters, thereby achieving precise calibration and ensuring the accuracy of data fusion, providing strong support for the further development of 3D modeling, autonomous driving, robot navigation and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0045] Figure 1 A schematic diagram of the structure of a multi-sensor fusion device for point cloud data acquisition provided by the present invention.
[0046] Figure 2 Schematic diagram of the knowledge distillation target recognition model structure of the stage attention mechanism provided by the present invention. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] The embodiment of the present invention discloses a multi-sensor fusion device for point cloud data acquisition, such as Figure 1 As shown, including:
[0049] LiDAR sensor, camera sensor, millimeter wave radar sensor, data acquisition module, calculation processing module and result generation module;
[0050] The laser radar sensor is used to obtain the three-dimensional spatial information of the target object with high precision, and constructs point cloud data by emitting laser beams and receiving reflected light; based on the laser radar sensor, detailed spatial structure information of the target object can be provided;
[0051] The camera sensor is used to capture visual information such as color and texture of the target object; based on the camera sensor, the visual understanding of the target object is enhanced;
[0052] The millimeter wave radar sensor is used to detect the distance and speed information of the target object; especially in adverse weather conditions (such as rain, fog, snow, etc.), the millimeter wave radar sensor still has good performance and ensures the reliability of the data;
[0053] The data acquisition module is responsible for receiving and storing data from the laser radar sensor, camera sensor and millimeter wave radar sensor; the data acquisition module has a high-speed data transmission interface and sufficient data caching capacity, and can cope with the real-time processing requirements of high data volumes;
[0054] The computing and processing module is used to perform fusion processing on the data received from the laser radar sensor, the camera sensor and the millimeter wave radar sensor through a fusion algorithm; improve the accuracy and completeness of the target object information obtained; the data of the laser radar sensor, the camera sensor and the millimeter wave radar sensor are synchronized in time and space;
[0055] The result generation module performs three-dimensional rendering on the fused data to obtain a three-dimensional visualization model of the target object, which is convenient for subsequent analysis, monitoring and decision-making.
[0056] Specifically, the laser radar sensor, the camera sensor and the millimeter wave radar sensor are located in the same plane.
[0057] Specifically, the calculation processing module includes:
[0058] The first computing unit is used to input the image collected by the camera sensor into the knowledge distillation target recognition model of the multi-stage attention mechanism, output the target object image with the region frame in sequence, and calculate the image coordinates of the center of the target object;
[0059] The second calculation unit is used to convert the image coordinates of the center of the target object into the space coordinates P in the world coordinate system. cam , and calculate the distance d from the center of the target object to the plane where the camera sensor is located cam ;
[0060] The third computing unit is used to obtain the spatial coordinates P of the center of the target object through the point cloud image of the laser radar sensor las , and the distance d between the center of the target object and the plane where the lidar sensor is located las ;
[0061] The fourth calculation unit is used to obtain the distance d according to the millimeter wave radar sensor. mil and speed v mil , calculate the distance deviation d 1 =|d las -d cam |,d 2 =|d mil -d cam |,d 3 =|d las -d mil |;
[0062] The judgment unit is used to determine if the distance deviation d 1 d 2 d 3 Any two of them are greater than or equal to the threshold d p , then re-collect data, if the distance deviation is d 1 d 2 d 3 Any two of them are less than the threshold d p , then the center coordinates of the target object are returned.
[0063] The center coordinates of the target object are: (λ cam P cam +λ las P las ),
[0064]
[0065] Among them, α is the weight adjustment factor, P las is the spatial coordinate of the center of the target object obtained by the lidar sensor, P camThe spatial coordinates of the center of the target space obtained by the camera sensor.
[0066] Specifically, the knowledge distillation target recognition model of the multi-stage attention mechanism is as follows: Figure 2 As shown, including:
[0067] A feature extraction unit is used to detect the image collected by the camera sensor through a feature extraction model to obtain a feature image; compare the coordinate points of the detected feature image with the recognition area, and if the coordinate points of the feature image are within the recognition area, intercept the feature image according to the coordinates;
[0068] The visual recognition unit is used to input the feature image into the knowledge distillation target recognition model of the multi-stage attention mechanism, use the teacher network distillation to train the student network, perform fine-grained recognition based on the student network, and obtain the visual recognition result of the current target object;
[0069] The trigger unit is used to determine whether the target object has reached the specified position based on the obtained limit switch detection signal. If so, it triggers the lidar sensor to obtain point cloud data to obtain the three-dimensional spatial information of the current vehicle body, and at the same time triggers the millimeter wave radar sensor to obtain the distance and speed information of the target object.
[0070] Specifically, the teacher network is a fine-grained target recognition network with a multi-stage attention mechanism with ResNet50 as the backbone network. The feature maps of the three stages Conv3_x, Conv4_x, and Conv5_x are extracted respectively, and an attention mechanism layer, a calibration layer, and a classification layer are additionally added. The calibration layer maps features of different channels and different sizes to specified channels and sizes through convolution operations, and fuses the output features of the three stages. The classifier is composed of a fully connected layer, which maps the feature map into an output category vector.
[0071] Specifically, the student network uses the MobinetV3 network as the backbone network, and pre-defines the MobinetV3 network in stages. The output feature maps of the three stages Bneck6, Bneck11, and Bneck15 are extracted respectively, and an attention layer, a calibration layer, and a classification layer are additionally added. The output value of each stage of the student network is obtained through the feature mapping function and network parameters of the student network.
[0072] Specifically, the feature extraction model adopts the yolov8 network structure.
[0073] Specifically, during the rendering process, the visualization tool uses OpenGL, Unity3D or Unreal Engine. GPU acceleration technology is used to achieve real-time rendering, improve rendering efficiency, and allow users to rotate, zoom and measure for more in-depth analysis.
[0074] The multi-sensor fusion device described in the embodiment of the present invention can be applied to fields such as autonomous driving, intelligent monitoring, robot navigation and environmental modeling.
[0075] In a specific embodiment of the present invention, the laser radar sensor is based on the principle of laser ranging. It emits a laser beam to the target object and accurately measures the time for the reflected light to return, thereby calculating the distance between the sensor and each point of the target object. This high-precision distance measurement capability enables it to construct point cloud data with high density and good accuracy. Moreover, the laser radar has various scanning modes, including rotary and solid-state types. The appropriate scanning mode can be selected according to different application scenarios to cover the required monitoring range. For example, in the application of self-driving cars, multi-line rotating laser radars are often used to fully perceive the environment around the vehicle, including road boundaries, other vehicles and pedestrians.
[0076] Camera sensors can capture details such as color and texture of target objects. This visual information is crucial for accurately identifying the type and attributes of objects. In practical applications, cameras can be divided into different types such as visible light cameras and infrared cameras. Visible light cameras can obtain high-quality color images in well-lit environments, but the effect may be limited in low-light or dark environments. Infrared cameras can work at night or in low-visibility conditions, and form images by detecting the thermal radiation of objects, providing additional visual information for the entire system and enhancing the system's adaptability in complex environments.
[0077] The operating frequency band of the millimeter wave radar sensor is in the millimeter wave range, with a higher frequency and a shorter wavelength, which makes it have better penetration than other sensors in severe weather conditions (such as rain, fog, and snow). It can accurately measure the distance and relative speed of the target object and can track multiple targets at the same time. In the field of autonomous driving, millimeter wave radar is used in conjunction with lidar, cameras, etc. When the lidar or camera is affected by severe weather, the millimeter wave radar can still reliably detect the dynamic information of nearby objects, providing an important basis for the safe driving of the vehicle.
[0078] The data acquisition module is equipped with a high-speed data transmission interface to meet the needs of fast transmission of large amounts of data generated by sensors such as laser radar, cameras, and millimeter-wave radar. For example, a high-resolution camera may generate tens or even hundreds of megabytes of data per second, and the amount of data from laser radars at high speed scanning is also considerable, and millimeter-wave radars also have continuous data output. At the same time, the data acquisition module has sufficient data caching capacity to cope with possible data transmission peaks or short data processing delays to ensure that data will not be lost. The data acquisition module uses high-performance storage chips or storage modules and is equipped with advanced data transmission protocols, such as high-speed Ethernet interfaces.
[0079] The computing and processing module receives data from the data acquisition module and integrates and processes the data through specific algorithms. In addition, the computing and processing module uses multi-core processors, GPU (graphics processing unit) acceleration and other technologies to improve data processing efficiency. It has strong computing capabilities and can meet application scenarios with high real-time requirements, such as real-time environmental perception and decision-making in autonomous driving.
[0080] The result generation module performs three-dimensional rendering on the processed data to obtain a three-dimensional visualization model of the target object.
[0081] When working, the laser radar sensor, camera sensor and millimeter wave radar sensor detect and collect data of the target object. The data generated by each of them is transmitted to the data acquisition module in real time through a high-speed data transmission link for temporary storage. The data acquisition module passes these data to the calculation and processing module in a certain format and order. After receiving the data, the calculation and processing module immediately starts the fusion algorithm to fuse the data of different sensors. After complex calculations and integration, the processed results are sent to the result generation module. The result generation module performs three-dimensional rendering based on the fused data, and finally generates a three-dimensional visualization model of the target object, completing the entire point cloud data collection and processing process. The multi-sensor fusion device of the embodiment of the present invention fully utilizes the advantages of different sensors through the close collaboration of various components, overcomes the limitations of a single sensor, and provides high-quality point cloud data and three-dimensional visualization models for many fields.
[0082] Specifically, a calibration method for a multi-sensor fusion device for point cloud data acquisition is performed at a known calibration site or using a standard calibration object (such as a chessboard, a calibration plate) to ensure that the sensor is fixed and does not move to reduce errors;
[0083] Install the LiDAR sensor, camera sensor, and millimeter-wave radar sensor in the predetermined positions and ensure that their relative positions and angles are constant;
[0084] Simultaneously start the LiDAR sensor, camera sensor, and millimeter wave radar sensor to collect point cloud data, image data, and distance and speed information including the benchmark;
[0085] The image collected by the camera sensor is input into the knowledge distillation target recognition model with a multi-stage attention mechanism, and the target object image with a region frame is output in sequence, and the image coordinates of the center of the target object are calculated;
[0086] Convert the image coordinates of the center of the target object into the spatial coordinates P in the world coordinate system cam , and calculate the distance d from the center of the target object to the plane where the camera sensor is located cam ;
[0087] The spatial coordinates P of the center of the target object are obtained through the point cloud image of the lidar sensor las , and the distance d between the center of the target object and the plane where the lidar sensor is located las ;
[0088] According to the distance d obtained by the millimeter wave radar sensor mil and speed v mil , calculate the distance deviation d 1 =|d las -d cam |,d 2 =|d mil -d cam |,d 3 =|d las -d mil |;
[0089] If the distance deviation d 1 d 2 d 3 Any two of them are greater than or equal to the threshold d p , then re-collect data, if the distance deviation is d 1 d 2 d 3 Any two of them are less than the threshold d p , then the center coordinates of the target object are returned.
[0090] Specifically, the knowledge distillation target recognition model of the multi-stage attention mechanism includes:
[0091] Detecting the image collected by the camera sensor through a feature extraction model to obtain a feature image;
[0092] The feature image is input into the knowledge distillation target recognition model of the multi-stage attention mechanism, the teacher network is used for distillation to train the student network, and fine-grained recognition is performed based on the student network to obtain the visual recognition result of the current target object;
[0093] According to the limit switch detection signal obtained, it is judged whether the target object has reached the specified position. If so, the lidar sensor is triggered to obtain point cloud data to obtain the three-dimensional spatial information of the current vehicle body, and the millimeter wave radar sensor is triggered to obtain the distance and speed information of the target object.
[0094] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0095] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-sensor fusion device for point cloud data acquisition, characterized in that: include: LiDAR sensor, camera sensor, millimeter wave radar sensor, data acquisition module, calculation processing module and result generation module; The laser radar sensor is used to obtain three-dimensional spatial information of the target object and construct point cloud data by emitting a laser beam and receiving reflected light; The camera sensor is used to capture visual information of the target object; The millimeter wave radar sensor is used to detect the distance and speed information of the target object; The data acquisition module is responsible for receiving and storing data from the laser radar sensor, the camera sensor and the millimeter wave radar sensor; The computing and processing module is used to perform fusion processing on the data received from the laser radar sensor, the camera sensor and the millimeter wave radar sensor through a fusion algorithm; The result generation module performs three-dimensional rendering on the fused data to obtain a three-dimensional visualization model of the target object.
2. The multi-sensor fusion device for point cloud data acquisition according to claim 1, characterized in that: The laser radar sensor, the camera sensor and the millimeter wave radar sensor are located on the same plane.
3. The multi-sensor fusion device for point cloud data acquisition according to claim 2, characterized in that: The computing and processing module comprises: The first computing unit is used to input the image collected by the camera sensor into the knowledge distillation target recognition model of the multi-stage attention mechanism, output the target object image with the region frame in sequence, and calculate the image coordinates of the center of the target object; The second calculation unit is used to convert the image coordinates of the center of the target object into the space coordinates P in the world coordinate system. cam , and calculate the distance d from the center of the target object to the plane where the camera sensor is located cam ; The third computing unit is used to obtain the spatial coordinates P of the center of the target object through the point cloud image of the laser radar sensor las , and the distance d between the center of the target object and the plane where the lidar sensor is located las ; The fourth calculation unit is used to obtain the distance d according to the millimeter wave radar sensor. mil and speed v mil , calculate the distance deviation d1 = |d las -d cam |,d2=|d mil -d cam |, d3=|d las -d mil |; A judgment unit is used to determine if any two of the distance deviations d1, d2, and d3 are greater than or equal to a threshold value d p , then re-collect data, if any two of the distance deviations d1, d2, and d3 are less than the threshold d p , then the center coordinates of the target object are returned.
4. The multi-sensor fusion device for point cloud data acquisition according to claim 3, characterized in that: The knowledge distillation target recognition model of the multi-stage attention mechanism includes: A feature extraction unit, used to detect the image collected by the camera sensor through a feature extraction model to obtain a feature image; The visual recognition unit is used to input the feature image into the knowledge distillation target recognition model of the multi-stage attention mechanism, use the teacher network distillation to train the student network, perform fine-grained recognition based on the student network, and obtain the visual recognition result of the current target object; The trigger unit is used to determine whether the target object has reached the specified position based on the obtained limit switch detection signal. If so, it triggers the lidar sensor to obtain point cloud data to obtain the three-dimensional spatial information of the current vehicle body, and at the same time triggers the millimeter wave radar sensor to obtain the distance and speed information of the target object.
5. The multi-sensor fusion device for point cloud data acquisition according to claim 4, characterized in that: The teacher network is a fine-grained target recognition network with a multi-stage attention mechanism and ResNet50 as the backbone network. It extracts feature maps of the three stages Conv3_x, Conv4_x, and Conv5_x respectively, and adds an attention mechanism layer, a calibration layer, and a classification layer. The calibration layer maps features of different channels and sizes to specified channels and sizes through convolution operations, and fuses the output features of the three stages. The classifier is composed of a fully connected layer, which maps the feature map into an output category vector.
6. The multi-sensor fusion device for point cloud data acquisition according to claim 4, characterized in that: The student network is based on the MobinetV3 network as the backbone network, and the MobinetV3 network is pre-defined in stages. The output feature maps of the three stages Bneck6, Bneck11, and Bneck15 are extracted respectively, and an attention layer, a calibration layer, and a classification layer are additionally added. The output value of each stage of the student network is obtained through the feature mapping function and network parameters of the student network.
7. The multi-sensor fusion device for point cloud data acquisition according to claim 4, characterized in that: The feature extraction model adopts the yolov8 network structure.
8. The multi-sensor fusion device for point cloud data acquisition according to claim 1, characterized in that: During the rendering process, the visualization tools use OpenGL, Unity3D or Unreal Engine.
9. A calibration method for a multi-sensor fusion device for point cloud data acquisition, characterized in that: include: Calibrate at a known calibration site or using standard calibration objects; Install the LiDAR sensor, camera sensor, and millimeter-wave radar sensor in the predetermined positions and ensure that their relative positions and angles are constant; Simultaneously start the LiDAR sensor, camera sensor, and millimeter wave radar sensor to collect point cloud data, image data, and distance and speed information including the benchmark; The image collected by the camera sensor is input into the knowledge distillation target recognition model with a multi-stage attention mechanism, and the target object image with a region frame is output in sequence, and the image coordinates of the center of the target object are calculated; Convert the image coordinates of the center of the target object into the spatial coordinates P in the world coordinate system cam , and calculate the distance d from the center of the target object to the plane where the camera sensor is located cam ; The spatial coordinates P of the center of the target object are obtained through the point cloud image of the lidar sensor las , and the distance d between the center of the target object and the plane where the lidar sensor is located las ; According to the distance d obtained by the millimeter wave radar sensor mil and speed v mil , calculate the distance deviation d1 = |d las -d cam |,d2=|d mil -d cam |, d3=|d las -d mil |; If any two of the distance deviations d1, d2, and d3 are greater than or equal to the threshold d p , then re-collect data, if any two of the distance deviations d1, d2, and d3 are less than the threshold d p , then the center coordinates of the target object are returned.
10. The multi-sensor fusion method for point cloud data acquisition according to claim 9, characterized in that: The knowledge distillation target recognition model of the multi-stage attention mechanism includes: Detecting the image collected by the camera sensor through a feature extraction model to obtain a feature image; The feature image is input into the knowledge distillation target recognition model of the multi-stage attention mechanism, the teacher network is used for distillation to train the student network, and fine-grained recognition is performed based on the student network to obtain the visual recognition result of the current target object; According to the limit switch detection signal obtained, it is judged whether the target object has reached the specified position. If so, the lidar sensor is triggered to obtain point cloud data to obtain the three-dimensional spatial information of the current vehicle body, and the millimeter wave radar sensor is triggered to obtain the distance and speed information of the target object.
Citation Information
Cited By
Multi-sensor information visualization method and device, electronic equipment and storage medium
CN121453025A