Multi-mode sensing fusion method, device and system and unmanned fire fighting truck
Through the multimodal perception fusion method, multiple sensor data checksum joint calibration, combined with deep learning models, the problem of insufficient perception of a single sensor is solved, and the high accuracy and stability perception of the autonomous driving system in complex environments is achieved.
Patent Information
- Application Number
- CN202510458490.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
The existing autonomous driving technology mainly relies on a single type of sensor and model for perception, resulting in insufficient accuracy and stability of perceived data, especially in complex environments, and the multimodal perception fusion system is seriously affected by sensor failure.
The multimodal perception fusion method is adopted to obtain data from multiple types of sensors, perform sensor verification, space-time synchronization, filtering and joint calibration, and combine deep learning models and multi-dimensional data association rules to dynamically select available sensors for data fusion.
It improves the accuracy and stability of perceived data, reduces the impact of sensor failure on the system, and achieves flexible and reliable perception in complex environments.
Smart Images

Figure CN120372537A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous driving technology, and in particular, to a multi-modal perception fusion method, apparatus, system, and unmanned fire truck. Background Art
[0002] With the rapid development of autonomous driving technology, the application scenarios of autonomous driving technology are increasing, such as in multiple aspects including firefighting, emergency rescue, and logistics transportation. Summary of the Invention
[0003] The inventors have found that: current autonomous driving technology mainly relies on sensors to perceive the surrounding environment to achieve autonomous driving. However, current perception algorithms mostly rely on a single type of sensor (for example, pure vision perception or pure lidar perception) and a single model (deep learning object detection) for perception. Therefore, there are certain limitations in accurately, flexibly, and stably obtaining perception data of the surrounding environment to ensure the intelligence of autonomous driving. Therefore, how to accurately, flexibly, and stably obtain perception data of the surrounding environment is a problem to be solved.
[0004] In view of this, the present disclosure proposes a multi-modal perception fusion method. According to some embodiments of the first aspect of the present disclosure, there is provided a multi-modal perception fusion method, including: acquiring multi-modal data collected by at least two types of sensors; verifying the data stream status of each sensor through sensor calibration metrics, and dynamically selecting available sensors based on decoupling conditions; performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multi-modal data collected by the available sensors to obtain calibrated perception data of each available sensor; in response to using a deep learning perception model to process the calibrated perception data, being able to obtain a first-level target perception result of each target detected by each available sensor, and adopting multi-dimensional data association rules to fuse the first-level target perception results to obtain a fused target result; outputting the fused target result.
[0005] In some embodiments, the multi-modal perception fusion method further includes: in response to using a deep learning perception model to process the calibrated perception data and not obtaining a first-level target perception result, processing the calibrated perception data based on an obstacle comprehensive perception model to obtain a second-level target perception result; outputting the second-level target perception result.
[0006] In some embodiments, the sensor calibration metrics include the timestamp update status, network transmission status, and device physical status. By means of the sensor calibration metrics, the data stream status of each sensor is verified. Dynamically selecting available sensors based on decoupling conditions includes: determining sensors with at least one of the timestamp update status, network transmission status, and device physical status meeting the fault condition as unavailable sensors; if the number of unavailable sensors of the same type is greater than the threshold, determining sensors of the same type as the unavailable sensors with the number greater than the threshold as unavailable sensors; and selecting available sensors from at least two types of sensors based on the unavailable sensors.
[0007] In some embodiments, verifying the data stream status of each sensor by means of the sensor calibration metrics and dynamically selecting available sensors based on decoupling conditions further includes: in the case where the timestamp update status, network transmission status, and device physical status of the unavailable sensors do not meet the fault condition, determining the unavailable sensors as available sensors.
[0008] In some embodiments, using multi-dimensional data association rules to fuse the first-level target perception results to obtain the fused target results includes: in the case where the available sensors include multiple types of sensors, using a multi-modal combination fusion method to fuse the first-level target perception results to obtain the fused target results; in the case where the available sensors include a single type of sensor, using a single-modal fusion method to fuse the first-level target perception results to obtain the fused target results.
[0009] In some embodiments, using multi-dimensional data association rules to fuse the first-level target perception results to obtain the fused target results further includes: based on a bird's-eye view, performing the multi-modal combination fusion method and the single-modal fusion method.
[0010] In some embodiments, jointly calibrating the multi-modal data collected by the available sensors includes: determining the initial installation pose of each available sensor; performing horizontal projection correction on the multi-modal data of each available sensor to obtain the projection data of each available sensor; selecting a main sensor from the available sensors, and performing sensor registration on the slave sensors in the available sensors based on the main sensor; performing data registration on the projection data of each available sensor to obtain the joint calibration transformation matrix parameters; and calibrating the projection data of each available sensor based on the joint calibration transformation matrix parameters to obtain the calibrated perception data of each available sensor.
[0011] In some embodiments, data registration is performed on the projection data of each available sensor to obtain the joint calibration transformation matrix parameters, including: based on the normal distribution transformation algorithm, performing first registration on the projection data until the first registration data that meets the first requirement is obtained; based on the iterative closest point algorithm, performing second registration on the first registration data until the second registration data that meets the second requirement is obtained; determining the joint calibration transformation matrix parameters based on the second registration data and the projection data.
[0012] In some embodiments, jointly calibrating the multimodal data collected by the available sensors further includes: periodically correcting the joint calibration transformation matrix parameters.
[0013] In some embodiments, based on the obstacle comprehensive perception model, processing the calibrated perception data to obtain the second-level target perception result includes: based on the calibrated perception data of each available sensor, identifying the geometric state, motion state, material, and material density of the target; based on the geometric state, motion state, material, and material density of the target, obtaining the second-level target perception result.
[0014] In some embodiments, based on the calibrated perception data of each available sensor, identifying the geometric state of the target includes: performing clustering calculation on the calibrated perception data to obtain the data clusters corresponding to the calibrated perception data; based on the data clusters, using the differentiable calculation optimization method to perform surface reconstruction on the target and determine the volume of the target; analyzing the surface smoothness of the target through the normal line and curvature of the surface of the target; determining the suspended state of the target based on the contour coordinates of the target; determining the geometric state of the target based on the volume, surface smoothness, and suspended state of the target.
[0015] In some embodiments, the available sensors include lidar and millimeter-wave radar. Based on the calibrated perception data of each available sensor, identifying the motion state of the target includes: fusing the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar to obtain the motion state of the target.
[0016] In some embodiments, the available sensors include lidar and millimeter-wave radar. Based on the calibrated perception data of each available sensor, identifying the material of the target includes: analyzing the average value of the reflection intensity in the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar to identify the material of the target.
[0017] In some embodiments, the multimodal perception fusion method further includes: determining an obstacle avoidance strategy based on the fusion target result or the second-level target perception result.
[0018] In some embodiments, the first-level target perception results include at least two of the motion attributes, detection box attributes, historical trajectory attributes, and appearance semantic attributes of each target. Using multi-dimensional data association rules, the first-level target perception results are fused to obtain the fused target results, including: calculating the similarity of at least two corresponding to the first-level target perception results among the motion attribute similarity, detection box attribute similarity, historical trajectory attribute similarity, and appearance semantic attribute similarity between every two targets detected by different sensors in the available sensors; performing weighted calculation on the similarity of at least two corresponding to the first-level target perception results among the motion attribute similarity, detection box attribute similarity, historical trajectory attribute similarity, and appearance semantic attribute similarity to obtain the similarity between every two targets detected by different sensors; determining the same target detected by different sensors based on the similarity between every two targets detected by different sensors; and performing weighted fusion on the multiple first-level target perception results of the same target to obtain the fused target result of the same target.
[0019] According to some embodiments of the second aspect of the present disclosure, there is provided a multimodal perception fusion device, including: an acquisition perception module configured to acquire multimodal data collected by at least two types of sensors; a data verification module configured to verify the data stream state of each sensor through sensor calibration metrics and dynamically select available sensors based on decoupling conditions; a joint calibration module configured to perform spatio-temporal synchronization, data filtering, and joint calibration on the multimodal data collected by the available sensors to obtain the calibrated perception data of each available sensor; a multimodal fusion module configured to process the calibrated perception data in response to using a deep learning perception model, be able to obtain the first-level target perception results of each target detected by each available sensor, and fuse the first-level target perception results using multi-dimensional data association rules to obtain the fused target results; and an output fusion module configured to output the fused target results.
[0020] According to some embodiments of the third aspect of the present disclosure, there is provided a multimodal perception fusion device, including: a memory and a processor coupled to the memory, the processor being configured to execute the multimodal perception fusion method in any of the above embodiments based on instructions stored in the memory.
[0021] According to some embodiments of the fourth aspect of the present disclosure, there is provided a multimodal perception fusion system, including: the multimodal perception fusion device in any of the above embodiments; a sensor combination including at least two different types of sensors, the sensor combination being configured to collect and send multimodal data to the multimodal perception fusion device.
[0022] In some embodiments, the sensor combination includes at least two of lidar, camera, millimeter wave radar, real-time kinematic (RTK), and inertial measurement unit (IMU).
[0023] According to some embodiments of the fifth aspect of the present disclosure, a driverless fire truck is provided, including: a multi-modal perception fusion system according to any one of the above embodiments; a sensor adaptation interface configured to be connected to the sensor combination; a target recognition database for fire fighting tasks configured to provide sample data for training a deep learning perception model; and a fusion weight adjustment policy library configured to provide weights for fusing multiple first-level target perception results of the same target.
[0024] According to some embodiments of the sixth aspect of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the multi-modal perception fusion method according to any one of the above embodiments is implemented.
[0025] According to some embodiments of the seventh aspect of the present disclosure, a computer program product is provided, including computer instructions, and when the computer instructions are executed by a processor, the multi-modal perception fusion method according to any one of the above embodiments is implemented.
[0026] In the above embodiments, by acquiring multi-modal data collected by at least two types of sensors, it provides feasibility for subsequent fusion of data collected by different types of sensors and ensures the accuracy of the acquired perception data; by using sensor calibration indicators to verify the data stream status of each sensor, and dynamically selecting available sensors based on decoupling conditions, it can flexibly select available sensors that can participate in the fusion according to the real-time working status of the sensors, realizing the decoupling mechanism in the perception fusion process. Multiple sensors can be both coupled and decoupled, reflecting the intelligent automatic judgment in the perception coupling process, ensuring the flexibility and stability in the perception fusion process, and reducing the risk that inaccurate fusion target results are caused by abnormal data collected by abnormal sensors during the perception fusion process; by performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multi-modal data collected by available sensors to obtain the calibrated perception data of each available sensor, it ensures the accuracy of the multi-modal perception fusion method and reduces the risk that the accuracy and stability of the multi-modal perception fusion method are affected by problems such as spatial inconsistency, time asynchrony, and multi-modal data conflicts; by processing the calibrated perception data through a deep learning perception model to obtain the first-level target perception results of each target detected by each available sensor, it realizes the differentiation of different targets and provides feasibility for subsequent fusion of the first-level target perception results for each target; by adopting multi-dimensional data association rules to fuse the first-level target perception results and obtain and output the fused target results, it realizes the fusion of multi-modal data for each target and ensures the accuracy of the fused target results of each target obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0028] With reference to the drawings, the present disclosure can be more clearly understood according to the following detailed description.
[0029] Figure 1 Schematic diagrams showing some embodiments of the multi-modal perception fusion method of the present disclosure.
[0030] Figure 2 Schematic diagrams showing some embodiments of the method for dynamically selecting available sensors of the present disclosure.
[0031] Figure 3 Schematic diagrams showing some embodiments of the method for jointly calibrating multi-modal data of the present disclosure.
[0032] Figure 4 Schematic diagrams showing some embodiments of the multi-dimensional data association rules of the present disclosure.
[0033] Figure 5 Schematic diagrams showing some embodiments of the method for the comprehensive obstacle perception model of the present disclosure.
[0034] Figure 6 Schematic diagrams showing some other embodiments of the multi-modal perception fusion method of the present disclosure.
[0035] Figure 7 Schematic diagrams showing some embodiments of the multi-modal perception fusion device of the present disclosure.
[0036] Figure 8 Schematic diagrams showing some other embodiments of the multi-modal perception fusion device of the present disclosure.
[0037] Figure 9 Schematic diagrams showing some embodiments of the multi-modal perception fusion system of the present disclosure.
[0038] Figure 10 Schematic diagrams showing some embodiments of the unmanned fire truck of the present disclosure. Detailed implementation manners
[0039] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.
[0040] Meanwhile, it should be understood that, for the sake of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship.
[0041] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present disclosure or its application or use.
[0042] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and devices should be regarded as part of the specification.
[0043] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.
[0044] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0045] With the rapid development of autonomous driving technology, the fusion of multi-modal perception data has become a key direction for improving the accuracy, stability of the perception data of the surrounding environment and the robustness of the autonomous driving system.
[0046] At present, for fire trucks in the construction machinery industry, the fusion method mainly focuses on the solution of using pure lidar for environmental perception, without the participation of cameras and millimeter-wave radars, which limits the stability to a certain extent during the fusion process. Regarding the limitations of single sensors, they are as follows: Cameras have degraded performance in low light or adverse weather conditions, and cannot provide accurate depth information, with significant deficiencies in long-distance perception and speed estimation. Lidar has insufficient accuracy in long-distance or high-reflectivity scenarios, and is affected by strong light, water mist, smoke, dust, snowflakes, thick fog, etc. Lidar has relatively weak capabilities in target classification, and the accuracy of lidar target classification is lower than that of cameras. Millimeter-wave radars also have the defects of weak classification ability, sparse point clouds, and unstable reflection intensity. In addition, millimeter-wave radars are greatly affected by the environment. For example, the wavelength of millimeter-wave radars is short and the penetration ability is weak. Therefore, in adverse weather such as rain and snow, their accuracy may be affected, reducing their detection ability.
[0047] Autonomous driving solutions are also evolving from single homogeneous sensors to heterogeneous multi-sensor fusion. Image algorithms have gone through 2D object detection, 3D object detection, BEV (Bird's Eye View), and then to the Semantic Occupancy Network, indicating the increasing demand for understanding the three-dimensional space. Among them, the Semantic Occupancy Network is a technology for 3D scene understanding and reconstruction. Millimeter-wave radar has also gone through signal processing (e.g., using Fast Fourier Transform (FFT) to estimate the distance and speed of targets, and using Beamforming or MUSIC (MUltiple SIgnal Classification) algorithms to estimate the angle of targets), then combined with the Kalman Filter Algorithm to achieve the multi-target tracking stage, and then to the CRF-Net (Conditional Random Field Network) proposed by combining deep learning with conditional random fields. In the future, it will also move towards the stage of fusing multi-modal data. LiDAR has evolved from traditional clustering analysis and geometric feature extraction analysis to deep learning MV3D (Multi-View 3D Object Detection Network), VoxelNet (Voxel-based 3D Detection Network), PointPillars (Point Cloud Processing via Pillar Encoding), FormerLocal3D (Transformer-based Local Attention for 3D Data Processing), etc., making full use of the three-dimensional geometric information of point clouds to improve detection accuracy. At the same time, deep learning also increases the computational complexity and has high requirements for hardware. Therefore, deep learning still faces some challenges in practical applications, such as the increasing demand for hardware vehicle cores, graphics cards, and computing power.
[0048] However, there may be conflicts in the target perception results of different sensors, and a large amount of labeled data is also required for the training of deep learning models. In addition, for the fusion of multi-modal perception data, whether it is data-level fusion or feature-level fusion, there is a common problem, that is, when a certain sensor is damaged or fails, the multi-modal perception fusion system (i.e., the system that executes the multi-modal perception fusion method) will be greatly affected, facing a serious decline in performance or directly stopping running. In severe cases, it will cause secondary accidents, which has a great impact on the entire autonomous driving system. These are all thorny problems that need to be solved urgently.
[0049] Furthermore, in the current post-fusion scheme of cameras and lidars, the lidar is projected into the image, or the pixels of the image are first projected into the camera coordinate system and then converted to the three-dimensional physical world coordinate system. These two fusion methods belong to simple post-fusion methods and have relatively strict requirements for joint calibration, and there are certain limitations in reducing fusion errors. At the same time, there are limitations in solving the problems of occlusion and large long-distance perception errors. The current perception scheme of the camera and lidar combination uses a pre-fusion algorithm to perceive the surrounding environment. Compared with the post-fusion algorithm, although the pre-fusion algorithm can make more full use of the data characteristics of the collected perception data, the pre-fusion algorithm also has relatively strict requirements for the hardware (i.e., the computing performance of the hardware) (for example, the pre-fusion algorithm usually requires dual orin chips to support the computing power of the hardware). At the same time, the pre-fusion algorithm has certain limitations in the investigation and positioning of targets in some special scenarios. The aforementioned post-fusion algorithm and pre-fusion algorithm both closely rely on the continuous input of the perception data collected by various sensors. Especially for the pre-fusion algorithm, the failure or damage of some sensors will have a greater impact on the fusion result and the safety and stability of the autonomous driving process. In addition, the pre-fusion algorithm has relatively strict requirements for the quality of the data collected by the sensors. Also, due to the limited number of target categories recognized by the pre-fusion algorithm, there is a potential safety hazard that the perception method fails for unfamiliar targets (obstacles).
[0050] Regarding how to accurately, flexibly, and stably obtain the perception data of the surrounding environment and reduce the risk of harm caused by inaccurate or unstable perception data of the surrounding environment, the following is as follows.
[0051] Figure 1 Schematic diagrams showing some embodiments of the multi-modal perception fusion method of the present disclosure.
[0052] As Figure 1 shown, the multi-modal perception fusion method includes step 110 to step 150, and this multi-modal perception fusion method is executed by a multi-modal perception fusion device.
[0053] In step 110, multi-modal data collected by at least two types of sensors is obtained.
[0054] For example, the sensors can be at least two of lidar, camera, millimeter-wave radar, Real-Time Kinematic (RTK), or Inertial Measurement Unit (IMU). Among them, lidar is used to obtain three-dimensional point cloud data of the environment, the camera is used to obtain two-dimensional image data of the environment, millimeter-wave radar is used to obtain the motion information of targets in the environment, RTK is used to obtain the high-precision positioning of the vehicle equipped with at least two types of sensors, and IMU is used to obtain the attitude information of the vehicle equipped with at least two types of sensors.
[0055] In step 120, through sensor verification metrics, the data stream status of each sensor is verified, and available sensors are dynamically selected based on decoupling conditions.
[0056] For example, the sensor verification metrics include timestamp update status (e.g., whether there is frame freezing, etc.), network transmission status (e.g., data packet loss rate (i.e., whether there is severe data packet loss), collision with no signal, transmission latency, etc.), and device physical status (e.g., whether the temperature is normal, whether the voltage is normal, whether there is power outage, etc.). Among them, the sensor verification metrics refer to the common metrics for at least two types of sensors (e.g., lidar, camera, millimeter-wave radar).
[0057] By using sensor verification metrics to verify the data stream status of each sensor, it is determined whether each sensor is working properly.
[0058] In step 130, the multi-modal data collected by available sensors is subjected to spatio-temporal synchronization, filtering, data cleaning, and joint calibration to obtain the calibrated perception data of each available sensor.
[0059] By performing spatio-temporal synchronization, filtering, and data cleaning on the multi-modal data collected by available sensors, good data alignment conditions are provided for the subsequent fusion of multi-modal data. In addition, by performing spatio-temporal synchronization, filtering, and data cleaning on the multi-modal data collected by available sensors, the noise and errors in the multi-modal data can be reduced, providing the possibility of improving the accuracy of perception fusion.
[0060] In step 140, in response to processing the calibrated perception data using a deep learning perception model, the first-level target perception results of each target detected by each available sensor can be obtained, and multi-dimensional data association rules are used to fuse the first-level target perception results to obtain the fused target results.
[0061] By fusing the first-level target perception results of the same target detected by each available sensor, a fused target result is obtained, fully exploiting the complementary information between the multi-modal data collected by different sensors and improving the accuracy and robustness of perception fusion.
[0062] In step 150, the fused target result is output.
[0063] In some embodiments, the fused target result with a unified format is output.
[0064] In the above embodiments, by obtaining multi-modal data collected by at least two types of sensors, it provides feasibility for subsequent fusion of data collected by different types of sensors and ensures the accuracy of the obtained perception data; by using sensor calibration metrics to verify the data stream status of each sensor and dynamically select available sensors based on decoupling conditions, it can flexibly select available sensors that can participate in the fusion according to the real-time working status of the sensors, realizing a decoupling mechanism in the perception fusion process. Multiple sensors can both couple and decouple, reflecting the intelligent automatic judgment in the perception coupling process, ensuring the flexibility and stability in the perception fusion process, and reducing the risk of inaccurate fused target results caused by abnormal data collected by abnormal sensors during the perception fusion process; by performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multi-modal data collected by available sensors to obtain the calibrated perception data of each available sensor, it ensures the accuracy of the multi-modal perception fusion method and reduces the risk of affecting the accuracy and stability of the multi-modal perception fusion method due to problems such as spatial inconsistency, time asynchrony, and multi-modal data conflicts; by processing the calibrated perception data through a deep learning perception model to obtain the first-level target perception results of each target detected by each available sensor, it realizes the differentiation of different targets and provides feasibility for subsequent fusion of the first-level target perception results for each target; by adopting multi-dimensional data association rules to fuse the first-level target perception results and obtain and output the fused target result, it realizes the fusion of multi-modal data for each target and ensures the accuracy of the fused target result of each obtained target.
[0065] In some embodiments, regarding verifying the data stream status of each sensor through sensor calibration metrics and dynamically selecting available sensors based on decoupling conditions, it is as follows: A sensor whose at least one of the timestamp update status, network transmission status, and device physical status satisfies the failure condition is determined as a non-available sensor; if the number of non-available sensors of the same type is greater than the threshold, the sensors of the same type as the non-available sensors with a number greater than the threshold are determined as non-available sensors; based on the non-available sensors, available sensors are selected from at least two types of sensors.
[0066] That is to say, if the number of non - available sensors of the same type is greater than the threshold α, the multi - modal perception fusion system will automatically decouple all sensors of this type and mark all sensors of this type (i.e., the available sensors and non - available sensors of this type) as non - available sensors. In addition, the signal used to indicate non - available sensors will be returned to the maintenance center in real time to remind relevant staff to repair the sensors. After the sensor status is updated, the multi - modal perception fusion system initializes the test. If the sensor calibration indicators of the sensors are all normal, the normal fusion mode will be restored, and the restored non - available sensors will be marked as available sensors.
[0067] The decoupling conditions include single - sensor decoupling conditions and multi - sensor decoupling conditions. Among them, the single decoupling condition is to determine a non - available sensor when at least one of the timestamp update status, network transmission status, and device physical status of a sensor meets the failure condition. That is, the data stream of the data collected by this sensor will be automatically filtered before fusion. The multi - sensor decoupling condition is that if the number of non - available sensors of the same type is greater than the threshold, the sensors of the same type as the non - available sensors with a number greater than the threshold will be determined as non - available sensors. That is, the data stream of the data collected by the sensors of the same type as the non - available sensors with a number greater than the value will be automatically filtered before fusion. Among them, the threshold is proportional to the total number of all sensors of the same type as a certain non - available sensor.
[0068] For example, a certain non - available sensor is a millimeter - wave radar, and the total number of millimeter - wave radars is 10. The threshold can be set according to the ratio of the number of non - available millimeter - wave radars to the total number of millimeter - wave radars. For example, if the ratio is 60%, the threshold is 6. When the number of non - available millimeter - wave radars is greater than 6, all millimeter - wave radars are determined as non - available sensors, and the data streams of the data collected by all millimeter - wave radars will be automatically filtered before fusion. Of course, taking the non - available millimeter - wave radar as an example, the threshold can also be set to 60%, directly judging whether the ratio of the number of non - available millimeter - wave radars to the total number of millimeter - wave radars is greater than 60%. If it is greater, all millimeter - wave radars will be determined as non - available radars.
[0069] In addition, in the case where the timestamp update status, network transmission status, and device physical status of the non - available sensor do not meet the failure conditions, the non - available sensor is determined as an available sensor.
[0070] For example, a non - available sensor is determined to be an available sensor if it meets the recovery conditions. The recovery conditions include single - sensor recovery conditions and multi - sensor recovery conditions. Among them, the single - sensor recovery condition is that the timestamp update status, network transmission status, and device physical status of a certain non - available sensor do not meet the fault conditions, and the non - available sensor is determined to be an available sensor. For example, when the weather temperature drops, the voltage of the non - available sensor becomes stable and normal, and the network signal becomes stronger, that is, the timestamp update status, network transmission status, and device physical status of the non - available sensor all return to normal, and this non - available sensor can be determined to be an available sensor. The multi - sensor recovery condition is that the timestamp update status, network transmission status, and device physical status of all non - available sensors of a certain type do not meet the fault conditions, then all non - available sensors of a certain type are determined to be available sensors. For example, all non - available sensors of a certain type are damaged sensors, and if all non - available sensors of this type are updated, then all non - available sensors of this type can be determined to be available sensors.
[0071] Another example is that during the process of determining all non - available sensors of a certain type as available sensors, an additional restriction can be added, that is, the proportion of the number of recovered non - available sensors of this type at least meets a specific threshold, such as 30%.
[0072] By decoupling the conditions and recovery conditions and dynamically selecting available sensors, it is possible to handle emergencies more flexibly. For example, when a sensor with too high temperature fails or a vehicle encounters a collision causing some sensors to be damaged, the available sensors can be adjusted in time, enabling the unmanned vehicle applying this multi - modal perception fusion method to make a reasonable sensor combination plan in various scenarios and flexibly activate the sensor combination plan suitable for the current scenario.
[0073] In some embodiments, when the available sensors include multiple types of sensors, a multi - modal combination fusion method is adopted to fuse the first - level target perception results to obtain a fused target result; when the available sensors include a single type of sensor, a single - modal fusion method is adopted to fuse the first - level target perception results to obtain a fused target result.
[0074] For example, when the available sensors include multiple types of sensors, a multi - modal BEV (Bird's Eye View) post - fusion method is adopted to fuse the data projected from the first - level target perception results onto the BEV coordinate system, and finally output the fused target result.
[0075] For example, based on BEV (Bird's Eye View), the multi - modal combination fusion method and the single - modal fusion method are executed.
[0076] For example, the multimodal combination fusion methods include the lidar-camera combination fusion method, the lidar-millimeter wave radar combination fusion method, the camera-millimeter wave radar combination fusion method, and the lidar-camera-millimeter wave radar combination fusion method, etc. The unimodal fusion methods include the pure vision (camera) fusion method, the pure lidar fusion method, and the pure millimeter wave radar fusion method, etc.
[0077] Figure 2 Schematic diagrams showing some embodiments of the method for dynamically selecting available sensors of the present disclosure.
[0078] As Figure 2 shown, the method for dynamically selecting available sensors is executed by the multi-sensor verification layer 201, and determining the fusion method according to the available sensors is executed by the executor 205.
[0079] The multi-sensor verification layer 201 verifies the data stream status of each sensor through the sensor verification metrics 202, and dynamically selects available sensors based on the decoupling conditions 203. As Figure 2 shown, the sensor verification metrics 202 include the timestamp update status 206, the network transmission status 207, and the device physical status 208. The decoupling conditions 203 include single-sensor decoupling 209 and multi-sensor decoupling 210. Corresponding to the decoupling conditions, there is also a recovery coupling condition 204, and the recovery coupling condition 204 includes single-sensor recovery coupling 211 and multi-sensor recovery coupling 212.
[0080] If the decoupling condition 203 is a situation where at least one of the sensor has an abnormal timestamp update status, an abnormal network transmission status, and an abnormal device physical status, then the sensor needs to be decoupled; if the recovery fusion condition 204 is a situation where the sensor has no abnormal timestamp update status, abnormal network transmission status, and abnormal device physical status, then the sensor can be restored from the state of an unavailable sensor to the state of an available sensor.
[0081] The fusion scheme determined by the executor 205 can be a pure vision scheme 213, a pure lidar scheme 214, a pure millimeter wave radar scheme 215, a lidar-camera fusion scheme 216, a lidar-millimeter wave radar fusion scheme 217, a lidar-millimeter wave radar-camera fusion scheme 218, or a camera-millimeter wave radar fusion scheme 219. For example, if the available sensor is only a camera, the fusion scheme determined by the executor 205 is the pure vision scheme 213; if the available sensors are a lidar and a camera, the fusion scheme determined by the executor 205 can be the lidar-camera fusion scheme.
[0082] In some embodiments, for the joint calibration of multimodal data collected by available sensors, the following is specifically carried out: determining the initial installation pose of each available sensor; performing horizontal projection correction on the multimodal data of each available sensor to obtain the projection data of each available sensor; selecting a master sensor from the available sensors and performing sensor registration on slave sensors in the available sensors based on the master sensor; performing data registration on the projection data of each available sensor to obtain the joint calibration transformation matrix parameters; and performing calibration on the projection data of each available sensor based on the joint calibration transformation matrix parameters to obtain the calibrated perception data of each available sensor. Among them, the joint calibration transformation matrix parameters are used to indicate the coordinate transformation relationship between the projection data and the calibrated perception data of each sensor.
[0083] By performing spatial alignment and projection at the raw data level, the data quality of joint calibration is improved.
[0084] For example, the following method is used to perform data registration on the projection data of each available sensor to obtain the joint calibration transformation matrix parameters: based on the normal distribution transformation algorithm, performing first registration on the projection data until first registration data that meets the first requirement is obtained; based on the iterative closest point algorithm, performing second registration on the first registration data until second registration data that meets the second requirement is obtained; and determining the joint calibration transformation matrix parameters based on the second registration data and the projection data.
[0085] In some embodiments, taking lidar and millimeter-wave radar as examples, for the joint calibration of each available sensor, the obtained joint calibration transformation matrix parameters include: based on the point cloud normal distribution transformation algorithm, performing registration on the projection data until the cumulative value of the joint probability density function is greater than a certain specific threshold, and then the first registration (i.e., the first-stage rough registration) ends; entering the second registration (i.e., the second-stage fine registration), using the iterative closest point algorithm (ICP) until the distance error meets another specific threshold, and then outputting the final calibration parameters, that is, obtaining the joint calibration transformation matrix parameters containing the rotation R and the translation parameter T. Among them, the point cloud normal distribution transformation algorithm is an improved adaptive normal distribution transformation algorithm.
[0086] Considering that different sensors (such as lidar, camera, millimeter-wave radar, IMU, etc.) have different data characteristics and working modes, resulting in great difficulties in joint calibration. Through the joint calibration method in the above embodiments, the manual intervention is reduced and it does not rely on a specific calibration board, providing an accuracy guarantee for subsequent perception fusion. High-precision joint calibration can improve the accuracy of perception fusion, especially in complex scenarios (such as at night, rainy or snowy weather, etc.). In addition, by performing parallel data registration on the projection data of each available sensor, the real-time performance and efficiency of joint calibration can be ensured.
[0087] In some embodiments, after the first registration is completed, it is possible to determine the parameters of the joint calibration transformation matrix with a relatively poor joint calibration effect. Subsequently, the second registration is performed to optimize the parameters of the joint calibration transformation matrix obtained previously to obtain the parameters of the joint calibration transformation matrix with a relatively better joint calibration effect.
[0088] For example, taking the joint calibration between a lidar and a millimeter-wave radar as an example, the first requirement may be that the cumulative joint probability density function value of the registration data (i.e., the first registration data) obtained from the first rough registration of the lidar or the millimeter-wave radar is greater than the initial threshold, and the second requirement may be that the distance error of the data (i.e., the second registration data) obtained from the second refined registration of the lidar or the millimeter-wave radar is less than the secondary threshold.
[0089] In some embodiments, regarding the conditions for stopping the registration in the first rough registration and the second refined registration, in addition to the first requirement and the second requirement, it is also possible to determine whether to stop the registration based on the number of registration times in the first registration and the number of registration times in the second registration.
[0090] On the basis of meeting the first requirement and the second requirement, the conditions for stopping the registration also take into account the limitation of the number of registration times. By adding the limitation of the number of registration times on the basis of meeting the first requirement and the second requirement, the risk of the registration process falling into an infinite loop is reduced, and the efficiency of joint calibration is guaranteed.
[0091] In some embodiments, taking the joint calibration between lidars as an example, the rotation angle of the lidar in the first registration process is greater than the rotation angle of the lidar in the second registration process.
[0092] In some embodiments, during the process of performing the second registration on the first registration data, in addition to performing the second registration through the iterative closest point algorithm, the second registration can also be achieved through manual adjustment.
[0093] For example, taking the joint calibration between lidars as an example, the parameters of the joint calibration transformation matrix include the rotation matrix R of the point cloud and the translation vector T of the point cloud.
[0094] In some embodiments, regarding the joint calibration of multi-modal data collected by available sensors, the parameters of the joint calibration transformation matrix can also be corrected periodically.
[0095] For example, periodically correcting the parameters of the joint calibration transformation matrix can be to check and correct the parameters of the joint calibration transformation matrix. If the error of the current parameters of the joint calibration transformation matrix is greater than a certain threshold, it will enter the automatic joint calibration calculation module to automatically replace the calibration file (update the parameters of the joint calibration transformation matrix), and then start the perception fusion algorithm.
[0096] Due to the vibration phenomenon that accompanies the long-term driving of vehicles, the calibrated sensors may experience slight changes in pose every quarter or half year, which will affect the accuracy of multi-modal data perception and fusion. Therefore, an automatic joint calibration and correction algorithm is proposed, that is, by periodically (for example, once a day) checking and correcting the parameters of the joint calibration transformation matrix, the effectiveness of the joint calibration transformation matrix parameters is guaranteed, thus ensuring the accuracy of the multi-modal perception and fusion method, and providing feasibility for the application of the multi-modal perception and fusion method in complex scenarios.
[0097] Figure 3 Schematic diagrams showing some embodiments of the method for jointly calibrating multi-modal data of the present disclosure.
[0098] As Figure 3 shown, the method for jointly calibrating multi-modal data includes steps 301 to 315, specifically as follows, taking lidar and millimeter-wave radar as available sensors as an example.
[0099] In step 301, measure the initial pose of each available sensor (lidar and millimeter-wave radar).
[0100] In step 302, perform horizontal projection on the multi-modal data collected by each available sensor to obtain the projection data of each available sensor.
[0101] In step 303, determine the rotation angle of each available sensor.
[0102] In step 304, preprocess the projection data of each available sensor, that is, filter and denoise.
[0103] In step 305, divide the space of each available sensor into grids and count the point clouds in the grids.
[0104] In step 306, calculate the mean and covariance of the grids and construct a Gaussian distribution.
[0105] In step 307, determine the main sensor (a certain lidar) and the slave sensor (millimeter-wave radar or lidar) among the available sensors, and convert the point cloud data (point cloud to be registered) of the slave sensor into the point cloud coordinate system of the main sensor (i.e., the reference point cloud coordinate system).
[0106] In step 308, perform the first registration on the projection data of each available sensor through the NDT algorithm to obtain the first registration data, where the first registration belongs to the first rough registration.
[0107] In step 309, calculate the value of the joint probability density function (also called the score) between the first registration data, and solve the rotation matrix R and the translation vector T of the point cloud.
[0108] In step 310, it is determined whether the score of the joint probability density between the first registration data is greater than the initial threshold. If it is greater, step 311 is executed. If it is less than or equal, step 308 is executed again, that is, the first registration is performed again.
[0109] In step 311, the better R and T are output.
[0110] In step 312, the first registration data is secondarily registered by a local ICP algorithm or manual fine-tuning to obtain second registration data, and the distance error between the second registration data is calculated. Among them, the second registration belongs to the second refined registration.
[0111] In step 313, it is determined whether the distance error between the second registration data is less than the secondary threshold. If it is less, step 314 is executed. If it is greater than or equal, step 308 is executed again, that is, the first registration is performed again.
[0112] In step 314, the optimal R and T are output.
[0113] In step 315, the joint calibration method is ended.
[0114] In some embodiments, the first-level target perception results include at least two of the motion attributes (e.g., position coordinates, speed, heading angle), detection box attributes, historical trajectory attributes, and appearance semantic attributes of each target. Regarding the use of multi-dimensional data association rules, the first-level target perception results are fused to obtain fused target results, specifically as follows: calculating the similarity of at least two corresponding to the first-level target perception results among the motion attribute similarity, detection box attribute similarity, historical trajectory attribute similarity, and appearance semantic attribute similarity between every two targets detected by different sensors in the available sensors; performing weighted calculation on the similarity of at least two corresponding to the first-level target perception results among the motion attribute similarity, detection box attribute similarity, historical trajectory attribute similarity, and appearance semantic attribute similarity to obtain the similarity between every two targets detected by different sensors; determining the same target detected by different sensors based on the similarity between every two targets detected by different sensors; performing weighted fusion on multiple first-level target perception results of the same target to obtain the fused target result of the same target. Among them, the detection box attribute similarity between every two targets detected by different sensors in the available sensors refers to the overlap degree of the detection boxes between every two targets detected by different sensors.
[0115] By performing weighted fusion on multiple first-level target perception results of the same target detected by different sensors, a fused target result of the same target is obtained, which fully combines the perception data of different sensors for the same target and improves the accuracy and stability of perception fusion.
[0116] For example, the fused target result of the same target may include target position, target speed, target category, target identifier, target confidence, target heading deviation angle, etc.
[0117] The multi-dimensional data association method includes: calculating the similarity of at least two corresponding to the first-level target perception results among the similarity of motion attributes, detection box attributes, historical trajectory attributes, and appearance semantic attributes between every two targets detected by available sensors; performing weighted calculation on the similarity of at least two corresponding to the first-level target perception results among the similarity of motion attributes, detection box attributes, historical trajectory attributes, and appearance semantic attributes to obtain the similarity between every two targets detected by different sensors.
[0118] For example, the motion attributes of each target include the position coordinates of each target, the speed of each target, and the heading angle of each target. The motion attributes of each target are mainly used to distinguish static targets and dynamic targets.
[0119] It is considered that if two targets are vehicles with similar heights traveling at the same speed and in the same direction at an equal distance, it is relatively difficult to distinguish between the two targets if only the motion attributes of the targets are considered during the process of distinguishing the two targets.
[0120] To supplement the deficiency of only considering the motion attributes of the target, a multi-dimensional data association method is provided. During the process of calculating the similarity between every two targets detected by different sensors, at least two of the motion attributes, detection box attributes, historical trajectory attributes, and appearance semantic attributes of each target are considered, which fully combines the spatial features and time features of the target, realizes the precise association between every two targets detected by different sensors, and provides a good foundation for multi-target tracking and multi-target motion trajectory prediction. In addition, by performing weighted fusion on multiple first-level target perception results of the same target to obtain the fused target result of the same target, the accuracy of the obtained fused target result is guaranteed, and the consistency between the estimated target motion state and the actual target motion state during the target tracking process is also guaranteed.
[0121] The attribute information of multiple dimensions of the target is considered to perform the matching between every two targets in different sensors (calculate the similarity between every two targets in different sensors), making full use of the data characteristics of the data collected by each sensor, ensuring the accuracy and stability of the target association process, adapting to the target matching in complex scenarios (for example, the situation where two targets are close in position, the same in appearance color, and similar in volume), solving problems such as short-term occlusion of the target, frequent switching of target identifiers (considering the motion attribute, appearance semantic attribute, detection box attribute, and historical trajectory attribute in the process of calculating the similarity between every two targets in different sensors helps target tracking), and unclear target distinguishability (considering the motion attribute, appearance semantic attribute, detection box attribute, and historical trajectory attribute in the process of calculating the similarity between every two targets in different sensors helps to distinguish targets from various dimensions), and improving the stability and robustness of the multi-modal perception fusion method.
[0122] In addition, the multi-dimensional data association rule can make the multi-target tracking performance more stable, output stable and unique target data for planning and decision-making, and is of great significance for promoting the development and simplifying the maintenance of high-order autonomous driving systems.
[0123] Regarding the detection box attribute between every two targets detected by different sensors, both the camera and the lidar can perform 3D detection on the target. The camera can use the LSS (Lift-Splat-Shoot) method to obtain the visual BEV perception result. However, for the millimeter-wave radar, whether it is 3D detection or 4D detection, the perception error of the volume and shape contour of the target is relatively large. Therefore, in terms of the perception of the detection box attribute of the target, it mainly relies on the fusion of lidar and camera data. The detection target results are projected onto the BEV space, and the overlap degree of the detection boxes between every two targets detected by different sensors in the available sensors is obtained through the IOU (Intersection over Union) calculation formula, which is used as one of the important references for the data association of different sensor perception results and also as an important reference for the multi-dimensional data association rule.
[0124] Regarding the historical trajectory attribute between every two targets detected by different sensors, that is, considering the time characteristics of each target, making full use of the historical motion trajectory information to increase the stability of target association. The motion trajectory of each target is regarded as a set of a series of points, and the similarity between every two targets detected by different sensors is determined by comparing the shapes of the motion trajectories of the targets.
[0125] In the process of calculating the similarity between every two targets detected by different sensors, considering the similarity of the historical trajectory attribute can more comprehensively perceive and understand the motion state of the target, and improve the accuracy of target matching.
[0126] For example, the DTW (Dynamic Time Warping) algorithm can be used to calculate the similarity of the motion trajectories of every two targets detected by different sensors to determine the similarity between every two targets detected by different sensors.
[0127] The DTW algorithm is an algorithm for comparing the similarity of two time series. The DTW algorithm can automatically align the points in the motion trajectories of every two targets detected by different sensors and calculate the minimum cumulative distance between the motion trajectories of every two targets detected by different sensors, taking into account the order and time interval of the points, so as to more accurately measure the similarity of the motion trajectories of every two targets detected by different sensors. The main advantage of the DTW algorithm is to process sequences with inconsistent lengths, and it can handle time delays or accelerations in the motion trajectories. This algorithm focuses on the shape similarity of the motion trajectories, not just the position of the points.
[0128] In some embodiments, the motion trajectories of target A and target B are represented as a set of time series points: and where is the position of target A at the i-th moment, is the position of target B at the j-th moment. The DTW algorithm is an algorithm for calculating the similarity between two time series, which can handle the problems of inconsistent time series lengths and time offsets, and calculate the distance matrix D of two motion trajectories T A and T B . D ij represents and the distance between, ||·||2 is the Euclidean distance.
[0129] The cumulative distance matrix C is calculated through dynamic programming, where C ij represents the minimum cumulative distance from (1, 1) to (i, j): C i,j = D i,j + min(C i-1,j , C i,j-1 , C i-1,j-1 ), and the initial condition is: C 1,1 = D 1,1 . The distance calculated by the DTW algorithm is the last element C n,m of the cumulative distance matrix C, that is, DTW(T A , T B ) = C n,m , and the similarity of the historical trajectory attributes between every two targets detected by different sensors is: Among them, the similarity of the historical trajectory attributes ranges from 0 to 1.
[0130] For the appearance semantic attributes between every two targets detected by different sensors, the appearance semantic attributes include the color attributes and texture attributes between every two targets detected by different sensors. The similarity between every two targets detected by different sensors is determined by the color attributes and texture attributes of the images of the targets. For example, the color histogram is a statistical method for describing the color distribution of a target, which can effectively capture the color attributes, semantic categories, etc. of the target.
[0131] In some embodiments, the color attribute between every two targets detected by different sensors may be, for example, a color histogram. Regarding how to obtain the color histogram, specifically as follows, the image collected by the camera is divided into several intervals, and the number of pixels in each interval is counted. The pixel of the image is I(x, y), and the color histogram H of this pixel is defined as: H(b) = ∑ x,y δ(b, Bin(I(x, y))), where b refers to the interval index of the color histogram, Bin(I(x, y)) is a function that maps the pixel I(x, y) to the corresponding interval, and δ is the Kronecker delta function.
[0132] In some embodiments, regarding how to obtain the texture attributes between every two targets detected by different sensors, texture attribute extraction can be performed through a deep learning model (such as a CNN (Convolutional Neural Network)).
[0133] In some embodiments, the similarity between every two targets detected by different sensors is determined by calculating the similarity of the color attributes and the similarity of the texture attributes between every two targets detected by different sensors. Among them, the color histograms of every two targets detected by different sensors are denoted as H1 and H2, then the similarity of the color attributes between every two targets detected by different sensors is: The texture feature vectors (texture attributes) of every two targets detected by different sensors are denoted as t1 and t2, then the similarity of the texture attributes between every two targets detected by different sensors is: That is, the similarity between every two targets detected by different sensors is: sim img (A, B) total = w colo( × sim(H A , H B ) + w texture × sim(t A , t B ), where A and B represent every two targets detected by different sensors, wcolo( , w t+xtu(+ are the weights of the color attribute similarity and the texture attribute similarity respectively, where w colo( + w t+xtu(+ = 1.
[0134] In some embodiments, taking the weighted calculation of the motion attribute similarity, the appearance semantic attribute similarity, the detection box attribute similarity, and the historical trajectory attribute similarity to obtain the similarity between every two targets detected by different sensors as an example, where the weights of the motion attribute similarity, the appearance semantic attribute similarity, the detection box attribute similarity, and the historical trajectory attribute similarity are 0.5, 0.15, 0.15, and 0.2 in sequence, and the similarity between every two targets detected by different sensors (i.e., the total target association score of every two targets detected by different sensors) is calculated. The higher the similarity, the higher the probability that the two targets are the same target.
[0135] For example, the sum of the weights of the motion attribute similarity, the appearance semantic attribute similarity, the detection box attribute similarity, and the historical trajectory attribute similarity is 1, that is, w1 + w2 + w3 + w4 = 1, where the similarity between every two targets detected by different sensors is: sim total (A, B) = w1 × sim motion (A, B) + w2 × sim img (A, B) + w3 × sim volume (A, B) + w4 × sim trace (A, B), where A and B represent every two targets detected by different sensors.
[0136] Figure 4 A schematic diagram showing some embodiments of the multi-dimensional data association rule of the present disclosure.
[0137] As Figure 4 shown, the multi-dimensional data association rule 401 mainly performs the association (matching) between every two targets detected by different sensors from the following four dimensions, namely the motion attribute 402, the appearance semantic attribute 403, the detection box attribute 404, and the historical trajectory attribute 405. After determining the similarities of the above four dimensions between every two targets detected by different sensors, multi-dimensional data association 417 is performed, that is, the similarities of the above four dimensions between every two targets detected by different sensors are weighted and summed through a linear weighting function to obtain the similarity between every two targets detected by different Changanqi (i.e., the multi-dimensional data association score 419), and finally the matching result 420 is output, that is, the same target detected by different sensors.
[0138] As Figure 4As shown, the motion attributes 402 of the target include the horizontal and vertical positions 406 of the target, the horizontal and vertical speeds 407 of the target, and the heading angle and direction 408 of the target. The appearance semantic attributes 403 of the target include the color attributes of the target (i.e., the color histogram 409) and the texture attributes (i.e., the semantic categories of pixel points 410). Through the color attributes and texture attributes of the target, the appearance semantic attribute similarity of every two targets detected by different sensors is determined. The detection box attributes 404 of the target refer to detecting the target through the 3D camera 412 and the 3D lidar 413, obtaining the detection box of the target through the deep learning perception model, and then obtaining the overlap degree 414 of the detection boxes of every two targets detected by different sensors through the IOU formula. The detection box attribute similarity of every two targets detected by different sensors is determined through the overlap degree. The historical trajectory attributes 405 of the target can calculate the historical trajectory similarity 416 of every two targets detected by different sensors through the dynamic time warping algorithm 415.
[0139] For example, for autonomous driving technology, based on the fused target results, an obstacle avoidance strategy can be determined for mechanical devices (such as unmanned fire trucks, unmanned trucks, etc.), that is, accurate decision-making information can be provided for the obstacle avoidance strategy.
[0140] In some embodiments, based on the fused target results, a target path can also be determined for the mechanical device to achieve the autonomous driving of the mechanical device, which helps the mechanical device to develop in the direction of intelligence.
[0141] In some embodiments, in response to processing the calibrated perception data using the deep learning perception model and not obtaining the first-level target perception result, the calibrated perception data is processed based on the obstacle comprehensive perception model to obtain the second-level target perception result; the second-level target perception result is output.
[0142] That is to say, in response to processing the calibrated perception data using the deep learning perception model and not obtaining the first-level target perception result, the multi-modal perception fusion system enters the obstacle comprehensive perception model stage. Based on a comprehensive analysis of the multi-modal data collected by each available sensor that has been spatio-temporally synchronized, filtered, data cleared, and jointly calibrated, the second-level target perception result is obtained, where the second-level target perception result is a supplement to the recognition process of the first-level target perception result.
[0143] As an extension of the deep learning perception model, the obstacle comprehensive perception model improves the problem that training samples cannot be exhausted, reduces the dependence of the multi-modal perception fusion method on labeled data, improves the generalization ability of the multi-modal perception fusion method, ensures the feasibility of applying the multi-modal perception fusion method to more roads and relatively complex transportation scenarios, enhances the adaptability of the multi-modal perception fusion method to complex environments, can identify newly added targets on the road, such as small animals, damaged manhole covers, balloons, and debris flows, etc., and also improves the target recognition ability of the multi-modal perception fusion method.
[0144] Regarding the obstacle comprehensive perception model, the calibrated perception data is processed to obtain the second-level target perception result, which is specifically as follows: based on the calibrated perception data of each available sensor, the geometric state, motion state, material, and material density of the target are identified; based on the geometric state, motion state, material, and material density of the target, the second-level target perception result is obtained.
[0145] By comprehensively considering the geometric state, motion state, material, and material density of the target, a detailed judgment of the target is realized, and the risk of missed detection of the target is reduced.
[0146] For example, the second-level target perception result may include target position, target speed, target category, target identifier, target confidence, target heading deviation angle, and may also include whether the target is suspended, whether the material of the target is metal or non-metal, etc.
[0147] In some embodiments, based on the obstacle comprehensive perception method (that is, when the first-level deep learning perception method fails, the second-level obstacle comprehensive perception method is activated), the purpose is to solve the long-tail problems of autonomous driving, such as missed detection and misdetection. Regarding the calibrated perception data of each available sensor, the geometric state of the target is identified, which is specifically as follows: clustering calculation is performed on the calibrated perception data to obtain the data clusters corresponding to the calibrated perception data; based on the data clusters, the surface of the target is reconstructed using a differentiable calculation optimization method, and the volume of the target is determined; the surface smoothness of the target is analyzed through the normal vector of the surface of the target and the curvature; the suspended state of the target is determined based on the contour coordinates of the target; based on the volume, surface smoothness, and suspended state of the target, the geometric state of the target is determined.
[0148] In some embodiments, if the change in the normal vector of the surface of the target and the curvature of the surface of the target is relatively small, it indicates that the surface of the target is relatively smooth. In addition, the braking of the corresponding vehicle or the change of the planned trajectory of the corresponding vehicle can be determined through the suspended state of the target.
[0149] For example, when the number of point clouds is small, the calibration perception data can be clustered by the OPTICS (Ordering Points To Identify the Clustering Structure) algorithm (such as the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm). By sorting the reachability distances, datasets with different densities can be processed. This method is applicable to geometric appearances with complex shapes. When the number of point clouds is large, the data space is divided into grid cells, and data clusters are formed by merging adjacent high-density cells. This method can quickly process large-scale calibration perception data and identify geometric shapes in the grid structure.
[0150] Clustering the calibration perception data by the DBSCAN algorithm does not require specifying the number of data clusters in advance, can handle data clusters with different densities, and has strong robustness to noise.
[0151] In some embodiments, regarding the use of differentiable computing optimization methods for surface reconstruction of an object, the following is specific: Triangulating the object's point cloud according to Delaunay triangulation, and combining the Greedy Projection Triangulation algorithm or Poisson algorithm in the PCL (Point Cloud Library) for surface reconstruction. The object can also be voxelized, and finally, the volume size information of the object is obtained by cumulative summation, as follows.
[0152] For the triangular mesh O = {o1, o2,..., o m}, the volume contribution of each triangle o j : where a, b, and c are the three vertex coordinates of the triangle o j , and the total volume is the sum of the volume contributions of all triangles: where m represents the number of triangular meshes.
[0153] In some embodiments, regarding calculating the normal and curvature of the surface of an object, for example, the normal of the surface of the object can be calculated based on principal component analysis (PCA). For each point cloud U i , the neighborhood point set is denoted as N i = {U i1 , U i2 ,..., U ik}, and then the centroid of the neighborhood points is calculated: For each point cloud U i , calculate and construct a covariance matrix: where k represents the number of point clouds in the neighborhood point set. Then perform eigenvalue decomposition on the covariance matrix Q to obtain λ1, λ2,..., λ m , and the corresponding eigenvectors are r1, r2,..., r m . Select the eigenvector corresponding to the smallest eigenvalue as the normal vector. For example, if λ m is the smallest eigenvalue, then n i = r m is the normal vector to determine the consistency of the normal direction. Calculate the curvature of the surface of the target using the eigenvalues, that is, calculate the minimum curvature corresponding to the covariance matrix: where it is assumed that λ m is the smallest eigenvalue, and λ1, λ2,..., λ m are the eigenvalues of the covariance matrix.
[0154] In some embodiments, regarding determining the suspended state of the target, first obtain the central point coordinates of the target and the outer contour coordinates of the target's top, bottom, left, and right (i.e., the contour coordinates of the target) (x i , y i , z i ), i = 1, 2, 3, 4. The suspended height of the target is: h = min(z i ) > 0. By determining whether the target is suspended, determine whether the corresponding vehicle brakes or changes the planned trajectory.
[0155] In some embodiments, the available sensors include lidar and millimeter-wave radar. Regarding the calibrated perception data based on each available sensor, identify the motion state of the target as follows: fuse the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar to obtain the motion state of the target.
[0156] For example, the lidar can obtain the position of the target, the millimeter-wave radar can obtain the motion speed of the target, and then combined with the Kalman filtering algorithm (Kalman filtering algorithm), the motion state of the target can be accurately obtained.
[0157] In some embodiments, the available sensors include lidar and millimeter-wave radar. Regarding the calibrated perception data based on each available sensor, identify the material of the target as follows: analyze the average value of the reflection intensity in the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar and apply it to identify the material of the target.
[0158] For example, the material of the target can be metal or non-metal. Analyzing the average value of the reflection intensity in the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar as described above can be to analyze the ratio of the average value of the reflection intensity to the square of the distance of the target, so as to identify the material of the target. For example, if the ratio is positive, the material of the target is metal; if the ratio is negative, the material of the target is non-metal. Among them, the reflection intensity is closely related to the material of the target, the surface roughness of the target, the incident angle of the target, and the distance of the target.
[0159] For the calibrated perception data based on each available sensor, identify the material density of the target to determine whether the target is porous. First, it is necessary to divide the point cloud into several small regions (such as voxels or grids), calculate the density, count the number of point clouds in each small region, calculate the point cloud density. If the point cloud density is less than the threshold of the point cloud density, then mark this small region as porous, otherwise, it is the same.
[0160] In some embodiments, regarding determining the material density of the target, specifically: divide the point cloud of the target into N voxel regions, and the volume size of each voxel is V vox+l =Δx×Δy×Δz. For the point cloud set of the target, it is denoted as E={e1,e2,...,e n}}, where the coordinates of any point cloud in the point cloud set of the target are E i =(x i ,y i ,z i ). Next, divide the point cloud control into voxels F={f ijk}}, where: f ijk ={(x,y,z)|x∈[iΔx,(i + 1)Δx],y∈[jΔy,(j + 1)Δy],z∈[jΔz,(j + 1)Δz]}. For each voxel f ijk , count the number of point clouds N ijk in this voxel, and calculate the point cloud density in this voxel
[0161] In addition, in the process of judging whether the target is porous, it is to judge whether each voxel of the target is porous. First, for each voxel, set the point cloud density threshold ρ th(+shold , and use S to represent whether this voxel is porous:
[0162] For example, for autonomous driving technology, based on the second-level target perception result, an obstacle avoidance strategy can be determined for mechanical equipment (such as unmanned fire trucks, unmanned trucks, etc.), that is, provide accurate decision-making information for the obstacle avoidance strategy.
[0163] The details are as follows: According to the geometric state of the target, the volume size, surface smoothness, and suspension state of the target can be obtained. If the target is small in volume and in a low-speed suspended state, it means that this target has little impact on the vehicle, and the planned trajectory of the corresponding vehicle can remain unchanged. If the target is a small stationary target on the ground, with a relatively smooth surface and a regular geometric shape, such as a tire or a manhole cover, the planned trajectory of the corresponding vehicle can remain unchanged, and the corresponding vehicle can pass by the target at a short distance horizontally. If the target is a moving target on the ground, with a relatively smooth surface and an irregular geometric shape, such as a small animal, at this time, the corresponding vehicle should decelerate or stop and wait for the target to pass before starting the corresponding vehicle. If the target is a large stationary target on the ground, with a relatively rough surface and an irregular geometric shape, such as a mudslide, the corresponding vehicle can decelerate or wait remotely for an operation instruction. Of course, with a high-level autonomous driving system embedded with a VLM (Vision-Language Model) large model, it can handle such situations better and basically does not require remote instruction control.
[0164] For example, for a stone pier by the roadside, the corresponding vehicle should decelerate and change the planned trajectory of the corresponding vehicle. If the target is a large stationary target on the ground, not suspended, with a relatively smooth surface and a regular geometric shape. For a metal obstacle, such as an iron rod by the roadside, the corresponding vehicle should decelerate and change the planned trajectory of the corresponding vehicle.
[0165] In some embodiments, based on the second-level target perception result, a target path can also be determined for the unmanned vehicle equipment to achieve the autonomous driving of the mechanical equipment, which helps the unmanned vehicle to develop in the direction of intelligence.
[0166] Figure 5 A schematic diagram showing some embodiments of the method of the obstacle comprehensive perception model of the present disclosure.
[0167] As Figure 5 shown, in the case where the deep learning target recognition fails 501, the obstacle comprehensive perception method 502 will be used to identify the target from the following four dimensions, which are: geometric state 503, motion state 504, material 505, and material density 506.
[0168] As Figure 5As shown, the geometric state 503 of the target is to obtain the shape 507 of the target according to a three-dimensional reconstruction method, determine the smoothness 508 of the surface of the target according to normal curvature analysis, and determine the suspended state 509 of the target according to the height of the lowest point of the target. The motion state 504 of the target is to determine the motion state information 512 of the target according to the position 510 of the target and the speed and motion direction 511 of the target. The reflection intensity of the target is closely related to the material, roughness, incident angle, and distance 513 of the target. According to the reflection intensity of the target, the material 505 of the target can be determined, that is, whether the target is a metal or a non-metal 514. The material density 506 of the target can determine whether the target is a porous material 516 (multi-void material) according to the relationship between the actual density of the point cloud and the point cloud density threshold 515.
[0169] As Figure 5 shown, according to the geometric state 503, motion state 504, material 505, and material density 506 of the target, the comprehensive judgment strategy 519 of the corresponding vehicle can be determined. Among them, according to the information 517 of the above four dimensions of the target, the second-level target perception result of the target is determined, and then the comprehensive judgment strategy 519 of the corresponding vehicle is determined according to the second-level target perception result of the target.
[0170] Figure 6 Schematic diagrams showing other embodiments of the multi-modal perception fusion method of the present disclosure.
[0171] As Figure 6 shown, the data input layer F1 of the multi-modal perception fusion method includes a front lidar (Front-Lidar), a left lidar (Left-Lidar), a right lidar (Rigte-Lidar), a rear lidar (Rear-Lidar), a front camera (Front-Camera), a rear camera (Rear-Camera), a front RTK (Front-RTK), a front millimeter-wave radar (Front-Rader), and a rear millimeter-wave radar (Rear-Rader). Among them, the lidar is used to obtain the three-dimensional point cloud data of the environment, the camera is used to obtain the two-dimensional image data of the environment, the millimeter-wave radar is used to obtain the motion state information of the target in the environment, and the RTK is used to obtain the high-precision positioning of the vehicle.
[0172] In addition, the data input layer may further include an IMU for obtaining the attitude information of the vehicle.
[0173] As Figure 6 shown, the sensor verification layer F2 verifies the data flow state of the data input layer (sensors) to dynamically select available sensors based on decoupling conditions. For example, the sensor verification layer can be a domain controller or an industrial computer. Figure 6A1, A2, A3, and A4 in it represent detecting (verifying) the data stream status of the sensor to determine whether the sensor is in a normal working state.
[0174] As Figure 6 shown, the preprocessing layer F3 performs spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multi-modal data collected by available sensors to obtain the calibrated perception data of each available sensor. Figure 6 The processing steps involved in M01 in it are the preprocessing performed by the preprocessing layer.
[0175] The processing steps involved in M01 include M01-01 to M01-15, which are as follows.
[0176] In step M01-01, the point clouds collected by the lidars in the front, rear, left, and right directions are preprocessed and de-distorted.
[0177] In step M01-02, the internal and external parameters of the data collected by the front and rear cameras are calculated.
[0178] In step M01-03, the data collected by the front and rear cameras are de-distorted.
[0179] In step M01-04, the data collected by the front RTK is preprocessed.
[0180] In step M01-05, the point clouds collected by the front and rear millimeter-wave radars are preprocessed.
[0181] In step M01-06, based on NTP (Network Time Protocol) or PTP (Precision Time Protocol), hard time synchronization is performed on the multi-modal data preprocessed in M01-01 to M01-05.
[0182] In step M01-07, soft algorithmic synchronization is performed on the multi-modal data preprocessed in M01-01 to M01-05.
[0183] Hard time synchronization and soft algorithmic synchronization are parallel steps, and either hard time synchronization or soft algorithmic synchronization can be selected for execution when performing spatio-temporal synchronization on the multi-modal data preprocessed in M01-01 to M01-05.
[0184] In step M01-08, joint calibration of lidar-RTK is performed.
[0185] In step M01-09, joint calibration of front camera-lidar is performed.
[0186] In step M01-10, joint calibration is performed on the rear camera - lidar.
[0187] In step M01-11, joint calibration is performed on the front camera - millimeter wave radar.
[0188] In step M01-12, joint calibration is performed on the rear camera - millimeter wave radar.
[0189] In step M01-13, joint calibration is performed on the lidar - lidar, where the joint calibration is an adaptive joint calibration, that is, self - perform joint calibration until the calibration result meets the requirements, and correction will also be performed periodically.
[0190] In step M01-14, the transformation matrix (i.e., the joint calibration transformation matrix parameter) is output.
[0191] In step M01-15, based on the transformation matrix, the point cloud to be registered is transformed so that the point cloud to be registered is transformed into the ego - vehicle coordinate system (reference point cloud coordinate system).
[0192] As Figure 6 shown, through the perception layer F4, the environment is perceived, Figure 6 and the steps involved in M02 in
[0193] The processing steps for M02 include M02-01 to M02-09, which are as follows.
[0194] In step M02-01, the point clouds detected by multiple lidars are stitched.
[0195] In some embodiments, stitching the point clouds detected by multiple lidars refers to calibrating the projection data corresponding to the point clouds detected by the lidars based on the transformation matrix obtained from the adaptive joint calibration to obtain the calibrated perception data of each lidar.
[0196] In step M02-02, the lidar performs 3D object detection to determine Target 1.
[0197] In step M02-03, the lidar performs 3D semantic segmentation to determine the road edge line.
[0198] In step M02-04, the front camera performs object detection.
[0199] In step M02-05, the rear camera performs object detection.
[0200] In step M02-06, the front and rear cameras perform BEV object detection on the images to determine Target 2.
[0201] In step M02-07, the front millimeter-wave radar performs 2D target detection.
[0202] In step M02-08, the rear millimeter-wave radar performs 2D target detection.
[0203] In step M02-09, the detections of the front and rear millimeter-wave radars are merged to determine target 3.
[0204] In the steps involved in M02, target 1, target 2, and target 3 are targets detected by different sensors, which can be the same target or different targets. In the subsequent fusion layer F5, target 1, target 2, and target 3 are considered to be the same target.
[0205] Such as Figure 6 shown, through the fusion layer F5, the multi-modal data after joint calibration can be fused.
[0206] Figure 6 The processing steps involved in M04 in are multi-modal perception fusion methods, including two methods, namely the B multi-sensor perception fusion method and the C obstacle comprehensive perception method. Among them, the B multi-sensor perception fusion method includes M04-01 and M04-02, as follows.
[0207] The jointly calibrated multi-modal data (i.e., the perception data corresponding to target 1, target 2, and target 3) is detected by a deep learning perception model. If the first-level target perception results of each target detected by each available sensor can be obtained, the B multi-sensor perception fusion method is executed. If the first-level target perception results of each target detected by each available sensor cannot be obtained, the C obstacle comprehensive perception method is executed. Among them, the target can be, for example, a pedestrian, a truck, a private car, etc.
[0208] Regarding the B multi-sensor perception fusion method, in step M04-01, the first-level target perception results of each target detected by the deep learning perception model are fused to obtain a fused target result. In step M04-02, the fused target result is merged with the road edge line.
[0209] In some embodiments, the multi-modal perception fusion method (M04) is executed first, and then SLAM (M03) is executed. During the process of constructing the SLAM cost map, the data after perception fusion (i.e., the fusion target result or the second-level target perception result), the calibrated perception data detected by the lidar, and the calibrated perception data detected by RTK are required. Among them, the process of constructing the SLAM cost map is as follows. First, a global map and a local map are constructed based on the calibrated perception data detected by the lidar and the calibrated perception data detected by RTK, and then the data after perception fusion (the target from a bird's-eye view) is projected onto the local map to obtain the SLAM cost map.
[0210] Figure 6 The processing steps involved in M03 are SLAM (Simultaneous Localization and Mapping), including M03-01 to M03-04, which are as follows.
[0211] In step M03-01, the position of the vehicle is accurately perceived through the integrated navigation real-time positioning technology. Among them, the vehicle is a vehicle equipped with a variety of sensors, such as an unmanned fire truck, an unmanned truck, etc.
[0212] In step M03-02, the pose of the vehicle is calculated.
[0213] In step M03-03, the foreground target and the background are separated through the point cloud foreground and background separation technology.
[0214] In step M03-04, the SLAM cost map is constructed. Among them, the SLAM cost map is a binary grid, and different colors indicate passable or non-passable. For example, black indicates non-passable, and white indicates passable.
[0215] As Figure 6 shown, the result of perception fusion can be displayed through the output layer F6, and the decision-making plan M05 (i.e., path planning) generated based on the result of perception fusion and the constructed SLAM cost map can also be displayed.
[0216] Figure 6 The processing steps involved in M06 are perception visualization, including M06-01 to M06-04. In step M06-01, the 2D target is visualized. In step M06-02, the 3D target is visualized. In step M06-03, the BEV corresponding to the fusion target result is visualized. In step M06-04, the road edge line is visualized.
[0217] In the above embodiments, by obtaining multimodal data collected by at least two types of sensors, it provides feasibility for subsequent fusion of data collected by different types of sensors and ensures the accuracy of the obtained perception data. The sensor verification layer verifies the data stream status of each sensor to dynamically select available sensors based on decoupling conditions, enabling flexible selection of available sensors that can participate in the fusion according to the real-time working status of the sensors, realizing a decoupling mechanism in the perception fusion process. Multiple sensors can both couple and decouple, reflecting the intelligent automatic judgment in the perception coupling process, ensuring the flexibility and stability in the perception fusion process, and reducing the risk of inaccurate fusion target results caused by abnormal data collected by abnormal sensors during the perception fusion process. By performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multimodal data collected by the available sensors to obtain the calibrated perception data of each available sensor, it ensures the accuracy of the multimodal perception fusion method and reduces the risk of affecting the accuracy and stability of the multimodal perception fusion method due to problems such as spatial inconsistency, time asynchrony, and multimodal data conflicts.
[0218] Figure 7 Schematic diagrams showing some embodiments of the multimodal perception fusion device of the present disclosure.
[0219] As Figure 7 shown, the multimodal perception fusion device 70 includes an acquisition perception module 71, a data verification module 72, a joint calibration module 73, a multimodal fusion module 74, and an output fusion module 75.
[0220] The acquisition perception module 71 is configured to acquire multimodal data collected by at least two types of sensors.
[0221] The data verification module 72 is configured to verify the data stream status of each sensor through sensor verification metrics and dynamically select available sensors based on decoupling conditions.
[0222] For example, the sensor verification metrics include timestamp update status, network transmission status, and device physical status.
[0223] The joint calibration module 73 is configured to perform spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multimodal data collected by the available sensors to obtain the calibrated perception data of each available sensor.
[0224] The multimodal fusion module 74 is configured to process the calibrated perception data in response to using a deep learning perception model, and can obtain the first-level target perception results of each target detected by each available sensor, and fuse the first-level target perception results using multi-dimensional data association rules to obtain the fusion target results.
[0225] The output fusion module 75 is configured to output a fused target result.
[0226] In the above embodiments, by acquiring multi-modal data collected by at least two types of sensors, it provides feasibility for subsequent fusion of data collected by different types of sensors and ensures the accuracy of the acquired perception data; by using sensor calibration metrics to verify the data stream status of each sensor, and dynamically selecting available sensors based on decoupling conditions, it can flexibly select available sensors that can participate in the fusion according to the real-time working status of the sensors, realizing the decoupling mechanism in the perception fusion process. Multiple sensors can be coupled or decoupled to work, reflecting the intelligent automatic judgment in the perception coupling process, ensuring the flexibility and stability in the perception fusion process, and reducing the risk of inaccurate fused target results caused by abnormal data collected by abnormal sensors during the perception fusion process; by performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multi-modal data collected by the available sensors to obtain the calibrated perception data of each available sensor, it ensures the accuracy of the multi-modal perception fusion method and reduces the risk of affecting the accuracy and stability of the multi-modal perception fusion method due to problems such as spatial inconsistency, time asynchrony, and multi-modal data conflicts; by processing the calibrated perception data through a deep learning perception model to obtain the first-level target perception results of each target detected by each available sensor, it realizes the differentiation of different targets and provides feasibility for subsequent fusion of the first-level target perception results for each target; by adopting multi-dimensional data association rules to fuse the first-level target perception results, obtaining and outputting the fused target result, it realizes the fusion of multi-modal data for each target and ensures the accuracy of the fused target result of each target obtained.
[0227] In some embodiments, the sensor calibration metrics include the timestamp update status, network transmission status, and device physical status. The data verification module 72 is further configured to determine a sensor whose at least one of the timestamp update status, network transmission status, and device physical status meets a failure condition as a non-available sensor; if the number of non-available sensors of the same type is greater than a threshold, then determine the sensors of the same type as the non-available sensors with a number greater than the threshold as non-available sensors; based on the non-available sensors, select available sensors from at least two types of sensors.
[0228] In some embodiments, the data verification module 72 is further configured to determine a non-available sensor as an available sensor when the timestamp update status, network transmission status, and device physical status of the non-available sensor do not meet the failure conditions.
[0229] In some embodiments, the joint calibration module 73 is further configured to determine the initial installation pose of each available sensor; perform horizontal projection correction on the multi-modal data of each available sensor to obtain the projection data of each available sensor; select a primary sensor from the available sensors, and perform sensor registration on slave sensors in the available sensors based on the primary sensor; perform data registration on the projection data of each available sensor to obtain joint calibration transformation matrix parameters; and perform calibration on the projection data of each available sensor based on the joint calibration transformation matrix parameters to obtain the calibrated perception data of each available sensor.
[0230] In some embodiments, the joint calibration module 73 is further configured to perform first registration on the projection data based on the normal distribution transformation algorithm until first registration data that meets the first requirement is obtained; perform second registration on the first registration data based on the iterative closest point algorithm until second registration data that meets the second requirement is obtained; and determine the joint calibration transformation matrix parameters based on the second registration data and the projection data.
[0231] In some embodiments, the joint calibration module 73 is further configured to periodically correct the joint calibration transformation matrix parameters.
[0232] In some embodiments, when the available sensors include multiple types of sensors, the multi-modal fusion module 74 is further configured to adopt a multi-modal combination fusion method to fuse the first-level target perception results to obtain a fused target result; when the available sensors include a single type of sensor, adopt a single-modal fusion method to fuse the first-level target perception results to obtain a fused target result.
[0233] In some embodiments, the multi-modal fusion module 74 is further configured to perform the multi-modal combination fusion method and the single-modal fusion method based on a bird's-eye view.
[0234] In some embodiments, the first-level target perception results include at least two of the motion attribute, detection box attribute, historical trajectory attribute, and appearance semantic attribute of each target. The multi-modal fusion module 74 is further configured to calculate the similarity of at least two corresponding to the first-level target perception results among the motion attribute similarity, detection box attribute similarity, historical trajectory attribute similarity, and appearance semantic attribute similarity between every two targets detected by different sensors in the available sensors; perform weighted calculation on the similarity of at least two corresponding to the first-level target perception results among the motion attribute similarity, detection box attribute similarity, historical trajectory attribute similarity, and appearance semantic attribute similarity to obtain the similarity between every two targets detected by different sensors; determine the same target detected by different sensors based on the similarity between every two targets detected by different sensors; and perform weighted fusion on multiple first-level target perception results of the same target to obtain the fused target result of the same target.
[0235] In some embodiments, the multimodal perception fusion device 70 further includes a processing module configured to process the calibrated perception data in response to using a deep learning perception model to obtain a first-level target perception result, and process the calibrated perception data based on an obstacle comprehensive perception model to obtain a second-level target perception result; and output the second-level target perception result.
[0236] In some embodiments, the processing module is further configured to identify the geometric state, motion state, material, and material density of the target based on the calibrated perception data of each available sensor; and obtain a second-level target perception result based on the geometric state, motion state, material, and material density of the target.
[0237] In some embodiments, the processing module is further configured to perform clustering calculation on the calibrated perception data to obtain data clusters corresponding to the calibrated perception data; based on the data clusters, use a differentiable computing optimization method to perform surface reconstruction on the target and determine the volume of the target; analyze the surface smoothness of the target through the normal and curvature of the surface of the target; determine the suspended state of the target based on the contour coordinates of the target; and determine the geometric state of the target based on the volume, surface smoothness, and suspended state of the target.
[0238] In some embodiments, the processing module is further configured to fuse the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar to obtain the motion state of the target.
[0239] In some embodiments, the processing module is further configured to analyze the average value of the reflection intensities in the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar to identify the material of the target.
[0240] In some embodiments, the multimodal perception fusion device 70 further includes a determination module configured to determine an obstacle avoidance strategy based on the fusion target result or the second-level target perception result.
[0241] Figure 8 A schematic diagram showing other embodiments of the multimodal perception fusion device of the present disclosure.
[0242] As Figure 8 shown, the multimodal perception fusion device 70 of this embodiment includes: a memory 81 and a processor 82 coupled to the memory 81. The processor 82 is configured to execute the multimodal perception fusion method in any of the foregoing embodiments based on instructions stored in the memory 81.
[0243] The memory 81 can include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.
[0244] The multi-modal perception fusion device 70 can also include an input / output interface 83, a network interface 84, a storage interface 85, etc. These interfaces 83, 84, 85, the memory 81, and the processor 82 can be connected through a bus 86, for example. Among them, the input / output interface 83 provides a connection interface for input / output devices such as a display, a mouse, a keyboard, a touch screen, a microphone, and a speaker. The network interface 84 provides a connection interface for various networking devices. The storage interface 85 provides a connection interface for external storage devices such as an SD card and a USB flash drive.
[0245] In the above embodiments, by acquiring multi-modal data collected by at least two types of sensors, it provides feasibility for subsequent fusion of data collected by different types of sensors and ensures the accuracy of the acquired perception data; by verifying the data stream status of each sensor through sensor verification metrics and dynamically selecting available sensors based on decoupling conditions, it can flexibly select available sensors that can participate in the fusion according to the real-time working status of the sensors, realizing the decoupling mechanism in the perception fusion process. Multiple sensors can be both coupled and decoupled, reflecting the intelligent automatic judgment in the perception coupling process, ensuring the flexibility and stability in the perception fusion process, and reducing the risk of inaccurate fusion target results caused by abnormal data collected by abnormal sensors during the perception fusion process; by performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multi-modal data collected by available sensors to obtain the calibrated perception data of each available sensor, it ensures the accuracy of the multi-modal perception fusion method and reduces the risk of affecting the accuracy and stability of the multi-modal perception fusion method due to problems such as spatial inconsistency, time asynchrony, and multi-modal data conflicts; by processing the calibrated perception data through a deep learning perception model to obtain the first-level target perception results of each target detected by each available sensor, it realizes the differentiation of different targets and provides feasibility for subsequent fusion of the first-level target perception results for each target; by adopting multi-dimensional data association rules to fuse the first-level target perception results and obtain and output the fusion target results, it realizes the fusion of multi-modal data for each target and ensures the accuracy of the fusion target results of each obtained target.
[0246] Figure 9 Schematic diagrams showing some embodiments of the multi-modal perception fusion system of the present disclosure.
[0247] As Figure 9As shown, the multimodal perception fusion system 90 includes the multimodal perception fusion device 70 in any of the above embodiments and a sensor combination 91.
[0248] The sensor combination 91 includes at least two different types of sensors and is configured to collect and send multimodal data to the multimodal perception fusion device.
[0249] In some embodiments, the sensor combination includes at least two of lidar, camera, millimeter-wave radar, real-time kinematic differential positioning, and inertial measurement unit.
[0250] In the above embodiments, by obtaining multimodal data collected by at least two types of sensors, it provides feasibility for subsequent fusion of data collected by different types of sensors and ensures the accuracy of the obtained perception data; by using sensor calibration metrics to verify the data stream status of each sensor and dynamically select available sensors based on decoupling conditions, it can flexibly select available sensors that can participate in the fusion according to the real-time working state of the sensors, realizing a decoupling mechanism in the perception fusion process. Multiple sensors can work both in a coupled and decoupled manner, reflecting the intelligent automatic judgment in the perception coupling process, ensuring the flexibility and stability in the perception fusion process, and reducing the risk of inaccurate fusion target results caused by abnormal data collected by abnormal sensors during the perception fusion process; by performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multimodal data collected by available sensors to obtain the calibrated perception data of each available sensor, it ensures the accuracy of the multimodal perception fusion method and reduces the risk of affecting the accuracy and stability of the multimodal perception fusion method due to problems such as spatial inconsistency, time asynchrony, and multimodal data conflict; by processing the calibrated perception data through a deep learning perception model to obtain the first-level target perception results of each target detected by each available sensor, it realizes the differentiation of different targets and provides feasibility for subsequent fusion of the first-level target perception results for each target; by adopting multi-dimensional data association rules to fuse the first-level target perception results and obtain and output the fused target results, it realizes the fusion of multimodal data for each target and ensures the accuracy of the fused target results of each obtained target.
[0251] Figure 10 Schematic diagrams showing some embodiments of the unmanned fire truck of the present disclosure.
[0252] As Figure 10 shown, the unmanned fire truck 100 includes the multimodal perception fusion system 90 in any of the above embodiments, a sensor adaptation interface 101, a target recognition database 102 for fire fighting tasks, and a fusion weight adjustment strategy library 103.
[0253] The sensor adaptation interface 101 is configured to be connected to the sensor combination.
[0254] The target recognition database 102 for fire protection tasks is configured to provide sample data for training the deep learning perception model.
[0255] The fusion weight adjustment policy library 103 is configured to provide weights for fusing multiple first-level target perception results of the same target.
[0256] In the above embodiments, by obtaining multi-modal data collected by at least two types of sensors, it provides feasibility for subsequent fusion of data collected by different types of sensors and ensures the accuracy of the obtained perception data; by using the sensor verification index to verify the data stream status of each sensor, and dynamically selecting available sensors based on the decoupling condition, it can flexibly select available sensors that can participate in the fusion according to the real-time working status of the sensors, realizing the decoupling mechanism in the perception fusion process. Multiple sensors can work both coupled and decoupled, reflecting the intelligent automatic judgment in the perception coupling process, ensuring the flexibility and stability in the perception fusion process, and reducing the risk that the fusion target result is inaccurate due to abnormal data collected by abnormal sensors in the perception fusion process; by performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multi-modal data collected by the available sensors to obtain the calibrated perception data of each available sensor, it ensures the accuracy of the multi-modal perception fusion method and reduces the risk that the accuracy and stability of the multi-modal perception fusion method are affected by problems such as spatial inconsistency, time asynchrony, and multi-modal data conflict; by processing the calibrated perception data through the deep learning perception model to obtain the first-level target perception result of each target detected by each available sensor, it realizes the differentiation of different targets and provides feasibility for subsequent fusion of the first-level target perception results for each target; by adopting multi-dimensional data association rules to fuse the first-level target perception results and obtain and output the fusion target result, it realizes the fusion of multi-modal data for each target and ensures the accuracy of the fusion target result of each obtained target.
[0257] In some embodiments, a computer program product is protected, including a computer program or instruction, which when executed by a processor implements the above multi-modal perception fusion method. The computer program product includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network by the multi-modal perception fusion device, or installed from the storage device, or installed from the ROM. When the computer program is executed by the CPU, it executes the above functions defined in the method of the embodiments of the present disclosure.
[0258] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0259] So far, the multi-modal perception fusion method, device, system, and unmanned fire truck of the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details known in the art are not described. Those skilled in the art can fully understand how to implement the technical solutions disclosed here based on the above description.
[0260] The methods and systems of the present disclosure can be implemented in many ways. For example, the methods and systems of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0261] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A multimodal perception fusion method, comprising: Obtaining multimodal data collected by at least two types of sensors; Verifying the data stream status of each sensor through sensor verification metrics, and dynamically selecting available sensors based on decoupling conditions; Performing spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multimodal data collected by the available sensors to obtain calibrated perception data for each available sensor; In response to processing the calibrated perception data using a deep learning perception model to obtain a first-level target perception result for each target detected by each available sensor, using multi-dimensional data association rules to fuse the first-level target perception results to obtain a fused target result; Outputting the fused target result.
2. The multimodal perception fusion method according to claim 1, further comprising: In response to processing the calibrated perception data using the deep learning perception model and not obtaining the first-level target perception result, processing the calibrated perception data based on an obstacle comprehensive perception model to obtain a second-level target perception result; Outputting the second-level target perception result.
3. The multimodal perception fusion method according to claim 1, wherein, The sensor verification metrics include timestamp update status, network transmission status, and device physical status, The verifying the data stream status of each sensor through sensor verification metrics and dynamically selecting available sensors based on decoupling conditions includes: Determining a sensor with at least one of the timestamp update status, the network transmission status, and the device physical status satisfying a failure condition as a non-available sensor; If the number of non-available sensors of the same type is greater than a threshold, determining sensors of the same type as the non-available sensors with the number greater than the threshold as the non-available sensors; Selecting the available sensors from the at least two types of sensors based on the non-available sensors.
4. The multimodal perception fusion method according to claim 3, wherein The verifying the data stream status of each sensor through sensor verification metrics and dynamically selecting available sensors based on decoupling conditions further includes: In the case where the timestamp update status, the network transmission status, and the device physical status of the non-available sensors do not satisfy the failure condition, determining the non-available sensors as the available sensors.
5. The multimodal perception fusion method according to claim 1, wherein, The using multi-dimensional data association rules to fuse the first-level target perception results to obtain a fused target result includes: In the case where the available sensors include multiple types of sensors, using a multimodal combination fusion method to fuse the first-level target perception results to obtain the fused target result; In the case where the available sensors include a single type of sensor, using a unimodal fusion method to fuse the first-level target perception results to obtain the fused target result.
6. The multimodal perception fusion method according to claim 5, wherein, The using multi-dimensional data association rules to fuse the first-level target perception results to obtain a fused target result further includes: Performing the multimodal combination fusion method and the unimodal fusion method from an aerial view perspective.
7. The multimodal perception fusion method according to claim 1, wherein, The performing joint calibration on the multimodal data collected by the available sensors includes: Determine the initial installation pose of each of the available sensors; Perform horizontal projection correction on the multi-modal data of each of the available sensors to obtain the projection data of each of the available sensors; Select a master sensor from the available sensors, and perform sensor registration on the slave sensors in the available sensors based on the master sensor; Perform data registration on the projection data of each of the available sensors to obtain the combined calibration transformation matrix parameters; Based on the combined calibration transformation matrix parameters, calibrate the projection data of each of the available sensors to obtain the calibrated perception data of each of the available sensors.
8. The multimodal perception fusion method according to claim 7, wherein, The performing data registration on the projection data of each of the available sensors to obtain the combined calibration transformation matrix parameters includes: Based on the normal distribution transformation algorithm, perform first registration on the projection data until first registration data that meets the first requirement is obtained; Based on the iterative closest point algorithm, perform second registration on the first registration data until second registration data that meets the second requirement is obtained; Based on the second registration data and the projection data, determine the combined calibration transformation matrix parameters.
9. The multimodal perception fusion method according to claim 7, wherein, The performing combined calibration on the multi-modal data collected by the available sensors further includes: Periodically correct the combined calibration transformation matrix parameters.
10. The multimodal perception fusion method according to claim 2, wherein, The processing the calibrated perception data based on the obstacle comprehensive perception model to obtain the second-level target perception result includes: Based on the calibrated perception data of each of the available sensors, identify the geometric state, motion state, material, and material density of the target; Based on the geometric state, motion state, material, and material density of the target, obtain the second-level target perception result.
11. The multimodal perception fusion method according to claim 10, wherein, The identifying the geometric state of the target based on the calibrated perception data of each of the available sensors includes: Perform clustering calculation on the calibrated perception data to obtain data clusters corresponding to the calibrated perception data; Based on the data clusters, use a differentiable computing optimization method to perform surface reconstruction on the target and determine the volume of the target; Analyze the surface smoothness of the target through the normal and curvature of the surface of the target; Based on the contour coordinates of the target, determine the suspended state of the target; Based on the volume, surface smoothness, and suspended state of the target, determine the geometric state of the target.
12. The multimodal perception fusion method according to claim 10, wherein, The available sensors include a lidar and a millimeter-wave radar, and the identifying the motion state of the target based on the calibrated perception data of each of the available sensors includes: Fuse the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar to obtain the motion state of the target.
13. The multimodal perception fusion method according to claim 10, wherein, The available sensors include a lidar and a millimeter-wave radar, and the identifying the material of the target based on the calibrated perception data of each of the available sensors includes: Analyze the average value of the reflection intensity in the calibrated perception data of the lidar and the calibrated perception data of the millimeter-wave radar to identify the material of the target.
14. The multi-modal perception fusion method according to claim 2, further comprising: Determine an obstacle avoidance strategy based on the fused target result or the second-level target perception result.
15. The multimodal perception fusion method according to any one of claims 1 to 14, wherein, The first-level target perception results include at least two of the motion attributes, detection box attributes, historical trajectory attributes, and appearance semantic attributes of each target. Using multi-dimensional data association rules to fuse the first-level target perception results, the obtained fused target results include: Calculating the similarity of at least two corresponding to the first-level target perception results among the motion attribute similarity, detection box attribute similarity, historical trajectory attribute similarity, and appearance semantic attribute similarity between every two targets detected by different sensors in the available sensors; Performing weighted calculation on the similarity of at least two corresponding to the first-level target perception results among the motion attribute similarity, the detection box attribute similarity, the historical trajectory attribute similarity, and the appearance semantic attribute similarity to obtain the similarity between every two targets detected by different sensors; Based on the similarity between every two targets detected by different sensors, determining the same target detected by different sensors; Performing weighted fusion on multiple first-level target perception results of the same target to obtain the fused target result of the same target.
16. A multimodal perception fusion device, comprising: An acquisition perception module configured to acquire multimodal data collected by at least two types of sensors; A data verification module configured to verify the data stream status of each sensor through sensor calibration metrics and dynamically select available sensors based on decoupling conditions; A joint calibration module configured to perform spatio-temporal synchronization, filtering, data cleaning, and joint calibration on the multimodal data collected by the available sensors to obtain the calibrated perception data of each available sensor; A multimodal fusion module configured to, in response to using a deep learning perception model to process the calibrated perception data, be able to obtain the first-level target perception results of each target detected by each available sensor, and use multi-dimensional data association rules to fuse the first-level target perception results to obtain fused target results; An output fusion module configured to output the fused target results.
17. A multimodal perception fusion device, comprising: A memory; And A processor coupled to the memory, the processor being configured to execute the multimodal perception fusion method according to any one of claims 1 to 15 based on instructions stored in the memory.
18. A multimodal perception fusion system, comprising: The multimodal perception fusion device according to claim 16 or 17; A sensor combination including at least two different types of sensors, the sensor combination being configured to collect and send multimodal data to the multimodal perception fusion device.
19. The multimodal perception fusion system according to claim 18, wherein, The sensor combination includes at least two of lidar, camera, millimeter-wave radar, real-time kinematic differential positioning, and inertial measurement unit.
20. An unmanned fire truck, comprising: Equipped with the multimodal perception fusion system according to claim 18 or 19; A sensor adaptation interface configured to be connected to the sensor combination; A target recognition database for fire fighting tasks configured to provide sample data for training the deep learning perception model. The fusion weight adjustment policy library is configured to provide weights for fusing multiple first-level target perception results of the same target.
21. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the multi-modal perception fusion method according to any one of claims 1 to 15.
22. A computer program product comprising computer instructions, which, when executed by a processor, implement the multi-modal perception fusion method according to any one of claims 1 to 15.