A Multi-Source Sensor Data Fusion Mapping Method and Vehicle
By using probability distribution processing and data fusion of active non-contact ranging sensors and vision sensors, the problems of high cost and low robustness in automatic parking systems are solved, achieving high-precision and low-cost environmental perception and map generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING CHANGAN AUTOMOBILE CO LTD
- Filing Date
- 2025-06-24
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, LiDAR and camera solutions are costly and require a lot of computing power in automatic parking systems, and the generated maps have low robustness and high uncertainty in single-frame data fusion mapping.
Active non-contact ranging sensors and vision sensors are used to generate grid maps with semantics and occupancy probabilities through probability distribution processing and data fusion. This includes a combination of ultrasonic radar and millimeter-wave radar with fisheye cameras, and data fusion is performed using Bayesian fusion and DS evidence synthesis methods.
It improves data accuracy and recognition precision, reduces costs, and achieves highly robust environmental perception and map generation on low-computing-power platforms.
Smart Images

Figure CN120351923B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of assisted driving environment perception technology, and in particular to a multi-source sensor data fusion mapping method and a vehicle. Background Technology
[0002] Automated parking technology aims to enable vehicles to automatically identify parking spaces and complete the parking operation. Automated parking control systems based on multi-information fusion are intelligent systems developed to solve a series of problems in the automated parking process.
[0003] In related technologies, LiDAR and camera solutions have been widely used in some intelligent driving solutions. However, LiDAR and vision solutions have high computing power requirements, and if a large coverage area is needed, multiple LiDAR blind spot radars are required, which not only increases costs but also further drives up computing power requirements. In addition, single-frame data fusion mapping has high uncertainty and the generated map has low robustness. Summary of the Invention
[0004] This application provides a multi-source sensor data fusion mapping method and vehicle to solve the problem of low robustness of maps generated in related technologies.
[0005] In a first aspect, embodiments of this application provide a multi-source sensor data fusion mapping method, which includes:
[0006] Acquire sensor data from active non-contact ranging sensors and vision sensors to obtain the probability distribution of measurement points in the current frame for each type of sensor;
[0007] Based on the probability distribution of the measurement points in the current frame of the visual sensor and the confidence distribution corresponding to each measurement point, the first probability distribution of the measurement points in the current frame of the visual sensor is obtained;
[0008] The probability distribution of the measurement points in the current frame of the active non-contact ranging sensor is fused with the probability distribution of the measurement points in the previous data frames of a set number of frames to obtain the first fused probability distribution of the active non-contact ranging sensor.
[0009] The first probability distribution of the measurement points of the current frame of the visual sensor is fused with the first probability distribution of the measurement points of the previous data frames of a set number of frames to obtain the first fused probability distribution of the visual sensor.
[0010] The first fusion probability distribution of each type of sensor is updated to the fusion data layer, and probability fusion is performed to obtain the second fusion probability distribution;
[0011] Based on the second fusion probability distribution and the target semantic information acquired by the visual sensor, a raster map with semantics and occupancy probability is obtained.
[0012] Furthermore, the active non-contact ranging sensor includes a first ranging sensor and a second ranging sensor, and the detection distance of the first ranging sensor is less than the detection distance of the second ranging sensor.
[0013] Furthermore, the first ranging sensor is an ultrasonic radar, and the second ranging sensor is a millimeter-wave radar.
[0014] Furthermore, sensor data from active non-contact ranging sensors and vision sensors are acquired to obtain the probability distribution of measurement points in the current frame for each type of sensor, including:
[0015] Acquire sensor data from active non-contact ranging sensors and vision sensors;
[0016] Based on the sensor data from the active non-contact ranging sensor and the vision sensor, the measurement point data of the active non-contact ranging sensor and the vision sensor are obtained.
[0017] Based on the measurement point data of the active non-contact ranging sensor and the vision sensor, and the corresponding uncertainty model, the probability distribution of the measurement points in the current frame of the active non-contact ranging sensor and the vision sensor is obtained.
[0018] Furthermore, before obtaining the probability distribution of measurement points in the current frame of the active non-contact ranging sensor and the vision sensor, the method further includes:
[0019] The measurement point data of the image captured by the vision sensor is obtained based on the sensor data of the vision sensor;
[0020] The type of target within the field of view of the visual sensor and the position of the target relative to the visual sensor are determined based on the measurement point data of the image.
[0021] The position of the target relative to the vehicle body is obtained by converting intrinsic and extrinsic parameters based on the position data of the target relative to the visual sensor and the position and angle of the visual sensor on the vehicle.
[0022] Based on the type of target within the field of view of the vision sensor, the target's position relative to the vehicle body, and the corresponding uncertainty model, the probability distribution of the target in the vehicle coordinate system is obtained.
[0023] Furthermore, the method also includes obtaining the confidence distribution, which specifically includes:
[0024] A confidence distribution corresponding to the measurement points is generated according to a first rule, which includes:
[0025] If the measurement point is a ground element, the greater the distance from the measurement point to the visual sensor, the lower the confidence level.
[0026] Furthermore, the method also includes obtaining the confidence distribution, which specifically includes:
[0027] The confidence distribution corresponding to the measurement points is generated according to the second rule, which includes:
[0028] The confidence level of measurement points in the ground-accessible area is lower when the foreground of the non-ground element target is isolated and the foreground is obscured at a higher position.
[0029] Furthermore, the foreground for separating non-ground element targets includes:
[0030] The probability distribution of the separation point between the target and the ground is obtained based on the preset obstacle separation model and the types of non-ground element targets;
[0031] Based on the probability distribution of the segmentation points, the non-ground element target near the ground is separated;
[0032] The height of the visual sensor's observation position is obtained based on the angle between the observation position of the visual sensor and the segmentation point.
[0033] The confidence level of measurement points in the drivable ground area that is obscured by the higher position of the foreground is configured to be lower.
[0034] Furthermore, after acquiring sensor data from the active non-contact ranging sensor and the vision sensor, and obtaining the probability distribution of measurement points in the current frame for each type of sensor, the method further includes:
[0035] Determine whether the type of target observed by the visual sensor belongs to the ground element target subclass or the non-ground element target subclass;
[0036] If the target subclass is a ground element target, then generate the probability distribution of the ground element target in the vehicle coordinate system.
[0037] Further, obtaining the first probability distribution of the measurement points in the current frame of the visual sensor based on the probability distribution of the measurement points and the confidence distribution corresponding to each measurement point includes:
[0038] Based on the probability of the measurement point in the current frame of the visual sensor, the confidence level corresponding to the measurement point, and the third rule, the first probability distribution of the measurement point in the current frame of the visual sensor is obtained;
[0039] The third rule includes: the lower the confidence level of the region, the lower the first probability that the measurement point corresponding to the current frame is occupied or empty, and the first probability is lower than the probability that the measurement point in the current frame is occupied or empty.
[0040] Further, updating the first fusion probability distribution of each type of sensor to the fusion data layer and performing probability fusion to obtain the second fusion probability distribution includes:
[0041] Based on the first fusion probability distribution of each type of sensor, the effective detection range of each type of sensor, the confidence level of each type of sensor in the space being occupied, and the conflict factor of the sensor observation results, the occupancy results of each measurement point in the fusion data layer are obtained, and the occupancy results include occupied, empty, and unknown.
[0042] A second fusion probability distribution is formed based on the occupancy results of each measurement point.
[0043] Furthermore, after obtaining the probability distribution of measurement points in the current frame for each type of sensor, the method further includes:
[0044] Based on the first moment of acquiring sensor data from the active non-contact ranging sensor and the vision sensor, and the vehicle motion data between the current moment and the first moment, the probability distribution of the active non-contact ranging sensor measurement points and the first probability distribution of the vision sensor measurement points at the first moment are corrected to the current moment.
[0045] Furthermore, after correcting the probability distribution of the active non-contact ranging sensor measurement points and the first probability distribution of the vision sensor measurement points at the first moment to the current moment, the method further includes:
[0046] Get the first position of the preset type of target in the world coordinate system in the current frame;
[0047] Obtain the second position of the target of the preset type in the world coordinate system in the previous frame;
[0048] The vehicle's attitude in the grid map is corrected based on the difference between the first position and the second position, wherein the position of the preset type target in the world coordinate system is obtained based on its position in the vehicle coordinate system and the position of the vehicle in the world coordinate system.
[0049] Furthermore, before obtaining a raster map with semantics and occupancy probability based on the second fusion probability distribution and the target semantic information acquired by the visual sensor, the method further includes:
[0050] Based on the target semantic information, determine whether the target is a dynamic target. If it is a dynamic target, then track the target's trajectory.
[0051] Furthermore, after determining whether the target is a dynamic target based on the target semantic information, and if it is a dynamic target, tracking the target's trajectory, the method further includes:
[0052] Based on the preset disappearance coefficient of the dynamic target, the occupancy probability of each grid on the track in the current frame of the dynamic target, and the occupancy probability of each grid on the track in the previous frame of the dynamic target after disappearance, the occupancy probability of each grid on the track in the current frame after disappearance is obtained.
[0053] Further, obtaining a grid map with semantics and occupancy probability based on the second fusion probability distribution and the target semantic information acquired by the visual sensor includes:
[0054] Based on the second fusion probability distribution, the boundary contour information of non-ground element targets is fitted;
[0055] Extract the location and type of ground-element targets and non-ground-element targets from the target semantic information obtained from visual sensors;
[0056] Construct a raster map with semantics and occupancy probability, wherein the raster map with semantics and occupancy probability includes boundary contour information of non-ground element targets, and the location and type of ground element targets and non-ground element targets.
[0057] Furthermore, the ground element targets include parking spaces, speed bumps, and wheel chocks, while the non-ground element targets include vehicles, people, and pillars.
[0058] Secondly, embodiments of this application provide a multi-source sensor data fusion mapping software architecture, which includes:
[0059] The preprocessing module is used to acquire sensor data from active non-contact ranging sensors and vision sensors, obtain the probability distribution of measurement points in the current frame for each type of sensor, and obtain the first probability distribution of measurement points in the current frame for the vision sensor based on the probability distribution of measurement points in the current frame and the confidence distribution corresponding to each measurement point.
[0060] The homogeneous fusion module is used to fuse the probability distribution of the measurement points in the current frame of the active non-contact ranging sensor with the probability distribution of the measurement points in a set number of previous data frames to obtain a first fused probability distribution of the active non-contact ranging sensor; and to fuse the first probability distribution of the measurement points in the current frame of the vision sensor with the first probability distribution of the measurement points in a set number of previous data frames to obtain a first fused probability distribution of the vision sensor.
[0061] The heterogeneous fusion module is used to update the first fusion probability distribution of each type of sensor to the fusion data layer and perform probability fusion to obtain a second fusion probability distribution; and based on the second fusion probability distribution and the target semantic information obtained by the visual sensor, to obtain a grid map with semantics and occupancy probability.
[0062] Furthermore, the preprocessing module also includes a visual preprocessing module, which is used to generate a confidence distribution corresponding to the measurement point according to a first rule, the first rule including: if the measurement point is a ground element, the greater the distance from the measurement point to the visual sensor, the lower the confidence.
[0063] The visual preprocessing module is also used to generate a confidence distribution corresponding to the measurement points according to a second rule, the second rule including: separating the foreground of non-ground element targets, the higher the position of the foreground, the lower the confidence of the measurement points in the drivable ground area.
[0064] Furthermore, it also includes an input module, a data management module, a map pose module, and an output module;
[0065] The input module is used to send sensor data from the active non-contact ranging sensor and the vision sensor to the preprocessing module and the map pose module.
[0066] The data management module is used to receive and store the first fusion probability distribution of the active non-contact ranging sensor and the first fusion probability distribution of the visual sensor from the same-source fusion module, and send them to the heterogeneous fusion module and the map pose module.
[0067] The data management module is also used to receive a raster map with semantics and occupancy probability from the heterogeneous fusion module and send it to the output module.
[0068] Furthermore, the map pose module is used to correct the probability distribution of the active non-contact ranging sensor measurement points and the first probability distribution of the visual sensor measurement points at the first moment to the current moment, based on the first moment when the sensor data of the active non-contact ranging sensor and the visual sensor are acquired, and the vehicle motion data between the current moment and the first moment.
[0069] Furthermore, the map pose module is also used for:
[0070] Obtain the first position of the preset type target in the world coordinate system in the current frame; obtain the second position of the preset type target in the world coordinate system in a previous frame;
[0071] The vehicle's attitude in the grid map is corrected based on the difference between the first position and the second position, wherein the position of the preset type target in the world coordinate system is obtained based on its position in the vehicle coordinate system and the position of the vehicle in the world coordinate system.
[0072] Thirdly, embodiments of this application provide a vehicle, including: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the vehicle to perform the method as described in any of the preceding claims.
[0073] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the preceding claims.
[0074] The technical solution provided in this application can bring at least the following beneficial effects:
[0075] This application utilizes two main types of sensors—active non-contact ranging sensors and visual sensors—as data sources. These two sensors can complement each other. Furthermore, by applying probability distribution processing to both the active non-contact ranging sensors and the visual sensors, the accuracy of the data is improved. In particular, this application also performs confidence processing on the visual data, avoiding problems such as inaccurate recognition caused by occlusion. Finally, data fusion is performed between data from the same type of sensor and data from different types of sensors, further improving the accuracy and robustness of recognition. Attached Figure Description
[0076] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0077] Figure 1 Flowchart of the multi-source sensor data fusion mapping method provided in this application;
[0078] Figure 2 This application provides a software architecture module diagram for multi-source sensor data fusion mapping.
[0079] Figure 3 The software architecture diagram for multi-source sensor data fusion mapping provided in this application. Detailed Implementation
[0080] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0081] See Figure 1 As shown in the figure, this application provides a multi-source sensor data fusion mapping method, which includes the following steps:
[0082] 100: Acquire sensor data from active non-contact ranging sensors and vision sensors to obtain the probability distribution of measurement points in the current frame for each type of sensor.
[0083] In this application embodiment, the sensor can be divided into two types: active non-contact ranging sensor and visual sensor.
[0084] Furthermore, in order to achieve better complementarity, some embodiments may include two active non-contact ranging sensors with different detection ranges. Specifically, these may include a first ranging sensor and a second ranging sensor, with the first ranging sensor having a shorter detection range than the second ranging sensor.
[0085] In some specific implementations, the first ranging sensor can be an ultrasonic radar, and the second ranging sensor can be a millimeter-wave radar.
[0086] Based on their different detection distances, millimeter-wave sensors are suitable for long-distance, high-precision detection in complex environments (such as advanced driver assistance systems and industrial monitoring), utilizing electromagnetic wave characteristics to achieve multi-parameter measurement. Ultrasonic sensors are suitable for short-distance, low-cost obstacle detection in non-vacuum environments (such as household appliances and simple industrial scenarios), relying on mechanical wave reflection and having low requirements for environmental details.
[0087] In the field of advanced driver assistance systems (ADAS), both ultrasonic radar and millimeter-wave radar play important roles. For example, ultrasonic radar has a detection range of only about 5 meters, while millimeter-wave radar can detect ranges up to 200 meters. Furthermore, ultrasonic radar can only measure distance, not speed, while millimeter-wave radar can measure speed. Ultrasonic radar has high accuracy at short ranges, while millimeter-wave radar has high accuracy at long ranges.
[0088] Furthermore, in some embodiments, the vision sensor can include commonly used automotive cameras, such as monocular cameras, fisheye cameras, etc. In this embodiment, to meet parking requirements, multiple fisheye cameras can be set, for example, using 4-6 fisheye lenses to achieve 360° coverage around the vehicle body without blind spots.
[0089] It should be understood that this application is not limited to any specific vision sensor; however, in this embodiment, a fisheye camera was selected to better suit the parking scenario.
[0090] In step 100, the sensor data of different types of sensors are processed to obtain the probability distribution of the measurement points in the current frame for each sensor. The current frame is represented in the form of an occupied grid.
[0091] Processing sensor data into a probability distribution can improve the robustness of perception and decision-making by quantifying uncertainty. Furthermore, the echoes of false detection points have a low probability in the probability distribution and are easily eliminated by filtering algorithms.
[0092] 200: Based on the probability distribution of the measurement points in the current frame of the visual sensor and the confidence distribution corresponding to each measurement point, the first probability distribution of the measurement points in the current frame of the visual sensor is obtained.
[0093] In the data measured by the visual sensor, since it can reflect the occlusion between targets in the image, the probability distribution of the current frame measurement point alone cannot completely and realistically reflect whether the measurement point is occupied or empty. Therefore, the confidence distribution corresponding to each measurement point is introduced, and the probability distribution is adjusted by the confidence distribution to ensure that the obtained first probability distribution can more realistically reflect the occupancy of the measurement point in the current frame of the visual sensor.
[0094] 300: The probability distribution of the measurement points in the current frame of the active non-contact ranging sensor is fused with the probability distribution of the measurement points in the previous data frames of a set number of frames to obtain the first fused probability distribution of the active non-contact ranging sensor.
[0095] In step 300, for example, if the active non-contact ranging sensor mainly uses millimeter-wave radar or ultrasonic radar, the probability distribution of the measurement points in the current frame of the millimeter-wave radar is fused with the probability distribution of the measurement points in the previous data frames of the millimeter-wave radar for a set number of frames to obtain the first fused probability distribution of the millimeter-wave radar.
[0096] The probability distribution of the measurement points in the current frame of the ultrasonic radar is fused with the probability distribution of the measurement points in the previous data frames of the ultrasonic radar for a set number of frames to obtain the first fused probability distribution of the ultrasonic radar.
[0097] Both step 300 and the next step 400 perform source data fusion on the current frame and previous historical data frames of the same type of sensor, and use a Bayesian fusion equation for source data fusion, wherein the Bayesian fusion equation is as follows:
[0098]
[0099] in:
[0100] : Indicates from frame 1 to frame 2 The observation of frames includes historical observations before fusion. For operational efficiency, the number of frames is usually set to 3, but other values such as 2, 4, or 5 can be used depending on actual needs.
[0101] The probability after fusing the observations of the current frame is the first fusion probability distribution of the active non-contact ranging sensor or the first fusion probability distribution of the visual sensor.
[0102] : The likelihood of observation when the grid is occupied in the current frame, i.e., the probability distribution of the measurement points of the current frame of the active non-contact ranging sensor or the first probability distribution of the measurement points of the current frame of the visual sensor.
[0103] Before fusion The prior probability after the frame is the probability distribution of the measurement points of the preceding data frame with a given number of frames.
[0104] : Normalization factor, set to 1.
[0105] 400: The first probability distribution of the measurement points of the current frame of the visual sensor is fused with the first probability distribution of the measurement points of the preceding data frames of a set number of frames to obtain the first fused probability distribution of the visual sensor.
[0106] In step 400, for example, if the visual sensor mainly uses a fisheye camera, the first probability distribution of the measurement points of the current frame of the fisheye camera is fused with the first probability distribution of the measurement points of the previous data frames of the fisheye camera for a set number of frames to obtain the first fused probability distribution of the fisheye camera.
[0107] 500: Update the first fusion probability distribution of each type of sensor to the fusion data layer and perform probability fusion to obtain the second fusion probability distribution.
[0108] In step 500, the DS evidence synthesis method is used to update the first fusion probability of the current frame of different types of sensors to the fusion data layer, and then perform probability fusion to achieve the second fusion probability distribution.
[0109] The DS evidence synthesis method is a theory for handling uncertain information. Compared with the Bayesian method, its advantage lies in its ability to effectively express the "unknown" (such as the uncertainty when the sensor cannot determine the target category).
[0110] In this application, the DS evidence synthesis method transforms the "subjective judgments" of various types of sensors into "collective decisions" (the fused probability distribution) through mathematical rules, while retaining the ability to express the "unknown." This process is similar to "expert voting"—different sensors, acting as "experts," provide judgments with uncertainty. The DS rules, by handling conflicts and normalization, ultimately output more reliable conclusions, providing a more robust environmental perception foundation for map generation in fields such as parking.
[0111] 600: Based on the second fusion probability distribution and the target semantic information obtained by the visual sensor, a grid map with semantics and occupancy probability is obtained.
[0112] This application uses measurement data from ranging sensors and vision sensors, and calculates the probability distribution of measurement data from each type of sensor. When fusing the probability distributions of sensors of the same type, a Bayesian fusion method is used for multi-frame temporal fusion. When fusing the probability distributions of sensors of different types, a DS evidence synthesis method is used for dynamic mapping and uncertainty processing, resulting in stronger robustness.
[0113] The mapping method provided in this application can use millimeter-wave radar, ultrasonic radar, and cameras, without the need for lidar. On the one hand, compared to the high-precision combination of lidar and cameras, which requires a high computing power platform, the sensor combination in this application can run on a low computing power platform, making the method more universally applicable in the market.
[0114] Secondly, compared to using LiDAR as the main sensor, which requires multiple LiDARs to fill blind spots and has a high cost, this application uses millimeter-wave radar, ultrasonic radar and cameras, which can greatly reduce costs. By using a combination of 4 fisheye cameras, 4 millimeter-wave radars and 12 ultrasonic radars, a balance can be achieved between cost control and 360° coverage, which is convenient for mass production.
[0115] Furthermore, it should be noted that, without considering cost issues, this application can also utilize lidar.
[0116] Further, in step 100 above, acquiring sensor data from the active non-contact ranging sensor and the vision sensor to obtain the probability distribution of the measurement points in the current frame for each type of sensor specifically includes:
[0117] 101: Acquire sensor data from active non-contact ranging sensors and vision sensors.
[0118] 102: Based on the sensor data of the active non-contact ranging sensor and the vision sensor, obtain the measurement point data of the active non-contact ranging sensor and the vision sensor.
[0119] In step 102, the original sensor data includes not only measurement point data related to the target, but also some noise data. After signal processing, calibration, decoding, and other processing, as well as noise processing to remove outliers (such as false echoes from radar), valid measurement points are retained to obtain measurement point data.
[0120] It should be understood that the measurement points in this application will vary depending on the different sensors. For example, millimeter-wave radar obtains a sparse point cloud, which contains three-dimensional data points including distance, azimuth, and relative velocity. Ultrasonic waves, on the other hand, produce a simple distance signal, while visual sensors generate a pseudo-3D point cloud after processing the data frame. Therefore, in advanced driver assistance systems (ADAS) applications, point signals are ultimately obtained. Furthermore, after processing, the visual signal not only forms a 3D point cloud but also a semantic understanding layer, which classifies the captured target, such as identifying whether it is a person, a vehicle, a lane line, a parking space line, a pillar, etc.
[0121] 103: Based on the measurement point data of the active non-contact ranging sensor and the vision sensor and the corresponding uncertainty model, the probability distribution of the measurement points of the active non-contact ranging sensor and the vision sensor in the current frame is obtained.
[0122] Step 103 combines the measurement point data from the active non-contact ranging sensor and the vision sensor with the corresponding uncertainty model to calculate the probability distribution of the measurement points in the current frame. Subsequent data fusion from different types of sensors can effectively compensate for the limitations of a single sensor. This is because active non-contact ranging sensors can provide accurate distance information but lack semantic recognition capabilities; vision sensors can identify target categories but are greatly affected by lighting and occlusion, and have low distance measurement accuracy. Fusing the data from both in subsequent steps allows for cross-validation using multi-source data, reducing measurement errors and improving the accuracy of target localization. Furthermore, the output of the probability distribution quantifies the uncertainty of environmental perception, making the system more robust to complex scenarios (such as occlusion and severe weather), providing more reliable and comprehensive environmental information for path planning and decision-making in applications such as advanced driver assistance systems and parking navigation, significantly enhancing the system's reliability and safety.
[0123] In summary, step 103, by independently constructing uncertainty models and generating probability distributions for the ranging and vision sensors, essentially transforms the hardware measurements of the sensors into probabilistic judgments at the cognitive level. This process not only preserves the unique advantages of each sensor (such as the geometric accuracy of ranging and the semantic understanding of vision), but also improves the accuracy of target localization by quantifying uncertainty.
[0124] In steps 101-103, for the input sensor data that is raw sensor data in the form of distance points, such as millimeter-wave radar, ultrasonic radar, and visual image feature points, an uncertainty Gaussian filter model (i.e. uncertainty model) based on the corresponding sensor position established based on historical sensor data is used to generate the probability distribution of the measurement points of the current frame of the sensor.
[0125] The probability is expressed as follows:
[0126]
[0127] Where p(z): probability distribution;
[0128] z: Observation value, i.e., the position of the grid near the measurement input by the sensor in the current frame.
[0129] μ: Observation expectation, which is the measurement value of the sensor input in the current frame (occupancy probability or empty probability).
[0130] Σ: Covariance matrix, which is the preset sensor measurement error covariance matrix of the sensor.
[0131] d: Dimension of the observation, i.e., the vector length of the observation.
[0132] Superscript T: indicates the transpose of the matrix.
[0133] -1: Indicates the inverse of the matrix.
[0134] |Σ|: Represents the determinant of matrix Σ.
[0135] Before obtaining the probability distribution of measurement points in the current frame of the active non-contact ranging sensor and the vision sensor, the method further includes step 104, wherein step 104 includes:
[0136] 1040: Obtain measurement point data of the image captured by the vision sensor based on the sensor data of the vision sensor;
[0137] 1041: Determine the type of target within the field of view of the visual sensor and the target's position relative to the visual sensor based on the measurement point data of the image;
[0138] 1042: Based on the position data of the target relative to the visual sensor and the position and angle of the visual sensor on the vehicle, perform intrinsic and extrinsic parameter conversion to obtain the position of the target relative to the vehicle body;
[0139] 1043: Based on the type of target within the field of view of the vision sensor, the position of the target relative to the vehicle body, and the corresponding uncertainty model, the probability distribution of the target in the vehicle coordinate system is obtained.
[0140] First, it should be noted that after establishing the vehicle's coordinate system, the probability distribution of the target in the world coordinate system can be obtained based on the vehicle's position in the world coordinate system.
[0141] Based on the above, coordinate system transformation can be obtained through the following calculation method:
[0142] The transformation equation is as follows:
[0143]
[0144] Among them, X ω Let X be the coordinates of the measured point in the world coordinate system. c K represents the coordinates of the measurement point in the camera coordinate system; K is the camera intrinsic parameter matrix (focal length, principal point); R is the rotation matrix from the world coordinate system to the camera coordinate system; t 平移 This is the translation vector from the world coordinate system to the camera coordinate system.
[0145] To use the above transformation equation to transform from the camera coordinate system to the vehicle coordinate system, simply change the X... c Divide by K.
[0146] In some embodiments, step 200, obtaining the first probability distribution of the measurement points in the current frame of the visual sensor based on the probability distribution of the measurement points in the current frame of the visual sensor and the confidence distribution corresponding to each measurement point, may further include the following steps:
[0147] 201: Obtain the confidence distribution, that is, obtain the confidence of the area around the vehicle captured by the vision sensor, which specifically includes the following steps:
[0148] 2011: Generate the confidence distribution corresponding to the measurement point according to the first rule, the first rule including: if the measurement point is a ground element, the greater the distance from the measurement point to the visual sensor, the lower the confidence.
[0149] It should be understood that setting a lower confidence level for points farther from the vision sensor results in a lower probability distribution for distant observation points. Specifically, the probability of classifying a point as occupied or empty is lower, while the probability of classifying it as unknown is higher. This better matches the vision sensor's recognition performance at a distance, as the sensor may misidentify points due to blurriness or diffuse reflection from dust. Therefore, the greater the distance from the measurement point to the vision sensor, the lower the confidence level. In some examples, the confidence level is set to 0.9 for observation points further away from the sensor, and 0.85 for even further away. It's important to note that we are only slightly reducing the confidence level for more distant points; the recognition results still have relatively good reference value.
[0150] In some optional implementations, the confidence distribution can also be obtained through the following step 2012, which specifically includes:
[0151] 2012: Generate a confidence distribution corresponding to the measurement points according to the second rule, which includes: separating the foreground of non-ground element targets, and the lower the confidence of the measurement points in the ground drivable area that are obscured by the higher position of the foreground.
[0152] First, visual sensors cannot identify information behind occluded objects. This application avoids misidentification of these areas by configuring the confidence level of the occluded parts to be lower.
[0153] Specifically, the foreground of non-ground element targets is first separated, including:
[0154] The probability distribution of the separation point between the target and the ground is obtained based on the preset obstacle separation model and the types of non-ground element targets;
[0155] It should be understood that the obstacle separation model is a pre-defined model, and different target types have different characteristics related to occlusion separation, such as their contours. Therefore, the target type needs to be involved in the foreground separation process. Specifically, foreground separation includes the formation of the dividing line between the foreground and the ground, as well as the formation of edge contours. For visual sensors, edge contours are relatively easy to identify, while the dividing points between the ground and the foreground are relatively more difficult to identify accurately. Applying a probability distribution to the dividing points can effectively improve accuracy.
[0156] Then, based on the probability distribution of the segmentation points, the portion of the non-ground element target that is close to the ground is separated; that is, the foreground of the non-ground element target is separated from the ground using a separation model.
[0157] Obstacle separation models can be pre-established using parameters such as the camera observation angle, distance to non-ground targets, and types of non-ground element targets. Since non-ground element targets differ in size, shape, and other semantics, the probability distribution of segmentation points for different types of non-ground element targets can be obtained based on the types of non-ground element targets and the obstacle separation model.
[0158] After separating the foreground, it is necessary to identify which observed points are most likely to be occluded by the foreground. In the previous step of forming the segmentation points, this application used a probability distribution method to increase the possibility of deviation. However, in some scenarios (such as when the foreground and ground colors are basically the same), there may still be some errors in the identification. Therefore, in order to further improve the accuracy, this application sets the confidence level of the measurement points occluded by the higher the segmentation point is, based on the fact that the higher the segmentation point is, the higher the probability that it is actually the foreground (occluder).
[0159] For the reasons mentioned above, this application obtains the height of the visual sensor observation position based on the angle between the observation position of the visual sensor and the segmentation point.
[0160] The confidence level of measurement points in the drivable ground area that is obscured by the higher position of the foreground is configured to be lower.
[0161] It should also be noted that during separation, since the non-ground target elements are too high up, the parts that are obscured by the vehicle's view are often in the air and do not affect driving. Therefore, when performing foreground separation, it is not necessary to form a complete target foreground to reduce computing power requirements.
[0162] After step 201, step 200 is performed. In some embodiments, step 200 further includes the following steps:
[0163] 202: Based on the probability of the measurement point in the current frame of the visual sensor, the confidence level corresponding to the measurement point, and the third rule, the first probability distribution of the measurement point in the current frame of the visual sensor is obtained;
[0164] The third rule includes: the lower the confidence level of the region, the lower the first probability that the measurement point corresponding to the current frame is occupied or empty, and the first probability is lower than the probability that the measurement point in the current frame is occupied or empty.
[0165] The third rule can be to multiply the first probability of each measurement point being occupied or empty by the corresponding confidence level.
[0166] Based on this, in some instances, the following algorithm can be used to determine whether the measurement points in the partially obscured drivable area are occupied or empty.
[0167]
[0168] in, The confidence level of the foreground portion near the ground that is visible in the current frame. The higher the height, the higher the confidence level. This can be calculated based on the angle from the segmentation point to the camera. This represents the probability distribution of drivable areas in the current frame. This means that the confidence level of measurement points in the drivable ground area that will be obscured by the higher position of the foreground is configured to be lower. This represents the first probability distribution that includes drivable areas in the current frame.
[0169] In some embodiments, after step 100, step 203 may be included, which includes the following steps:
[0170] Determine whether the type of target observed by the visual sensor belongs to the ground element target subclass or the non-ground element target subclass;
[0171] If the target subclass is a ground element target, then generate the probability distribution of the ground element target in the vehicle coordinate system.
[0172] As mentioned earlier, the confidence level of ground-based targets does not decrease significantly with increasing distance. Therefore, the impact of ground-based target confidence can be disregarded, further reducing the computational power requirement. Of course, whether to disregard ground-based target confidence can be chosen based on circumstances, such as weather conditions.
[0173] By weighted fusion of the probability distribution and confidence level of measurement points in the current frame of a visual sensor, a refined expression of the reliability of measurement results is achieved. Specifically, the original probability distribution of each measurement point reflects the likelihood that it belongs to a certain target or background, while the confidence level characterizes the credibility of the measurement result. By multiplying the two, the basic probability information of the measurement points is preserved, and the weights are adjusted according to the confidence level: the probability of measurement points with high confidence is strengthened, while the probability of low confidence is weakened. This processing method effectively solves the problem of unreliable measurements by visual sensors in complex environments (such as changes in lighting and occlusion), enabling the system to more accurately distinguish between real targets and noise or false detection results. The resulting first probability distribution not only more realistically reflects the environmental state but also provides more reliable data support for subsequent target recognition, localization, and decision-making, significantly improving the robustness and reliability of the visual perception system.
[0174] After step 203 is executed, step 300 is executed, which involves fusing the probability distribution of the measurement points in the current frame of the active non-contact ranging sensor with the probability distribution of the measurement points in the previous data frames of a set number of frames to obtain the first fused probability distribution of the active non-contact ranging sensor.
[0175] In some possible implementations, step 500 updates the first fusion probability distribution for each type of sensor to the fusion data layer and performs probability fusion to obtain a second fusion probability distribution, including:
[0176] 501: Based on the first fusion probability distribution of each type of sensor, the effective detection range of each type of sensor, the confidence of each type of sensor in the space being occupied, and the conflict factor of the sensor observation results, the occupancy result of each measurement point in the fusion data layer is obtained, and the occupancy result includes occupied, empty, and unknown.
[0177] 502: Based on the occupancy results of each measurement point, a second fusion probability distribution is formed.
[0178] The above steps mainly utilize the DS evidence synthesis method to obtain the final occupation result, which can be expressed by the following formula:
[0179]
[0180] in:
[0181] m1(occ): Occupancy probability of sensor 1;
[0182] m2(occ): Occupancy probability of sensor 2;
[0183] m1(free): The probability of sensor 1 being empty;
[0184] m2(free): The probability of sensor 2 being empty;
[0185] K 1,occ : Confidence coefficient of sensor 1 for occupation;
[0186] K 2,occ : Confidence coefficient of sensor 2 for occupation;
[0187] K 1,free : Confidence coefficient of sensor 1 for the air;
[0188] K 2,free : Confidence coefficient of sensor 2 for the air;
[0189] m 1,2 (occ): The occupancy probability output after fusion;
[0190] m 1,2 (free): The empty probability output after fusion.
[0191] Understandably, different types of sensors have different effective detection ranges, and therefore different levels of confidence in the occupancy of space. Because of these differing levels of confidence, conflicting observations may occur between different types of sensors. By combining the detection results from different types of sensors, the occupancy results for each measurement point can be obtained.
[0192] The above-mentioned scheme comprehensively considers the probability distribution characteristics, detection range, confidence level, and conflict factor of sensors during multi-sensor fusion, achieving accurate characterization and efficient integration of environmental information. Specifically, the first fusion probability distribution reflects the measurement confidence of sensors of the same type after preliminary processing; the effective detection range defines the reliable operating area of sensor data, avoiding interference from invalid data; the confidence level quantifies the historical performance and data reliability of each type of sensor, providing a weighting basis for fusion; and the conflict factor is used to measure the degree of contradiction between the observation results of different types of sensors. Based on these parameters, the measurement points of the fusion data layer are classified to determine whether they belong to the "occupied," "empty," or "unknown" state, effectively filtering redundant and contradictory data; and then the discrete occupancy results are transformed into a continuous second fusion probability distribution, forming a unified parking environment probability model.
[0193] This fusion strategy significantly enhances the environmental perception capabilities of multi-sensor systems: First, by constraining multi-dimensional parameters, it reduces the limitations of single sensors, such as the influence of lighting on visual sensors and the measurement blind spots of ranging sensors; Second, the handling of conflicting data avoids misjudgments in the fusion results, enabling the system to stably output accurate probability distributions even in complex scenarios (such as occlusion); Third, the final second fusion probability distribution provides more robust environmental information for subsequent decision-making.
[0194] In addition, the applicant discovered that several sensor timing mismatch issues exist in the relevant technologies. For example, millimeter-wave radar, ultrasonic radar, and visual signals have different feedback periods. Therefore, it is difficult to ensure that the multiple signals being processed are at the same time, and a unified time scale is required for subsequent data fusion.
[0195] Therefore, in some embodiments, after step 100, the method further includes step 700, which includes:
[0196] Based on the first moment of acquiring sensor data from the active non-contact ranging sensor and the vision sensor, and the vehicle motion data between the current moment and the first moment, the probability distribution of the active non-contact ranging sensor measurement points and the first probability distribution of the vision sensor measurement points at the first moment are corrected to the current moment.
[0197] It should be understood that the initial moments of the sensor data from the active non-contact ranging sensor and the vision sensor are different; each sensor has its own timestamp. However, for overall consistency, they all need to be unified to the current time. The applicant corrected the probability distribution of each sensor based on the differences in their respective timestamps and the motion data under these differences. The corrected data then has a unified timestamp.
[0198] Specifically, step 700 can be achieved through the following calculation method:
[0199] The synchronization equation is as follows:
[0200]
[0201] Among them, X ti Let T be the coordinates of the sensor at the first moment ti; ti t Let X be the pose transformation matrix from time ti to the current time t; t To ensure uniform compensation to the coordinates at the current time t.
[0202] Furthermore, relying solely on integral calculations from sensors such as IMUs and wheel speedometers (e.g., calculating the number of wheel rotations or integrating acceleration) can lead to an accumulation of vehicle pose errors over time (e.g., a positional deviation of 0.5 meters after traveling 100 meters). Measurements from vision / radar sensors at different times contain noise, which, if not corrected, can cause inconsistencies between mapping and localization (e.g., the same obstacle "drifting" in position across different frames).
[0203] Therefore, changes may occur between the vehicle's position and targets such as lane markings and parking spaces, causing map drift issues that require correction.
[0204] In some embodiments, after correcting the probability distribution of the active non-contact ranging sensor measurement points at the first moment and the first probability distribution of the visual sensor measurement points to the current moment, the method further includes step 800, which specifically may consist of steps 801-803:
[0205] 801: Get the first position of the preset type of target in the world coordinate system in the current frame.
[0206] Among them, the preset target types mainly include marking lines, parking space lines, etc.
[0207] 802: Obtain the second position of the preset type target in the world coordinate system in the previous frame.
[0208] 803: Correct the vehicle's attitude in the grid map based on the difference between the first position and the second position, wherein the position of the preset type target in the world coordinate system is obtained based on its position in the vehicle coordinate system and the position of the vehicle in the world coordinate system.
[0209] It is understandable that step 800 can be executed after step 400 or after step 700.
[0210] In addition, it should be emphasized that difference correction can move the entire map, that is, the probability distribution of all sensors, and can also move the vehicle's position. In order to reduce the computing power requirements, this application prioritizes the method of moving the vehicle.
[0211] By inferring pose error from the difference in target position across frames, the cumulative error of the sensor is cleared in real time, which can effectively reduce map drift and improve stability.
[0212] The single-point matching error is defined as follows:
[0213]
[0214] e i (T): The residual between the feature points of the preset type of target in the visual semantic layer of the current frame and the feature points of the preset type of target in the visual semantic layer of the previous frame under the pose transformation matrix T from the previous frame to the current frame, i.e., the distance difference. The preset type of target mainly includes road markings, parking lines, etc. The feature points refer to the feature points of road markings / parking lines, which are represented by grids in a grid map.
[0215] P cur,i : The position of the i-th grid cell of the preset type target in the visual semantic layer of the current frame, that is, the first position of the preset type target in the world coordinate system in the current frame.
[0216] P map,i:The position of the i-th grid of the preset type target in the visual semantic layer of the previous frame is the second position of the preset type target in the world coordinate system in the previous frame.
[0217] T: Pose transformation matrix from the previous frame to the current frame.
[0218] Based on the grids of predefined target categories in the visual semantic layers of the current frame and previous frames, the grids corresponding to the current frame and previous frames are combined into a set of matching feature point pairs, and then the single-point matching error e is calculated. i (T) is the minimum total error value under this set. At this point, the total error reaches its minimum value. The corresponding pose transformation matrix T is denoted as , This is the pose transformation matrix from the previous frame to the current frame that needs to be determined. This pose transformation matrix is also the data basis for the vehicle's pose transformation.
[0219]
[0220] N: Number of matching feature point pairs.
[0221] : This indicates finding the appropriate value among different input values of the pose transformation matrix T. To reduce the single-point matching error e i (T) The fitting method used when the total error is minimized under this set.
[0222] The applicant discovered that some sensor data, such as millimeter-wave radar, leaves residues on the tracks of moving objects. In reality, the historical position of the moving object on the track is no longer occupied, which affects the real-time performance of the data.
[0223] Specifically, this application also includes step 900. Step 900 is performed before obtaining a grid map with semantics and occupancy probability based on the second fusion probability distribution and the target semantic information acquired by the visual sensor. Step 900 includes determining whether the target is a dynamic target based on the target semantic information. If it is a dynamic target, the trajectory of the target is tracked.
[0224] Specifically, potential dynamic targets mainly include people, cars, and two-wheeled vehicles (such as electric bicycles and bicycles). Kalman filtering is applied to potential dynamic obstacles, such as people, cars, and two-wheeled vehicles, to track their motion states. The state equation for the tracked target is as follows:
[0225]
[0226] in:
[0227] x k : The motion state (including position and velocity) of the dynamic target in the k-th frame;
[0228] F: State transition matrix;
[0229] : Indicates state estimation noise;
[0230] z k : Current frame sensor observations;
[0231] H: Sensor-corresponding observation model, also known as the measurement matrix;
[0232] The current frame sensor observation noise, also known as measurement noise.
[0233] Furthermore, for dynamic targets, such as people, animals, and vehicles, their positions on the map may change over time, or even disappear from the map. In order to eliminate the problem of reduced map accuracy caused by the positional changes of dynamic targets due to their movement, this application sets a disappearance coefficient for different dynamic targets. Its core function is to help the system efficiently manage target tracking in complex environments (such as changes in lighting and occlusion in parking scenarios) by quantifying the "survival status" or "existence probability" of dynamic targets, thereby reducing false detections and redundant calculations.
[0234] Specifically, in order to remove the residue of dynamic targets, this application further includes step 1000, which includes: obtaining the occupancy probability of the dynamic target in the current frame after the elimination of each grid on the track based on the preset elimination coefficient of the dynamic target, the occupancy probability of each grid on the track of the current frame of the dynamic target in the current frame, and the occupancy probability of each grid on the track of the dynamic target in the previous frame after elimination.
[0235] Specifically, the following algorithm can be used to calculate the dynamic elimination of targets:
[0236]
[0237] p occ (t): The probability that a grid cell is occupied after the cell has been eliminated at time t;
[0238] p obs (t): The probability of this grid cell being occupied at the current time t;
[0239] p occ (t-1): The probability that the previous frame will be occupied after the extinction calculation;
[0240] : The extinction coefficient generated by the dynamic target type.
[0241] It is understandable that the process of eliminating dynamic targets, i.e., step 900 above, can be executed after step 500 (i.e., obtaining the second fusion probability distribution), or after step 400 (i.e., obtaining the first fusion probability distribution) and before step 500 (i.e., obtaining the second fusion probability distribution).
[0242] Further, in step 600, based on the second fusion probability distribution and the target semantic information acquired by the visual sensor, a grid map with semantics and occupancy probability is obtained, including:
[0243] 601: Based on the second fusion probability distribution, fit the boundary contour information of non-ground element targets;
[0244] 602: Extract the location and type of ground element targets and non-ground element targets from the target semantic information obtained from the visual sensor;
[0245] 603: Construct a raster map with semantics and occupancy probability, wherein the raster map with semantics and occupancy probability includes boundary contour information of non-ground element targets, and the location and type of ground element targets and non-ground element targets.
[0246] The ground-based target elements include parking spaces, speed bumps, and wheel chocks, while the non-ground-based target elements include vehicles, people, and pillars.
[0247] In this embodiment, information is extracted from the second fusion probability distribution to construct ground elements such as parking spaces, speed bumps, and wheel chocks. The boundaries and contours of various types of obstacles, such as cars and pillars, are fitted to form accurate perceptual information in the form of a grid map with semantics and occupancy probability.
[0248] After Gaussian filtering, the occupied and empty probabilities in the second fusion probability distribution are extracted to remove the regions whose distributions are greater than a certain value, such as 0.8, thus smoothing the occupied and empty outputs in the second fusion probability distribution. The Gaussian filtering algorithm is as follows:
[0249]
[0250] It should be understood that the second fusion probability distribution of this application is projected onto a grid map, where x and y are the coordinate offsets of neighboring grids relative to this grid, and σ is the standard deviation.
[0251] Based on this, ground elements and non-ground elements in the second fusion probability distribution can be clustered, for example, by using 8-neighborhood connectivity clustering, and then using the RANSAC random sampling consensus algorithm for boundary fitting to obtain the final map output.
[0252] Overall, this application significantly improves the accuracy of the final grid map by processing the probability distribution of different sensor signals, the confidence level of visual signals, and the fusion of the probability distributions of the same source and the overall probability distribution. Based on these improvements, the computational power requirement can be reduced by adjusting the density of the probability distribution grid while maintaining the same precision or accuracy.
[0253] Please refer to Figure 2 This application also provides a software architecture module diagram for multi-source sensor data fusion mapping, which includes:
[0254] The preprocessing module is used to acquire sensor data from active non-contact ranging sensors and vision sensors, obtain the probability distribution of measurement points in the current frame for each type of sensor, and obtain the first probability distribution of measurement points in the current frame for the vision sensor based on the probability distribution of measurement points in the current frame and the confidence distribution corresponding to each measurement point.
[0255] The homogeneous fusion module is used to fuse the probability distribution of the measurement points in the current frame of the active non-contact ranging sensor with the probability distribution of the measurement points in a set number of previous data frames to obtain a first fused probability distribution of the active non-contact ranging sensor; and to fuse the first probability distribution of the measurement points in the current frame of the vision sensor with the first probability distribution of the measurement points in a set number of previous data frames to obtain a first fused probability distribution of the vision sensor.
[0256] The heterogeneous fusion module is used to update the first fusion probability distribution of each type of sensor to the fusion data layer and perform probability fusion to obtain a second fusion probability distribution; and based on the second fusion probability distribution and the target semantic information obtained by the visual sensor, to obtain a grid map with semantics and occupancy probability.
[0257] For further information, please refer to [link / reference]. Figure 3 The software architecture diagram for multi-source sensor data fusion mapping is as follows: Specifically, the preprocessing module also includes a visual preprocessing module, which is used to generate the confidence distribution corresponding to the measurement points according to a first rule. The first rule includes: if the measurement point is a ground element, the greater the distance from the measurement point to the visual sensor, the lower the confidence level.
[0258] The visual preprocessing module is also used to generate a confidence distribution corresponding to the measurement points according to a second rule, the second rule including: separating the foreground of non-ground element targets, the higher the position of the foreground, the lower the confidence of the measurement points in the drivable ground area.
[0259] In some embodiments, this application further includes an input module, a data management module, a map pose module, and an output module;
[0260] The input module is used to send sensor data from the active non-contact ranging sensor and the vision sensor to the preprocessing module and the map pose module.
[0261] The data management module is used to receive and store the first fusion probability distribution of the active non-contact ranging sensor and the first fusion probability distribution of the visual sensor from the same-source fusion module, and send them to the heterogeneous fusion module and the map pose module.
[0262] The data management module is also used to receive a raster map with semantics and occupancy probability from the heterogeneous fusion module and send it to the output module.
[0263] In some embodiments, the map pose module of this application is used to correct the probability distribution of the active non-contact ranging sensor measurement points and the first probability distribution of the visual sensor measurement points at the first moment to the current moment, based on the first moment when the sensor data of the active non-contact ranging sensor and the visual sensor are acquired, and the vehicle motion data between the current moment and the first moment.
[0264] In some implementations, the map pose module is further configured to: obtain the first position of a preset type of target in the world coordinate system in the current frame; and obtain the second position of the preset type of target in the world coordinate system in a previous frame.
[0265] The vehicle's attitude in the grid map is corrected based on the difference between the first position and the second position, wherein the position of the preset type target in the world coordinate system is obtained based on its position in the vehicle coordinate system and the position of the vehicle in the world coordinate system.
[0266] It should be noted that the explanation of the aforementioned method embodiment for multi-source sensor data fusion mapping also applies to the software architecture of this embodiment, and will not be repeated here.
[0267] This application also provides a vehicle including a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to enable the vehicle to perform the multi-source sensor data fusion mapping method of this application.
[0268] Corresponding to the above-described multi-source sensor data fusion mapping method, this application embodiment also provides a storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps of the above embodiments. It should be noted that the storage medium in this application embodiment can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. Computer-readable storage media can be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or apparatus.
[0269] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0270] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0271] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0272] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for mapping based on multi-source sensor data fusion, characterized in that, It includes: Acquire sensor data from active non-contact ranging sensors and vision sensors to obtain the probability distribution of measurement points in the current frame for each type of sensor; Based on the probability distribution of the measurement points in the current frame of the visual sensor and the confidence distribution corresponding to each measurement point, the first probability distribution of the measurement points in the current frame of the visual sensor is obtained; The probability distribution of the measurement points in the current frame of the active non-contact ranging sensor is fused with the probability distribution of the measurement points in the previous data frames of a set number of frames to obtain the first fused probability distribution of the active non-contact ranging sensor. The first probability distribution of the measurement points of the current frame of the visual sensor is fused with the first probability distribution of the measurement points of the previous data frames of a set number of frames to obtain the first fused probability distribution of the visual sensor. The first fusion probability distribution of each type of sensor is updated to the fusion data layer, and probability fusion is performed to obtain the second fusion probability distribution; Based on the second fusion probability distribution and the target semantic information obtained by the visual sensor, a grid map with semantics and occupancy probability is obtained. Before obtaining a raster map with semantics and occupancy probability based on the second fusion probability distribution and the target semantic information acquired by the visual sensor, the method further includes: Based on the target semantic information, determine whether the target is a dynamic target; if it is a dynamic target, then track the target's trajectory. The probability of the dynamic target occupying each grid point on the track in the current frame is obtained based on the preset elimination coefficient of the dynamic target, the occupancy probability of the dynamic target in each grid point on the track in the current frame, and the occupancy probability of the dynamic target in each grid point on the track in the previous frame after elimination. The method further includes obtaining the confidence distribution, which specifically includes: generating a confidence distribution corresponding to the measurement points according to a second rule; the second rule includes: separating the foreground of non-ground element targets, and the measurement points of the ground drivable area that are obscured by the higher position of the foreground have lower confidence. The foreground from which non-ground element targets are separated includes: The probability distribution of the separation point between the target and the ground is obtained based on the preset obstacle separation model and the types of non-ground element targets; Based on the probability distribution of the segmentation points, the non-ground element target near the ground is separated; The height of the visual sensor's observation position is obtained based on the angle between the observation position of the visual sensor and the segmentation point. The confidence level of measurement points in the drivable ground area that is obscured by the higher the foreground position, the lower the confidence level should be; The step of updating the first fusion probability distribution of each type of sensor to the fusion data layer and performing probability fusion to obtain the second fusion probability distribution includes: Based on the first fusion probability distribution of each type of sensor, the effective detection range of each type of sensor, the confidence level of each type of sensor in the space being occupied, and the conflict factor of the sensor observation results, the occupancy results of each measurement point in the fusion data layer are obtained, and the occupancy results include occupied, empty, and unknown. A second fusion probability distribution is formed based on the occupancy results of each measurement point.
2. The multi-source sensor data fusion mapping method as described in claim 1, characterized in that: The active non-contact ranging sensor includes a first ranging sensor and a second ranging sensor, and the detection distance of the first ranging sensor is less than the detection distance of the second ranging sensor.
3. The multi-source sensor data fusion mapping method as described in claim 2, characterized in that: The first ranging sensor is an ultrasonic radar, and the second ranging sensor is a millimeter-wave radar.
4. The multi-source sensor data fusion mapping method as described in claim 1, characterized in that: Acquire sensor data from active non-contact ranging sensors and vision sensors to obtain the probability distribution of measurement points in the current frame for each type of sensor, including: Acquire sensor data from active non-contact ranging sensors and vision sensors; Based on the sensor data from the active non-contact ranging sensor and the vision sensor, the measurement point data of the active non-contact ranging sensor and the vision sensor are obtained. Based on the measurement point data of the active non-contact ranging sensor and the vision sensor, and the corresponding uncertainty model, the probability distribution of the measurement points in the current frame of the active non-contact ranging sensor and the vision sensor is obtained.
5. The multi-source sensor data fusion mapping method as described in claim 1, characterized in that: Before obtaining the probability distribution of measurement points in the current frame from the active non-contact ranging sensor and the vision sensor, the method further includes: The measurement point data of the image captured by the vision sensor is obtained based on the sensor data of the vision sensor; The type of target within the field of view of the visual sensor and the position of the target relative to the visual sensor are determined based on the measurement point data of the image. The position of the target relative to the vehicle body is obtained by converting intrinsic and extrinsic parameters based on the position data of the target relative to the visual sensor and the position and angle of the visual sensor on the vehicle. Based on the type of target within the field of view of the vision sensor, the target's position relative to the vehicle body, and the corresponding uncertainty model, the probability distribution of the target in the vehicle coordinate system is obtained.
6. The multi-source sensor data fusion mapping method as described in claim 1, characterized in that: The method further includes obtaining the confidence distribution, which specifically includes: A confidence distribution corresponding to the measurement points is generated according to a first rule, which includes: If the measurement point is a ground element, the greater the distance from the measurement point to the visual sensor, the lower the confidence level.
7. The multi-source sensor data fusion mapping method as described in claim 1, characterized in that: After acquiring sensor data from active non-contact ranging sensors and vision sensors, and obtaining the probability distribution of measurement points in the current frame for each type of sensor, the method further includes: Determine whether the type of target observed by the visual sensor belongs to the ground element target subclass or the non-ground element target subclass; If the target subclass is a ground element target, then generate the probability distribution of the ground element target in the vehicle coordinate system.
8. The multi-source sensor data fusion mapping method as described in claim 1, characterized in that: The step of obtaining the first probability distribution of the measurement points in the current frame of the visual sensor based on the probability distribution of the measurement points in the current frame of the visual sensor and the confidence distribution corresponding to each measurement point includes: Based on the probability of the measurement point in the current frame of the visual sensor, the confidence level corresponding to the measurement point, and the third rule, the first probability distribution of the measurement point in the current frame of the visual sensor is obtained; The third rule includes: the lower the confidence level of the region, the lower the first probability that the measurement point corresponding to the current frame is occupied or empty, and the first probability is lower than the probability that the measurement point in the current frame is occupied or empty.
9. The multi-source sensor data fusion mapping method as described in claim 1, characterized in that: After obtaining the probability distribution of measurement points in the current frame for each type of sensor, the method further includes: Based on the first moment of acquiring sensor data from the active non-contact ranging sensor and the vision sensor, and the vehicle motion data between the current moment and the first moment, the probability distribution of the active non-contact ranging sensor measurement points and the first probability distribution of the vision sensor measurement points at the first moment are corrected to the current moment.
10. The multi-source sensor data fusion mapping method as described in claim 9, characterized in that: After correcting the probability distribution of the active non-contact ranging sensor measurement points and the first probability distribution of the vision sensor measurement points at the first moment to the current moment, the method further includes: Get the first position of the preset type of target in the world coordinate system in the current frame; Obtain the second position of the target of the preset type in the world coordinate system in the previous frame; The vehicle's attitude in the grid map is corrected based on the difference between the first position and the second position, wherein the position of the preset type target in the world coordinate system is obtained based on its position in the vehicle coordinate system and the position of the vehicle in the world coordinate system.
11. The multi-source sensor data fusion mapping method as described in claim 1, characterized in that: The step of obtaining a grid map with semantics and occupancy probability based on the second fusion probability distribution and the target semantic information acquired by the visual sensor includes: Based on the second fusion probability distribution, the boundary contour information of non-ground element targets is fitted; Extract the location and type of ground-element targets and non-ground-element targets from the target semantic information obtained from visual sensors; Construct a raster map with semantics and occupancy probability, wherein the raster map with semantics and occupancy probability includes boundary contour information of non-ground element targets, and the location and type of ground element targets and non-ground element targets.
12. The multi-source sensor data fusion mapping method as described in claim 11, characterized in that: The ground-based target elements include parking spaces, speed bumps, and wheel chocks, while the non-ground-based target elements include vehicles, people, and pillars.
13. A multi-source sensor data fusion mapping system, characterized in that, It includes: The preprocessing module is used to acquire sensor data from active non-contact ranging sensors and vision sensors, and obtain the probability distribution of measurement points in the current frame for each type of sensor. Based on the probability distribution of the measurement points in the current frame of the visual sensor and the confidence distribution corresponding to each measurement point, the first probability distribution of the measurement points in the current frame of the visual sensor is obtained. The homogeneous fusion module is used to fuse the probability distribution of the measurement points in the current frame of the active non-contact ranging sensor with the probability distribution of the measurement points in a set number of previous data frames to obtain a first fused probability distribution of the active non-contact ranging sensor; and to fuse the first probability distribution of the measurement points in the current frame of the vision sensor with the first probability distribution of the measurement points in a set number of previous data frames to obtain a first fused probability distribution of the vision sensor. The heterogeneous fusion module is used to update the first fusion probability distribution of each type of sensor to the fusion data layer and perform probability fusion to obtain a second fusion probability distribution; and based on the second fusion probability distribution and the target semantic information obtained by the visual sensor, a grid map with semantics and occupancy probability is obtained. Before obtaining the raster map with semantics and occupancy probability based on the second fusion probability distribution and the target semantic information acquired by the visual sensor, the process also includes: Based on the target semantic information, determine whether the target is a dynamic target; if it is a dynamic target, then track the target's trajectory. The probability of the dynamic target occupying each grid point on the track in the current frame is obtained based on the preset elimination coefficient of the dynamic target, the occupancy probability of the dynamic target in each grid point on the track in the current frame, and the occupancy probability of the dynamic target in each grid point on the track in the previous frame after elimination. The preprocessing module further includes a visual preprocessing module, which is used to generate a confidence distribution corresponding to the measurement points according to a second rule; the second rule includes: separating the foreground of non-ground element targets, and the lower the confidence of the measurement points in the drivable ground area that are occluded at a higher position by the foreground. The foreground from which non-ground element targets are separated includes: The probability distribution of the separation point between the target and the ground is obtained based on the preset obstacle separation model and the types of non-ground element targets; Based on the probability distribution of the segmentation points, the non-ground element target near the ground is separated; The height of the visual sensor's observation position is obtained based on the angle between the observation position of the visual sensor and the segmentation point. The confidence level of measurement points in the drivable ground area that is obscured by the higher the foreground position, the lower the confidence level should be; The step of updating the first fusion probability distribution of each type of sensor to the fusion data layer and performing probability fusion to obtain the second fusion probability distribution includes: Based on the first fusion probability distribution of each type of sensor, the effective detection range of each type of sensor, the confidence level of each type of sensor in the space being occupied, and the conflict factor of the sensor observation results, the occupancy results of each measurement point in the fusion data layer are obtained, and the occupancy results include occupied, empty, and unknown. A second fusion probability distribution is formed based on the occupancy results of each measurement point.
14. The multi-source sensor data fusion mapping system as described in claim 13, characterized in that, The visual preprocessing module is used to generate a confidence distribution corresponding to the measurement point according to a first rule, the first rule including: if the measurement point is a ground element, the greater the distance from the measurement point to the visual sensor, the lower the confidence.
15. The multi-source sensor data fusion mapping system as described in claim 13, characterized in that, It also includes an input data management module, a map pose module, and an output module; The data management module is used to receive and store the first fusion probability distribution of the active non-contact ranging sensor and the first fusion probability distribution of the visual sensor from the same-source fusion module, and send them to the heterogeneous fusion module and the map pose module. The data management module is also used to receive a raster map with semantics and occupancy probability from the heterogeneous fusion module and send it to the output module.
16. The multi-source sensor data fusion mapping system as described in claim 15, characterized in that, The map pose module is used to correct the probability distribution of the active non-contact ranging sensor measurement points and the first probability distribution of the visual sensor measurement points at the first moment to the current moment, based on the first moment when the sensor data of the active non-contact ranging sensor and the vision sensor are acquired, and the vehicle motion data between the current moment and the first moment.
17. The multi-source sensor data fusion mapping system as described in claim 15, characterized in that, The map pose module is also used for: Obtain the first position of the preset type target in the world coordinate system in the current frame; obtain the second position of the preset type target in the world coordinate system in a previous frame; The vehicle's attitude in the grid map is corrected based on the difference between the first position and the second position, wherein the position of the preset type target in the world coordinate system is obtained based on its position in the vehicle coordinate system and the position of the vehicle in the world coordinate system.
18. A vehicle, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the vehicle to perform the method as described in any one of claims 1 to 12.
19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Target identification method based on data fusion of airborne radar and infrared imaging sensor
CN101697006A
Vision target tracking method and equipment
CN108320298A
Sensor data fusion method and device, electronic equipment and storage medium
CN114528941A
Obstacle risk field environment modeling method and device and related product
CN114782912A
Parking map construction method and device, electronic equipment and storage medium
CN115690733A