A multi-sensor fusion perception method, apparatus and vehicle
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]相关技术中针对低速AEB的方案要么是仅依靠单一传感器采集感知数据并解算障碍物的位置、运动速度等状态信息,然而,单一传感器受自身感知机理限制,难以在全工况场景下满足AEB系统对障碍物检测鲁棒性、检测准确性及时效性的要求;要么是采用摄像头与毫米波雷达等多种传感器进行数据融合处理,以此获取障碍物的位置、运动速度等状态信息,然而该方式存在障碍物定位不准、传感器融合可靠性不足的问题,难以满足低速复杂场景下的安全制动需求
[0006]根据上述技术手段,首先,根据车辆状态自适应确定障碍物目标检测范围,能够贴合实际行车工况动态调整感知区域,避免无效区域检测,降低算法运算开销,提升障碍物检测的针对性与实时性。其次,对鱼眼图像同步进行图像处理与目标检测,输出鱼眼深度图与包含类别、距离、边界框、障碍物特征在内的图像检测结果,一次性完成视觉深度感知与目标属性提取,为多源数据融合提供完备的图像基础信息。然后,在目标检测范围内将超声波雷达感知的第一感知数据与鱼眼深度图进行融合处理,联合构建近场占用网格地图,可精准表征近场障碍物的位置与轮廓信息,弥补单一超声波雷达无轮廓、鱼眼深度误差大的缺陷,大幅提升车辆近场环境建模的完整性与精度。进一步,通过毫米波雷达感知的第二感知数据与图像检测结果进行特征融合,生成包含位置、速度、类别的中远场障碍物列表,充分发挥毫米波雷达测速测距稳定、鱼眼相机分类识别精准的互补优势,提升中远场障碍物检测、分类及运动状态估计的准确性。最后,综合第一感知数据、第二感知数据、图像检测结果、近场占用网格地图及中远场障碍物列表多维度信息,统一融合生成车辆周围完整的目标障碍物列表,实现近场与中远场障碍物全覆盖、无遗漏感知,避免单一传感器漏检、误检问题。如此,基于全域融合后的目标障碍物列表实施车辆制动控制,感知输入更全面、位置速度类别信息更可靠,能够精准预判碰撞风险并匹配合理制动策略,有效提升车辆主动安全防护能力,兼顾行车安全性与驾乘舒适性。
Smart Images

Figure CN122551322A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle-assisted driving technology, and in particular to a multi-sensor fusion perception method, device and vehicle. Background Technology
[0002] Automatic Emergency Braking (AEB) is a core function of advanced driver assistance systems and is crucial for improving vehicle active safety performance. In low-speed driving conditions such as urban traffic jams, automatic parking, and passing other vehicles on narrow roads, vehicles continuously face static obstacles such as pillars, traffic barriers, and stationary vehicles, as well as dynamic obstacles such as pedestrians, cyclists, and vehicles traveling in the opposite direction, posing a high risk of collision.
[0003] In related technologies, solutions for low-speed AEB either rely solely on a single sensor to collect sensing data and calculate the position, speed, and other state information of obstacles. However, a single sensor, limited by its own sensing mechanism, cannot meet the requirements of robustness, accuracy, and timeliness of obstacle detection for AEB systems in all operating conditions. Alternatively, they employ data fusion processing using multiple sensors such as cameras and millimeter-wave radar to obtain the position, speed, and other state information of obstacles. However, this approach suffers from inaccurate obstacle localization and insufficient reliability of sensor fusion, making it difficult to meet the safety braking requirements in low-speed and complex scenarios. Summary of the Invention
[0004] This application provides a multi-sensor fusion perception method, device, and vehicle that can improve the accuracy of obstacle motion information to meet the safety braking requirements in low-speed complex scenarios.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides a multi-sensor fusion sensing method, the method including: The vehicle status during driving is obtained and perception data collected by various types of sensors, including: first perception data collected by ultrasonic radar array, second perception data collected by millimeter-wave radar and fisheye images collected by fisheye camera. The target detection range for detecting obstacles around the vehicle is determined based on the vehicle's status. Image processing and target detection are performed on fisheye images to obtain fisheye depth maps and image detection results, wherein the image detection results include at least one of the following: the category of the first obstacle, the measured distance, the bounding box, and the obstacle features; Within the target detection range, the first perception data and the fisheye depth map are fused to obtain a near-field occupancy grid map of near-field obstacles of the vehicle; wherein, the near-field occupancy grid map includes at least one of the following: the position and outline of the near-field obstacles; Within the target detection range, feature fusion processing is performed on the second perception data and image detection results to obtain a mid-to-far field obstacle list for the vehicle; wherein, the mid-to-far field obstacle list includes at least one of the following: the position, speed, and category of the mid-to-far field obstacle; Based on the first perception data, the second perception data, the image detection results, the near-field occupancy grid map, and the mid-to-far-field obstacle list, a list of target obstacles around the vehicle is determined, and the vehicle is braked based on the list of target obstacles.
[0006] Based on the aforementioned technical methods, firstly, the obstacle detection range is adaptively determined according to the vehicle's state. This allows for dynamic adjustment of the perception area to match actual driving conditions, avoiding invalid area detection, reducing algorithm computational overhead, and improving the targeting and real-time performance of obstacle detection. Secondly, image processing and target detection are performed simultaneously on the fisheye image, outputting a fisheye depth map and image detection results including category, distance, bounding box, and obstacle features. This completes visual depth perception and target attribute extraction in one step, providing comprehensive image foundation information for multi-source data fusion. Then, within the target detection range, the first perception data from the ultrasonic radar and the fisheye depth map are fused to jointly construct a near-field occupancy grid map. This accurately represents the position and contour information of near-field obstacles, compensating for the shortcomings of single ultrasonic radar (lack of contour) and large fisheye depth error, significantly improving the completeness and accuracy of vehicle near-field environment modeling. Furthermore, feature fusion is performed between the second sensing data obtained by millimeter-wave radar and the image detection results to generate a mid-to-far field obstacle list containing location, speed, and category. This fully leverages the complementary advantages of the stable speed and distance measurement of millimeter-wave radar and the accurate classification and recognition of fisheye cameras, improving the accuracy of mid-to-far field obstacle detection, classification, and motion state estimation. Finally, by integrating multi-dimensional information from the first sensing data, second sensing data, image detection results, near-field occupancy grid map, and mid-to-far field obstacle list, a complete target obstacle list around the vehicle is generated, achieving full coverage and no omissions in near-field and mid-to-far field obstacle perception, avoiding the problems of missed detections and false detections by a single sensor. Thus, vehicle braking control is implemented based on the fully fused target obstacle list, with more comprehensive sensing input and more reliable location, speed, and category information. This enables accurate prediction of collision risks and matching of appropriate braking strategies, effectively improving the vehicle's active safety protection capabilities while balancing driving safety and ride comfort.
[0007] This application provides a multi-sensor fusion sensing device, including: The acquisition module is used to acquire the vehicle status during vehicle operation and the perception data collected by various types of sensors. The perception data collected by various types of sensors includes: first perception data collected by ultrasonic radar array, second perception data collected by millimeter-wave radar and fisheye images collected by fisheye camera. The determination module is used to determine the target detection range when detecting obstacles around the vehicle based on the vehicle status; The processing module performs image processing and target detection on the fisheye image to obtain a fisheye depth map and image detection results. The image detection results include at least one of the following: the category, measurement distance, bounding box, and obstacle features of a first obstacle. Within the target detection range, the module performs data fusion processing on the first perception data and the fisheye depth map to obtain a near-field occupancy grid map of near-field obstacles for the vehicle. The near-field occupancy grid map includes at least one of the following: the position and outline of near-field obstacles. Within the target detection range, the module performs feature fusion processing on the second perception data and the image detection results to obtain a mid-to-far-field obstacle list for the vehicle. The mid-to-far-field obstacle list includes at least one of the following: the position, velocity, and category of mid-to-far-field obstacles. The determination module is also used to determine the list of target obstacles around the vehicle based on the first perception data, the second perception data, the image detection results, the near-field occupancy grid map, and the mid-to-far-field obstacle list; The control module is used to control the vehicle's braking based on a list of target obstacles.
[0008] This application provides a vehicle, the vehicle comprising: Memory is used to store executable instructions or computer programs. When the processor executes computer-executable instructions or computer programs stored in the memory, it implements the multi-sensor fusion sensing method provided in the embodiments of this application.
[0009] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the multi-sensor fusion sensing method provided in this application.
[0010] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the multi-sensor fusion sensing method provided in this application. Attached Figure Description
[0011] Figure 1 This is a first flowchart illustrating the multi-sensor data fusion method provided in this application embodiment; Figure 2This is a schematic diagram of the second process of the multi-sensor data fusion method provided in the embodiments of this application; Figure 3 This is an application scenario diagram of the multi-sensor data fusion method provided in the embodiments of this application; Figure 4 This is a flowchart illustrating the method for point cloud optimization and dynamic obstacle localization of ultrasonic point clouds provided in the embodiments of this application. Figure 5 This is a flowchart illustrating the multi-level fusion decision-making method provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of the multi-sensor data fusion device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the vehicle structure provided in the embodiments of this application.
[0012] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] Reference Figure 1 As shown, Figure 1 This is a schematic diagram illustrating the implementation process of a multi-sensor data fusion method provided in an embodiment of this application. The method may include the following steps: Step 101: Obtain the vehicle status and perception data collected by various types of sensors during the vehicle's operation.
[0015] The sensing data collected by various types of sensors includes: first sensing data collected by ultrasonic radar array, second sensing data collected by millimeter-wave radar, and fisheye images collected by fisheye camera.
[0016] In this embodiment of the application, the vehicle status can be the vehicle's own real-time operating parameters and working status during driving. The vehicle status includes, but is not limited to, the vehicle's own native operating data such as vehicle speed, wheel speed, throttle opening, braking status, steering wheel angle, gear information, vehicle posture, driving mode, headlight status, and chassis condition.
[0017] In one specific embodiment of this application, the multi-sensor fusion sensing device adopts a multi-modal sensor fusion architecture, and its hardware configuration includes ultrasonic radar, millimeter-wave radar, and image sensors. The image sensors include, but are not limited to, fisheye cameras, monocular cameras, or binocular cameras. The hardware deployment of this device is as follows: For near-field perception, ultrasonic radar arrays are positioned in the front and rear bumper areas of the vehicle. These arrays are arranged in a fan shape to cover the near-field blind spots in front of and behind the vehicle, used for obstacle distance measurement when the vehicle is traveling at low speeds or reversing. To balance positioning accuracy and cost control, the ultrasonic radar array is preferably configured with 6 units; however, it can be expanded to 8 or 10 units depending on actual needs.
[0018] For mid-to-long-range perception, a 77 GHz millimeter-wave radar is integrated at the front of the roof, while two radars of the same specification are positioned at the rear of the roof. This group of sensors is mainly responsible for the detection and tracking of dynamic targets at mid-to-long range, thereby enhancing the system's environmental perception capabilities in low-speed scenarios.
[0019] In terms of panoramic visual perception, four fisheye cameras, also known as fisheye cameras, are deployed around the vehicle to build a 360-degree panoramic surround view system, enabling a complete visual reconstruction of the vehicle's surrounding environment.
[0020] It should be noted that the raw data collected by the aforementioned multi-source sensors can be transmitted to the central domain controller via the vehicle's Controller Area Network (CAN) bus or a dedicated high-speed network. This central domain controller employs an embedded computing platform with Artificial Intelligence (AI) acceleration capabilities, responsible for executing data fusion, environmental modeling, and decision-making algorithms.
[0021] In this embodiment, the perception data includes first perception data, second perception data, and fisheye images. The first perception data may include initial measurement distances of obstacles around the vehicle by different ultrasonic radars. The second perception data includes measurement distances between the vehicle and obstacles collected by different millimeter-wave radars, radial velocities between the obstacles and the vehicle, azimuth angles of the obstacles relative to the mounting axis of the millimeter-wave radars, and radar cross-sections, wherein the radar cross-section is used to distinguish obstacle categories. The fisheye images may include panoramic information about the vehicle's surrounding environment, lane lines, obstacles, curbs, pedestrians, and other visual information.
[0022] In this embodiment of the application, when the vehicle is traveling at low speed, the vehicle status under this condition is collected in real time, and the first perception data collected by the ultrasonic radar array mounted on the vehicle, the second perception data collected by the millimeter-wave radar, and the fisheye image collected by the fisheye camera are obtained, thereby providing raw data for subsequent multi-source data fusion processing.
[0023] Step 102: Determine the target detection range when detecting obstacles around the vehicle based on the vehicle status.
[0024] In this embodiment, the target detection range can be a surrounding space detection area defined with the vehicle as the center, used to limit the effective spatial boundary for the sensor to detect obstacle targets.
[0025] In this embodiment, the vehicle status may include vehicle speed and wheel speed, and the target detection range may be dynamically and adaptively adjusted according to the vehicle's speed, wheel speed, and other vehicle status.
[0026] In one feasible approach, determining the target detection range for detecting obstacles around the vehicle based on the vehicle's status can be achieved through the following process: Pre-calibrating the lateral detection width and longitudinal detection length corresponding to different vehicle speed and wheel speed ranges; matching real-time vehicle speed and wheel speed to the corresponding preset ranges; and obtaining the detection boundary parameters for that range. Using the vehicle's center or body contour as a reference, the target detection range for obstacles around the vehicle is delineated according to the retrieved boundary parameters. Thus, by pre-calibrating the correspondence between vehicle speed and wheel speed ranges and lateral and longitudinal detection parameters, the obstacle target detection range can be adaptively and dynamically adjusted according to the vehicle's real-time driving status. Target detection is only performed within the matched detection boundary range, effectively detecting and filtering obstacles perceived by ultrasonic radar, millimeter-wave radar, and fisheye cameras, ignoring invalid targets outside the range, reducing the computational power consumption of the algorithm in invalid areas, and improving the real-time performance of detection. At low speeds, the detection range is narrowed to focus on near-field obstacles; at high speeds, the detection range is expanded to accommodate distant risk targets, improving the accuracy of obstacle detection around the vehicle and the reliability of environmental perception, thus ensuring driving safety.
[0027] In another possible implementation, determining the target detection range for detecting obstacles around the vehicle based on the vehicle's state can also be achieved through the following process: determining the vehicle's minimum braking distance based on the vehicle's state, and determining the target detection range for detecting obstacles around the vehicle at the minimum braking distance.
[0028] In this embodiment, the minimum braking distance under low-speed conditions is obtained by looking up tables or calculating through a preset braking model based on the collected vehicle state parameters such as vehicle speed and wheel speed. Using this minimum braking distance as a constraint benchmark, a corresponding spatial detection boundary is set around the vehicle, thereby determining the target detection range for the vehicle to perceive and detect obstacles around the vehicle under the constraint of this minimum braking distance. Thus, by adaptively calculating the minimum braking distance based on the vehicle's real-time state and then matching the target detection range with this benchmark, the detection range closely matches the vehicle's actual braking capability, avoiding setting the detection range too large or too small. Simultaneously, associating the target detection range with the minimum braking distance allows for precise coverage of areas with safety risks during vehicle braking, improving the targeting and effectiveness of obstacle detection.
[0029] Step 103: Obtain the fisheye depth map and image detection results obtained by image processing and target detection of the fisheye image.
[0030] The image detection results include at least one of the following: the category of the first obstacle, the measured distance, the bounding box, and the obstacle features.
[0031] In this embodiment, the fisheye depth map can be obtained by a fisheye camera or an image processor on a vehicle, using monocular depth estimation or stereo vision to calculate the depth of each pixel in the fisheye image.
[0032] In this embodiment of the application, the image detection results include, but are not limited to, the category of the first obstacle, the measured distance between the first obstacle and the first obstacle, the bounding box of the first obstacle, and the obstacle features of the first obstacle; wherein, the first obstacle may be one or more obstacles contained in the fisheye image.
[0033] In this embodiment, the image detection result can be obtained by using a fisheye camera or an image processor on a vehicle to detect fisheye images through a target detection network. It should be noted that the target detection network includes, but is not limited to, the YOLOv8 network or an optimized version of the YOLOv8 network.
[0034] Step 104: Within the target detection range, perform data fusion processing on the first perception data and the fisheye depth map to obtain a near-field occupancy grid map of near-field obstacles of the vehicle.
[0035] The near-field occupancy grid map includes at least one of the following: the location and outline of near-field obstacles.
[0036] In this embodiment, the near-field occupancy grid map can be a rasterized environmental representation map that divides the near-field space of the vehicle into grid cells with the vehicle as the center, and uses the grid occupancy status to characterize whether there are obstacles in the local area, and records the location, outline and other information of the corresponding obstacles; the near-field occupancy grid map is used to structurally describe the distribution of obstacles around the near field of the vehicle.
[0037] In this embodiment, within a defined target detection range, the first sensing data output by the ultrasonic radar array and the fisheye depth map are subjected to multi-source data fusion processing to construct a near-field occupancy grid map corresponding to near-field obstacles of the vehicle. This near-field occupancy grid map contains at least one or more of the location and contour information of near-field obstacles. This combines the ranging advantage of ultrasonic radar with the visual contour advantage of fisheye depth maps, compensating for the deficiencies in perception dimensions and detection accuracy of single sensors, and improving the completeness and reliability of near-field obstacle detection. Simultaneously, data fusion within a limited target detection range effectively filters irrelevant data in the mid-to-far field, reducing computational load and improving system real-time performance and resource utilization. Finally, the structured representation of obstacle positions and contours in the form of a near-field occupancy grid map accurately depicts the distribution of near-field obstacles around the vehicle, adapting to complex near-field conditions such as low-speed driving and reversing, significantly improving vehicle driving safety and environmental perception capabilities.
[0038] Step 105: Within the target detection range, perform feature fusion processing on the second perception data and image detection results to obtain a mid-to-far field obstacle list for the vehicle.
[0039] The list of mid-to-far field obstacles includes at least one of the following: the location, speed, and category of the mid-to-far field obstacle.
[0040] In this embodiment, within a defined target detection range, feature-level fusion processing is performed on the second sensing data acquired by the millimeter-wave radar and the image detection results output by the fisheye camera. Through the mutual fusion and complementarity of the two types of sensor information, a mid-to-far-field obstacle list is generated to characterize the vehicle's mid-to-far-field obstacle information. This list includes at least one or more of the following information: the position, speed, and obstacle category of the mid-to-far-field obstacles. Thus, fusing the feature information of the millimeter-wave radar and the visual image can compensate for the shortcomings of a single sensor in mid-to-far-field target recognition, speed detection, and category determination, effectively improving the perception accuracy and detection stability of mid-to-far-field obstacles. Simultaneously, performing feature fusion and target generation only within a limited target detection range avoids redundant processing of irrelevant data, reduces the computational overhead of the algorithm, and further improves the real-time output of the perception results.
[0041] Step 106: Based on the first perception data, the second perception data, the image detection results, the near-field occupancy grid map, and the mid-to-far-field obstacle list, determine the target obstacle list around the vehicle, and perform braking control on the vehicle according to the target obstacle list.
[0042] In this embodiment of the application, the target obstacle list includes, but is not limited to: the location, category, speed, and target confidence of the target obstacle, wherein the obstacle is an obstacle.
[0043] In this embodiment, within the target detection range, feature fusion processing is performed on the second sensing data and image detection results to obtain a mid-to-far field obstacle list for the vehicle. Based on the first sensing data and the second sensing data, the image detection results, near-field occupancy grid map, and mid-to-far field obstacle list are integrated to complete the parsing of grid information of near-field obstacles and the extraction of target information of mid-to-far field obstacles. Cross-sensor target association, duplicate target deduplication, and spatiotemporal consistency verification are performed on near-field and mid-to-far field obstacles to comprehensively determine and generate a target obstacle list covering the entire perimeter of the vehicle.
[0044] In this embodiment, the target obstacle list includes the location, category, and speed of the target obstacles. Braking control of the vehicle based on the target obstacle list can be achieved in the following way: determining the minimum braking distance of the vehicle under the current operating conditions based on the vehicle status; determining the target distance between the target obstacle and the vehicle based on the location of the target obstacle; classifying the driving risk according to the preset risk level judgment logic based on the relationship between the target obstacle category, speed, target distance, and minimum braking distance; and executing the corresponding graded braking control strategy for the vehicle according to the risk level.
[0045] In this embodiment, firstly, the category and speed of the target obstacle, the target distance between the vehicle and the obstacle, and the minimum braking distance of the vehicle under the current operating condition are uniformly converted to the same vehicle coordinate system for quantitative representation, establishing the correlation constraint relationship of the four types of parameters. Secondly, based on the obstacle category, different safety distance thresholds, speed thresholds, and braking distance thresholds are pre-calibrated for different categories; and multiple driving risk levels are divided according to the degree of danger from low to high. Then, the obstacle category, speed, target distance, and minimum braking distance acquired in real time are compared and calculated; the difference between the target distance and the minimum braking distance is compared, combined with the relative speed change trend of the obstacle, and the danger weight of the obstacle category is matched, and substituted into the preset risk level judgment logic to comprehensively determine the current driving risk level. Furthermore, braking control strategies are pre-configured for each risk level, including but not limited to different control levels such as safety warning, gradual deceleration braking, moderate deceleration braking, and automatic emergency braking; based on the actual risk level determined, the matched braking control strategy is invoked, and the corresponding braking torque, deceleration, and braking response timing are output to implement graded braking closed-loop control of the vehicle. In this way, risk assessment is conducted by combining multiple dimensions of parameters, including obstacle type, speed, target distance, and vehicle minimum braking distance, avoiding the one-sidedness of judgment based on a single parameter and significantly improving the accuracy and rationality of driving risk level classification. Different judgment thresholds and risk weights are set for different obstacle types, which can distinguish the degree of danger of different targets such as pedestrians, vehicles, and fixed obstacles, adapting to the safety protection needs of different types of obstacles. A graded risk assessment matching graded braking control strategy is adopted, which can sequentially execute gradient control such as warning, gradual deceleration, moderate braking, and automatic emergency braking according to the level of risk. The braking intervention is smooth and orderly, taking into account driving comfort and passenger experience.
[0046] In some embodiments, step 104 involves performing data fusion processing on the first perception data and the fisheye depth map within the target detection range to obtain a near-field occupancy grid map of near-field obstacles of the vehicle, which can be achieved through the following process.
[0047] The first sensing data within the target detection range is projected and associated with the fisheye depth map to match the ultrasonic measurement points of the same near-field obstacle in the first sensing data with the pixel regions in the fisheye depth map; using the initial measurement distance in the first sensing data, the depth estimation of the corresponding near-field obstacle in the fisheye depth map is corrected for near-field error to obtain the corrected fisheye depth map; based on the corrected fisheye depth map and the first sensing data, a near-field occupancy grid map is generated.
[0048] In this embodiment, firstly, by spatially projecting and associating the first sensing data within the target detection range with the fisheye depth map, precise registration of ultrasonic measuring points and image pixel regions is achieved, providing a spatial alignment basis for multi-source data fusion correction. Furthermore, data fusion is carried out within a limited target detection range, effectively filtering irrelevant data in the mid-to-far field region, reducing computational load, and improving system real-time performance and resource utilization. Secondly, leveraging the high accuracy of ultrasonic ranging, error correction is performed on the fisheye depth, effectively suppressing deviations in fisheye depth estimation and improving near-field depth detection accuracy. Finally, the fused and corrected depth map and the first sensing data jointly construct a grid map, which can more realistically and completely restore the position and contour distribution of near-field obstacles, improve environmental modeling accuracy, and provide reliable sensing input for subsequent obstacle risk assessment and vehicle graded braking control.
[0049] In some embodiments, step 105 involves performing feature fusion processing on the second perception data and image detection results within the target detection range to obtain a mid-to-far field obstacle list for the vehicle, which can be achieved through the following process.
[0050] The second sensing data within the target detection range is projected onto the fisheye image plane through spatial transformation and matched with the bounding box of the first obstacle to obtain the position of the mid-to-far field obstacle. A unified obstacle feature descriptor is constructed using the second sensing data and the obstacle features of the first obstacle. The second sensing data includes at least one of the following: the measured distance, radial velocity, azimuth angle, and radar cross-section of the second obstacle. The obstacle feature descriptor is input into a lightweight fusion neural network or a traditional classifier for joint target recognition and state estimation to obtain a mid-to-far field obstacle list, wherein the mid-to-far field obstacle list includes the velocity and category of the mid-to-far field obstacles.
[0051] In this embodiment, the second obstacle can be an obstacle identified by millimeter-wave radar.
[0052] In this embodiment, the second sensing data acquired by the millimeter-wave radar includes the measured distance, radial velocity, azimuth angle, and radar cross-section of the second obstacle. The image detection result includes the category, measured distance, bounding box, and obstacle features of the first obstacle. The first and second obstacles at least partially overlap. This embodiment first maps the second sensing data corresponding to the millimeter-wave radar to the fisheye image plane through spatial coordinate transformation, achieving cross-modal correlation matching between the radar-detected target and the bounding box of the first obstacle in the image, thereby obtaining stable position information of the mid-to-far field obstacle. Thus, by using spatial projection and bounding box correlation matching, cross-modal spatial alignment of the millimeter-wave radar data and the fisheye image is achieved, accurately matching the same mid-to-far field obstacle and improving the accuracy of target correlation. Then, one or more of the distance, radial velocity, azimuth angle, and radar cross-section of the second obstacle acquired by the millimeter-wave radar are fused with the image obstacle features to construct a unified obstacle feature descriptor. This fully leverages the advantages of stable radar ranging and velocity measurement and accurate visual classification and recognition, resulting in more comprehensive feature representation. Finally, the feature descriptor is input into a lightweight fusion neural network or a traditional classifier to complete cross-modal joint target recognition and motion state estimation, ultimately obtaining a mid-to-far field obstacle list containing obstacle speed and category. In this way, by using a lightweight fusion neural network or a traditional classifier for joint target recognition and state estimation, the computational overhead is reduced while ensuring recognition accuracy through collaborative perception of multi-source information from radar and vision, thus adapting to the real-time deployment requirements of vehicle-mounted terminals.
[0053] In some embodiments, the process of determining the list of target obstacles around the vehicle in step 106 based on the first sensing data, the second sensing data, the image detection result, the near-field occupancy grid map, and the mid-to-far-field obstacle list is combined with Figure 2 The following explanation is provided.
[0054] Step 201: Process the first perception data to obtain the first obstacle processing result around the vehicle.
[0055] The results of the first obstacle processing include the position and estimated speed of the third obstacle.
[0056] In this embodiment, the third obstacle may be a dynamic obstacle detected by ultrasonic radar.
[0057] In this embodiment, the estimated velocity of the third obstacle can be estimated by measuring the positional changes of the third obstacle in consecutive frames acquired by the ultrasonic radar. Alternatively, the estimated velocity of the third obstacle can be obtained by using an unscented Kalman filter (UKF) to estimate and track the state of the third obstacle.
[0058] In this embodiment of the application, step 201 processes the first perception data to obtain the first obstacle processing result around the vehicle, which can be achieved through the following process.
[0059] Step 211: Based on the initial measurement distance in the first perception data, determine the initial velocity of each third obstacle and the position of the target ultrasonic radar closest to each third obstacle.
[0060] In this embodiment of the application, the target ultrasonic radar can be a certain number of ultrasonic radars closest to the third obstacle. The number of target ultrasonic radars can be 1, 2 or 3. For example, the number of target ultrasonic radars can be 3.
[0061] In this embodiment, the initial velocity can be determined based on the initial measured distance in the first sensing data and the acquisition time of the first sensing data. For example, the initial velocity of the third obstacle can be obtained as follows: Multiple frames of first sensing data are continuously acquired, and the temporal data of the initial measured distance corresponding to each third obstacle in each frame is extracted; continuous frame target association is performed on the same third obstacle, and the ranging information of the same obstacle at different times is matched to establish a temporal sequence of distance changes over time; based on the difference in initial measured distance between adjacent sampling times and the sampling time interval, the relative initial velocity of each third obstacle relative to the vehicle is calculated through differential operation and temporal filtering algorithms; combined with the vehicle's own speed, wheel speed, and other vehicle state parameters, coordinate system calculation and motion compensation are performed to convert the absolute initial velocity of each third obstacle; the calculated initial velocity is subjected to outlier removal and smoothing filtering to correct the velocity estimation error caused by ranging jitter, and a stable and reliable initial velocity of each third obstacle is output.
[0062] In some embodiments, before determining the initial velocity of each third obstacle based on the initial measurement distance in the first sensing data in step 211, the following steps may also be performed: obtaining historical frame sensing data collected by the ultrasonic radar array; using a random sampling consensus algorithm to fit and filter the static background in the continuous frames of the historical frame sensing data to obtain static background sensing data; and removing the static background sensing data from the first sensing data.
[0063] In this embodiment, firstly, multiple frames of historical sensing data from the ultrasonic radar array are obtained. Then, the Random Sample Consensus (RANSAC) algorithm is used to perform static background fitting and filtering on the continuous frame data, accurately constructing an environmental static background sensing model. This effectively suppresses ultrasonic radar ranging outliers, environmental clutter, and random noise interference, improving the accuracy and robustness of static background modeling. Furthermore, this static background sensing data is removed from the first sensing data, separating and removing fixed static interference targets such as walls, curbs, and the ground, retaining only the effective sensing data of dynamic obstacles and newly added obstacles. This filters out redundant information from fixed facilities, reducing the computational overhead of subsequent data processing; simultaneously, it reduces interference with dynamic obstacle detection, significantly improving the accuracy and reliability of subsequent obstacle location calculation, velocity estimation, and near-field occupancy grid map construction.
[0064] Step 212: Based on the position of the target ultrasonic radar and the initial measurement distance of the target ultrasonic radar to the corresponding third obstacle, construct a system of spherical equations.
[0065] Step 213: Calculate the spherical equations using the nonlinear least squares method to obtain the initial positions of each third obstacle.
[0066] In this embodiment, the vehicle obtains the initial speed of each third obstacle based on the initial measured distance in the first perception data. Further, for the same third obstacle, a first number of target ultrasonic radars closest to it are selected, and the actual installation position of each target ultrasonic radar is obtained. It should be noted that the positions of the first number of target ultrasonic radars are adjacent; then, based on the installation position of the target ultrasonic radars... And the initial measurement distance of the corresponding third obstacle obtained by ultrasonic radar detection of each target. Establish the system of equations for the sphere as shown in formula (1) below.
[0067]
[0068] Furthermore, the nonlinear least squares method is used to solve the spherical equations to obtain the initial positions of each third obstacle. In this way, the one-dimensional distance information of the third obstacles acquired by the ultrasonic array is transformed into two-dimensional planar positions.
[0069] Step 214: Based on the position and velocity of each third obstacle at the previous moment, use the unscented Kalman filter state estimation and tracking algorithm to continuously smooth the initial position and initial velocity of the third obstacle at the current moment, and obtain the smoothed position and velocity of each third obstacle at the current moment.
[0070] In this embodiment of the application, the nonlinear least squares method is used to calculate the spherical equation system to obtain the initial position of each third obstacle. Then, the state vector is designed, the state transition equation is established using the uniform motion model, and the measurement equation is established. The state vector can be represented by the following formula (2), the state transition equation can be represented by the following formula (3), and the measurement method can be represented by the following formula (4).
[0071]
[0072] in, The location of the third obstacle. The speed of the third obstacle.
[0073]
[0074] in, Let K be the state vector of the third obstacle at time k-1. Let be the state vector of the third obstacle at time k, A be the state transition matrix, and w be the process noise.
[0075]
[0076] Where h(.) is the nonlinear measurement function, v is the measurement noise, and n is the number of target sensors.
[0077] Furthermore, based on the position and velocity of each third obstacle at the previous moment, an initial state vector is determined, and this initial state vector is input into the state transition equation to determine the next state vector of the third obstacle. Further, based on the next state vector, the state vector is mapped back to the reference measurement distance of the ultrasonic radar corresponding to the third obstacle using the measurement equation. Then, based on the initial state vector, the unscented Kalman filter (UKF) state estimation and tracking algorithm is used to predict the state vector and covariance according to the state transition model, and the predicted measurement distance of the third obstacle is calculated. The predicted measurement distance is compared with the reference measurement distance, the Kalman gain is calculated, and the state vector and covariance are updated. Continuously running UKF yields smooth and stable smoothed positions and velocities of each third obstacle at the current moment. Thus, by using UKF to process nonlinear measurements through unscented transformation (UT), continuous and smooth estimation of the position and velocity of dynamic obstacles is achieved, effectively filtering out measurement noise and predicting obstacle trajectories. Compared to traditional triangulation, this method significantly improves positioning accuracy and dynamic tracking capabilities.
[0078] In some embodiments, step 214 uses an unscented Kalman filter state estimation and tracking algorithm to continuously smooth the initial position and initial velocity of the third obstacle at the current moment based on the position and velocity of each third obstacle at the previous moment, so as to obtain the smoothed position and velocity of each third obstacle at the current moment. This can also be achieved through the following process.
[0079] Historical frame sensing data acquired by an ultrasonic radar array is obtained; an improved density-based spatial clustering algorithm is used to cluster and track spatially adjacent points in consecutive frames of the historical frame sensing data to generate dynamic obstacle clusters; for each dynamic obstacle cluster, based on the position and velocity of each dynamic obstacle at the previous moment, an unscented Kalman filter state estimation and tracking algorithm is used to continuously smooth the initial position and initial velocity of the dynamic obstacle at the current moment to obtain the smoothed position and velocity of the dynamic obstacle at the current moment; wherein, the third obstacle includes dynamic obstacles.
[0080] In this embodiment, firstly, historical frame sensing data of the ultrasonic radar array is obtained. An improved Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is used to cluster adjacent measurement points within consecutive frames using the time dimension, dividing and forming independent dynamic obstacle clusters. This effectively clusters ultrasonic measurement points, eliminates discrete clutter measurement points, and achieves accurate clustering and continuous tracking of dynamic obstacles, reducing the probability of missed and false detections. Furthermore, using the position and velocity state of the dynamic obstacle at the previous moment, the UKF algorithm is continuously run to perform state estimation and temporal tracking. The obtained initial position and initial velocity are then temporally smoothed and iteratively updated to output stable and accurate current position and velocity information. This effectively suppresses position and velocity estimation errors caused by ranging jitter and measurement noise, providing high-precision and high-reliability state input for subsequent obstacle distance calculation, risk level determination, and vehicle graded braking control, thereby improving the overall vehicle's active safety protection performance.
[0081] In some embodiments, after using UKF to estimate the state of the third obstacle and obtain the smoothed velocity, the velocity can also be estimated using the position changes of the third obstacle in continuous frames collected by ultrasonic radar. This velocity is used as a pseudo-measurement and is checked for consistency with the smoothed velocity, thereby eliminating false obstacles with abnormal velocity.
[0082] Step 202: Cluster and extract targets from the second perception data to obtain the processing results of the second obstacle around the vehicle.
[0083] The second obstacle processing result includes at least one of the following: the position, speed, trajectory and category of the second obstacle, and the second obstacle and the third obstacle are at least partially different.
[0084] In this embodiment, point cloud clustering algorithm and target extraction algorithm are used to cluster and extract targets from the second perception data, respectively, so as to obtain the processing result of the second obstacle in the far field of the vehicle.
[0085] Step 203: Determine the target obstacle list based on the first obstacle processing result, the second obstacle processing result, the image detection result, the near-field occupancy grid map, and the mid-far-field obstacle list.
[0086] In some embodiments, step 203, which determines the target obstacle list based on the first obstacle processing result, the second obstacle processing result, the image detection result, the near-field occupancy grid map, and the mid-far-field obstacle list, can be implemented through the following process.
[0087] Step 231: Cluster the first obstacle processing result, the second obstacle processing result, the image detection result, the near-field occupancy grid map, and the mid-far-field obstacle list according to their spatial location to obtain the target obstacles in the same spatial location.
[0088] In this embodiment of the application, after obtaining the first obstacle processing result, the second obstacle processing result, the image detection result, the near-field occupancy grid map, and the mid-to-far-field obstacle list, the first obstacle processing result, the second obstacle processing result, the image detection result, the near-field occupancy grid map, and the mid-to-far-field obstacle list are clustered and associated according to their spatial locations. If the same spatial location highly overlaps, the obstacles at that spatial location are determined to be the same obstacle, and that obstacle is determined to be the target obstacle.
[0089] Step 232: Configure the corresponding confidence levels for the first obstacle processing result, the second obstacle processing result, and the image detection result, respectively.
[0090] The confidence level is determined based on the characteristics of each type of sensor, the class matching degree of the target obstacle, the stability of distance measurement, and the continuity of historical frame tracking.
[0091] In this embodiment, based on the inherent characteristics of each type of sensor, the category matching degree of the target obstacle, the stability of distance measurement, and the continuity of historical frame tracking, the confidence level of the results of each type of sensor is dynamically assigned. Specifically, a first confidence level corresponding to the first obstacle processing result, a second confidence level corresponding to the second obstacle processing result, and a third confidence level corresponding to the image detection result are dynamically assigned. For example, ultrasonic radar has high confidence in near-field static / low-speed obstacles (also known as targets), millimeter-wave radar has high confidence in the radial velocity measurement of dynamic obstacles, and fisheye camera has high confidence in obstacle type identification.
[0092] Step 233: For each target obstacle, select the target obstacle with the highest confidence level from multiple confidence levels, and determine the category of the target obstacle with the highest confidence level as the target category.
[0093] In this embodiment of the application, for each target obstacle, the target obstacle with the highest confidence level is selected from the first confidence level corresponding to the first obstacle processing result, the second confidence level corresponding to the second obstacle processing result, and the third confidence level corresponding to the image detection result, and the category of the target obstacle with the highest confidence level is determined as the target category.
[0094] For example, taking the pedestrian perception scenario as an example, at time t, 3 meters directly in front of the vehicle, the fisheye camera detects a "pedestrian" bounding box with a confidence level of 0.9; the millimeter-wave radar clusters a target at this location with a radial velocity of 0.5 m / s (moving away) and a confidence level of 0.8; the ultrasonic radar UKF tracker has a stable trajectory at this location with a velocity (vx, vy) = (0, 0.5 m / s) and a confidence level of 0.7. At this point, the fisheye vision has the highest confidence level and determines the category as "pedestrian".
[0095] Step 234: Determine the position weight of each type of sensor for the target obstacle based on multiple confidence levels.
[0096] In this embodiment, the position weights of each type of sensor relative to the target obstacle can be the confidence levels corresponding to the processing results of each sensor. Of course, the position weights of each type of sensor relative to the target obstacle can also be determined based on the confidence levels corresponding to the processing results of each sensor, as well as the distance measurement variance in the processing results. This application does not impose specific limitations in this regard.
[0097] Step 235: Based on multiple position weights, perform weighted fusion of the positions of the target obstacle in the first obstacle processing result, the second obstacle processing result, and the image detection result to determine the target position of the target obstacle.
[0098] Step 236: Perform weighted fusion of the estimated velocity and the radial velocity in the second perception data to determine the target velocity of the target obstacle; wherein, the target obstacle list includes the target category, target position, target velocity and target confidence of the target obstacle, and the target confidence includes the highest confidence.
[0099] In this embodiment, after determining the position weights of each type of sensor for the target obstacle based on multiple confidence levels, the positions of the target obstacle in the first obstacle processing result, the second obstacle processing result, and the image detection result are weighted and fused according to multiple position weights to determine the target position of the target obstacle. The predicted velocity and the radial velocity in the second sensing data are weighted and fused to determine the target velocity of the target obstacle.
[0100] In some embodiments, if at least two of the first obstacle processing result, the second obstacle processing result, and the image detection result include the target obstacle, the highest confidence level is increased, and the increased highest confidence level is used as the target confidence level. Conversely, if at most one of the first obstacle processing result, the second obstacle processing result, and the image detection result includes the target obstacle, the highest confidence level is decreased, and the decreased highest confidence level is used as the target confidence level.
[0101] In this embodiment, the vehicle performs weighted fusion calculations on the target obstacle position output by the first perception data (ultrasonic radar), the target obstacle position output by the second perception data (millimeter-wave radar), and the target obstacle position information in the image detection results, based on the determined position weights corresponding to each type of sensor. By assigning weights, the advantages of the position data from highly reliable sensors are highlighted, and the precise initial position of the target obstacle is finally determined. At the same time, the estimated speed and radial speed are weighted and fused according to the confidence weights of the corresponding sensors, based on the solved speed of the target obstacle and the radial speed extracted from the second perception data, to offset the error of measurement by a single sensor. Finally, a stable and accurate target speed of the target obstacle is output, providing complete state parameter support for subsequent risk assessment and braking control.
[0102] As described above, the embodiments of this application employ a weighted fusion method, fully combining the advantages of different sensors. Ultrasonic radar provides accurate short-range ranging, millimeter-wave radar offers stable velocity measurement, and visual image positioning is intuitive. By assigning weights, the measurement data from highly reliable sensors plays a dominant role, effectively avoiding the limitations of single-sensor measurements and improving the accuracy of determining the position and velocity of the target obstacle. Position and velocity are weighted and fused separately to specifically address the jitter and deviation problems that may occur in single-sensor measurements, making the motion state of the target obstacle more realistic and avoiding misjudgments caused by single-sensor failures or measurement errors.
[0103] The process of the method provided in this application will be illustrated below through a specific embodiment.
[0104] Automatic emergency braking (AEB) systems, as a core function of advanced driver assistance systems (ADAS), are crucial for enhancing vehicle active safety performance. Especially in low-speed scenarios (such as urban traffic jams, automatic parking, and passing on narrow roads), vehicles frequently need to address the collision risks of static obstacles (such as pillars, barriers, and stopped vehicles) and dynamic obstacles (such as pedestrians, cyclists, and cross-traffic vehicles). Existing low-speed AEB systems mostly rely on single or a few sensors, which presents at least the following problems: ultrasonic radar is low-cost and reliable for short-range measurements, but its point cloud is sparse and its azimuth resolution is low, making it difficult to accurately locate dynamic obstacles and distinguish between static and dynamic backgrounds; millimeter-wave radar excels in speed and distance measurement, but its ability to classify static targets is weak and it is prone to clutter; fisheye cameras offer a wide field of view and provide rich semantic information, but they are greatly affected by lighting and weather conditions and lack depth information. A single sensor cannot meet the robustness, accuracy, and timeliness requirements of AEB systems for obstacle detection under all operating conditions.
[0105] To overcome the limitations of single sensors, multi-sensor fusion has become an inevitable trend. Existing fusion technologies are mostly concentrated in medium-to-high-speed scenarios, employing data-level or feature-level fusion, which places high demands on computing resources. For low-speed AEB (Autonomous Emergency Braking), the system needs to react at extremely close ranges (typically less than 5 meters) and within extremely short timeframes (typically less than 1 second), placing special demands on the real-time performance, reliability, and accurate tracking capabilities of the perception algorithm. Related technologies either employ a fusion of ultrasonic radar and camera perception, however, this approach suffers from weak dynamic obstacle recognition and low positioning accuracy; or a fusion of lidar and camera perception, however, this approach suffers from high lidar costs and poor economic efficiency.
[0106] To address the issues of inaccurate dynamic obstacle localization and insufficient sensor fusion reliability in low-speed AEB systems in related technologies, this application provides a low-speed AEB multi-sensor fusion perception system (corresponding to the aforementioned multi-sensor fusion perception device) and a low-speed AEB multi-sensor fusion perception method. This method, through hierarchical fusion and dynamic optimization, achieves redundant and accurate obstacle localization beyond the minimum braking distance, ensuring driving safety in low-speed scenarios.
[0107] The system hardware of the low-speed AEB multi-sensor fusion perception system in this embodiment includes: six ultrasonic radars mounted on the front bumper to cover a fan-shaped area in front, and six ultrasonic radars mounted on the rear bumper to cover a fan-shaped area behind, used to measure the distance to obstacles when the vehicle is moving forward or reversing; one 77GHz millimeter-wave radar mounted on the front of the roof and two 77GHz millimeter-wave radars mounted on the rear of the roof, used for mid-to-long-range target detection; and four fisheye cameras mounted around the vehicle body to construct a panoramic perception system. It should be noted that all sensor data is transmitted to the central domain controller via the vehicle's CAN bus or a dedicated high-speed network, wherein the controller employs an embedded system with AI acceleration capabilities.
[0108] Figure 3 This application provides a specific application scenario diagram of a low-speed AEB multi-sensor fusion sensing method, as illustrated in the embodiments of this application. Figure 3 As shown, the application scenario includes: a bottom sensor layer 310, a vehicle status assistance channel 320, an intermediate fusion layer 330, and a decision output layer 340.
[0109] The bottom sensor layer includes: an ultrasonic radar array 311 and a corresponding ultrasonic data processing module 312, a fisheye camera 313 and a corresponding visual processing module 314, and a millimeter-wave radar 315 and a corresponding millimeter-wave radar data processing module 316.
[0110] The ultrasonic data processing module 312 optimizes the ultrasonic point cloud (corresponding to the first sensing data mentioned above) collected by the ultrasonic radar array 311 and performs dynamic obstacle localization to obtain near-field data around the vehicle (corresponding to the first obstacle processing result mentioned above). Here, the ultrasonic data processing module includes a spatiotemporal synchronization module, a point cloud optimization module, and a UKF localization module. The spatiotemporal synchronization module is used for spatiotemporal synchronization of multiple ultrasonic arrays; the point cloud optimization module optimizes the accuracy of the ultrasonic point cloud to avoid noise caused by environmental interference and multiple reflections; and the UKF localization module locates dynamic obstacles in the ultrasonic point cloud, using the resulting dynamic obstacle tracking list containing the position and velocity of the dynamic obstacles as near-field data.
[0111] In one feasible approach, continue to refer to Figure 3 As shown, the process of using the ultrasonic data processing module to optimize the ultrasonic point cloud and locate dynamic obstacles is as follows: Figure 4 As shown.
[0112] Step 401: Obtain the ultrasonic point cloud acquired by the ultrasonic radar array.
[0113] Step 402: Use the time-space synchronization module to align and synchronize the ultrasonic point cloud with timestamps to obtain the aligned ultrasonic point cloud.
[0114] The original ultrasonic point cloud includes the measured distance of each obstacle from each ultrasonic radar.
[0115] Here, the spatiotemporal synchronization module can trigger multiple ultrasonic radars deployed on the front and rear bumpers of the vehicle at a fixed frequency, such as by using hardware triggering or software timestamp synchronization mechanisms, to ensure that the data from each sensor are on the same time reference, laying the foundation for subsequent triangulation measurements.
[0116] Step 403: Based on the aligned ultrasonic point cloud, the initial position of each obstacle is calculated using triangulation to obtain the ultrasonic position point cloud.
[0117] It should be noted that for multiple echoes from the same obstacle, a spherical equation set is established as described above (1) using the known sensor installation location and measurement distance. The nonlinear equation set is solved using the least squares method or analytical method to obtain the initial position estimate of each obstacle in the vehicle coordinate system. In this way, the one-dimensional distance information is transformed into a two-dimensional planar position.
[0118] In a feasible scenario, consider a pedestrian crossing laterally in front of a vehicle. Three adjacent ultrasonic radars, 0, 1, and 2, simultaneously receive echoes from the pedestrian, obtaining measured distances d0, d1, and d2. The system uses hardware synchronization to ensure that the three measured distances correspond to the same moment. Based on the installation positions of the three ultrasonic radars (x0, y0), (x1, y1), and (x2, y2), three spherical equations are constructed. These equations are solved using the nonlinear least squares method to obtain the pedestrian's initial position estimate P0 = (x0, y0).
[0119] Step 404: Use the RANSAC algorithm in the point cloud optimization module to filter the static background point cloud in the ultrasonic location point cloud.
[0120] Here, cluster analysis can be performed on the ultrasonic location point cloud of multiple consecutive frames. For example, the RANSAC algorithm is used to fit the main plane (such as the ground or wall), and the fitted points are treated as static background and removed, thereby obtaining the ultrasonic location point cloud of dynamic obstacles and reducing interference with the detection of dynamic obstacles.
[0121] Step 405: Use the improved DBSCAN clustering algorithm in the point cloud optimization module to perform spatiotemporal clustering on the ultrasonic location point cloud to obtain dynamic obstacle candidate clusters.
[0122] Here, an improved DBSCAN clustering algorithm is adopted, which introduces the time dimension to perform correlation clustering on spatially adjacent points in consecutive frames, forming dynamic obstacle candidate clusters.
[0123] Step 406: For each candidate cluster, based on the ultrasonic location point cloud, use the UKF positioning module to track and locate the dynamic obstacle.
[0124] Here, the processing steps of the UKF localization module include: designing a state vector X = [x, y, vx, vy]^T, containing the position and velocity of the dynamic obstacle, and an initial state vector X_0 = [x0, y0, 0, 0]^T. Based on the initial state vector, the next state vector (i.e., position and velocity) of the obstacle is determined using the state transition equation as shown in formula (3) above. In the subsequent frame k, based on the state vector, the state is mapped back to the distance measurement values of multiple sensors using the measurement equation as shown in formula (4) above, thereby obtaining the actual measured distance. After that, UKF performs a prediction step, specifically, predicting X_k|k-1 and covariance P_k|k-1 according to the state transition model. Then, an update step is performed, that is, calculating the predicted measured distance and comparing it with the actual measured distance, calculating the Kalman gain, and updating the state X_k and covariance P_k. By continuously running UKF, the smooth position (x_k, y_k) and velocity (vx_k, vy_k) of the dynamic obstacle can be obtained. Meanwhile, the point cloud optimization module will continuously filter out ground echoes and transient clutter to ensure that the input to the UKF is a valid dynamic target point cloud. It should be noted that for each candidate cluster, the above-mentioned UKF is applied for tracking. Only clusters that can be stably tracked for a certain number of frames are identified as effective dynamic obstacles, thereby suppressing instantaneous noise.
[0125] It is understandable that standard ultrasonic radar lacks Doppler velocity measurement capability. In some feasible approaches, after obtaining the velocity of a dynamic obstacle using UKF (Ultra-Knowledge Fault Factor), Doppler velocity can be used for verification. For example, velocity can be estimated using positional changes across consecutive frames. This velocity is then used as a spurious measurement and its consistency with the UKF state-predicted velocity is checked to eliminate false targets with anomalous velocities.
[0126] Step 407: Based on the frame number threshold, check whether UKF tracking is stable.
[0127] Here, the number of frames in which the position change of the dynamic obstacle meets the change requirements is obtained. If the number of frames is greater than or equal to the frame number threshold, it indicates that the UFK tracking is stable, and step 408 can be executed; if the number of frames is less than the frame number threshold, it indicates that the UFK tracking is unstable, and step 409 can be executed.
[0128] Step 408: Output the dynamic obstacle tracking list.
[0129] The dynamic obstacle tracking list includes the position and speed of the dynamic obstacles.
[0130] Step 409: Remove the dynamic obstacle.
[0131] Here, if the number of frames is less than the frame number threshold, it indicates that UFK tracking is unstable, and the dynamic obstacle can be removed as a false target.
[0132] The millimeter-wave radar data processing module 316 performs point cloud clustering and target extraction on the millimeter-wave point cloud collected by the millimeter-wave radar 315 to obtain a list of millimeter-wave targets.
[0133] The aforementioned visual processing module 314 performs image processing, target detection, and depth estimation on the fisheye image acquired by the fisheye camera 313, thereby obtaining the target detection result (corresponding to the image detection result mentioned above) and the depth estimation map. The visual processing module includes an image distortion correction module, a target detection module, and a depth estimation module. Specifically, the image distortion correction module performs distortion correction on the fisheye image; the target detection module detects targets in the fisheye image, using the obtained target detection result as the visual feature of the fisheye image; the target detection module can utilize a target detection network such as the YOLOv8 network to output the target detection result; and the depth estimation module estimates the pixel depth in the fisheye image, using the obtained depth estimation map as the visual data of the fisheye image; the depth estimation module can use a monocular depth estimation algorithm or a stereo vision algorithm to estimate the depth of the fisheye image.
[0134] The vehicle status auxiliary channel 320 is used to determine the minimum braking distance of the vehicle based on the vehicle status, such as vehicle speed and wheel speed, through the minimum braking distance calculation module 321, and generate a constraint detection range for obstacles based on the minimum braking distance.
[0135] The intermediate fusion layer 330 includes a data-level fusion module 331, a feature-level fusion module 332, and a decision-level fusion module 333.
[0136] The aforementioned data-level fusion module 331 is used for near-range redundancy fusion (e.g., less than 3 meters) of ultrasonic radar and fisheye camera. Within the constrained detection range, the sparse point cloud of the ultrasonic radar is projected and correlated with the depth estimation map of the fisheye camera; the distance values in the ultrasonic point cloud are used to correct the near-field error of obstacles in the visual depth estimation map, thereby generating a dense near-field obstacle occupancy grid map in the vehicle coordinate system. This fusion layer provides the highest confidence near-range obstacle contours.
[0137] The aforementioned feature-level fusion module 332 is used for mid-to-long-range complementary fusion of millimeter-wave radar and fisheye camera. The millimeter-wave radar outputs a millimeter-wave point cloud containing the range, angle, and radial velocity of obstacles (also known as targets). The fisheye camera outputs a feature vector containing obstacle category, bounding box, and obstacle information through a target detection network. Within the constrained detection range, the millimeter-wave point cloud is projected onto the image plane of the fisheye camera through spatial transformation, and associated with and matched with the obstacle bounding box in the fisheye image to obtain information on the stable position of the obstacle. Using the radial velocity and range features measured by the radar, as well as the obstacle features from the camera, a unified obstacle feature descriptor is constructed and input into a lightweight fusion neural network or a traditional classifier (such as SVM) for joint target recognition and state estimation to obtain the target category of the obstacle. Thus, a mid-to-long-range target list containing obstacle category, range, and obstacle bounding box is obtained.
[0138] The aforementioned decision-level fusion module 333, acting as the final arbitration layer, unifies the near-field occupancy grid from data-level fusion, the mid-to-far-field target list from feature-level fusion, and the detection results processed independently by each sensor, all within the vehicle coordinate system. A rule-based fusion strategy based on confidence and safety priority is employed. Specifically, the multi-level fusion decision-making process combines... Figure 5 Please provide an explanation.
[0139] Step 501: Obtain the detection results processed independently by each sensor.
[0140] Here, the detection results obtained by each sensor independently include: the dynamic obstacle tracking list of ultrasonic waves, the millimeter-wave target list of millimeter-wave radar, and the target detection results of fisheye cameras.
[0141] Step 502: Calculate the minimum braking distance at the current vehicle speed.
[0142] Step 503: Determine whether the distance to the obstacle is less than the minimum braking distance.
[0143] Here, if the obstacle distance in the detection result is less than the minimum braking distance D_min, then step 504 is executed; if the obstacle distance in the detection result is greater than or equal to the minimum braking distance, then step 505 is executed.
[0144] Step 504: Directly trigger the emergency warning to enter the AEB decision module.
[0145] Step 505: Convert all detection results to the vehicle coordinate system.
[0146] Here, all detection results include: the near-field occupancy grid output by the data-level fusion module, the mid- and far-field target list output by the feature-level fusion module, and the detection results processed independently by each sensor.
[0147] Step 506: Cluster the detection results from different sources according to their spatial location to obtain the target cluster.
[0148] Here, target clusters can be understood as target obstacles. Detection results from different sources are clustered and associated according to their spatial location. Targets with highly overlapping spatial locations are identified as the same target cluster (corresponding to the target obstacles mentioned above).
[0149] Step 507: For each target cluster, dynamically assign confidence to the detection results of each sensor.
[0150] Here, the confidence level is determined by a combination of factors, including the sensor's inherent characteristics, target type matching degree, distance measurement stability, and historical frame tracking continuity. For example, ultrasonic waves have high confidence in near-field static / low-speed targets, millimeter-wave radar has high confidence in measuring the radial velocity of dynamic targets, and fisheye cameras have high confidence in target type identification.
[0151] Step 508: For each target cluster, determine the target cluster category in the detection result with the highest confidence as the target category.
[0152] Step 509 determines whether at least two sensors have detected the same target cluster.
[0153] Here, if at least two sensors detect the same target cluster, proceed to step 510; if no at least two sensors detect the same target cluster, proceed to step 511.
[0154] Step 510: Increase the highest confidence level to obtain the overall confidence level.
[0155] Step 511: Reduce the highest confidence level to obtain the overall confidence level.
[0156] Step 512: For each target cluster, perform a weighted average of the positions of the target cluster in the three detection results to obtain the target position of the target cluster.
[0157] Here, the weight of each detection result is determined by its respective confidence level and distance measurement variance. Ultrasonic and millimeter-wave distance measurements are more accurate, as are visual orientation information.
[0158] Step 513: For each target cluster, fuse the two-dimensional velocity provided by UKF in the detection results output by the ultrasonic radar with the radial velocity in the detection results output by the millimeter-wave radar to obtain the target velocity of that target cluster.
[0159] Step 514: Output the list of fused obstacles (corresponding to the target obstacle list mentioned above).
[0160] Here, the list of fused obstacles includes target cluster category, location, velocity, and overall confidence level.
[0161] It should be noted that if multiple sensors detect obstacles or target clusters at the same spatial location or region, a "weighted confidence voting" or "worst-case" strategy is employed (e.g., for braking decisions, the minimum distance is taken). If one sensor detects an obstacle within its detection range while another does not, a "trust degradation" mechanism is activated, combining sensor health status and historical records for judgment. The final output is a fused obstacle list including location, velocity, size, category, and overall confidence level. This multi-level fusion ensures that, beyond the minimum braking distance D_min, at least two sensors (e.g., near-field: ultrasonic + fisheye; mid-to-far-field: millimeter-wave + fisheye) can cross-validate and locate obstacles, achieving functional redundancy.
[0162] As can be seen from the above, the embodiments of this application adopt a multi-level fusion architecture of "data level-feature level-decision level", which ensures the reliability and accuracy of positioning while taking into account real-time performance.
[0163] For example, taking pedestrian perception as an example, the decision-level fusion logic is illustrated. At time t, 3 meters directly in front of the vehicle, a fisheye camera detects a "pedestrian" bounding box; a millimeter-wave radar clusters a target at this location with a radial velocity of 0.5 m / s (moving away); an ultrasonic UKF tracker has a stable trajectory at this location with a velocity (vx, vy) = (0, 0.5 m / s). At this point, the detection results from different sources are clustered and associated according to their spatial locations. Since the spatial locations of the three are highly overlapping, they are determined to belong to the same target cluster. Based on the sensor's own characteristics, target type matching degree (i.e., matching degree with the above target cluster), distance measurement stability, and historical frame tracking continuity, the confidence level of the detection results of each sensor is dynamically determined. For example, the confidence level of the fisheye camera's detection result is 0.9, the confidence level of the millimeter-wave radar's detection result is 0.8, and the confidence level of the ultrasonic radar's detection result is 0.7. Since the fisheye vision has the highest confidence level, the target category of the obstacle is determined to be "pedestrian". The pedestrian's target location was obtained by weighted averaging of the positions from the three detection results. Additionally, the millimeter-wave radar provided a radial velocity v_radial = 0.5 m / s. The ultrasonic UKF provided a two-dimensional velocity (0, 0.5), with its radial component also at 0.5 m / s, showing good consistency. Fisheye vision estimated the velocity through optical flow or inter-frame displacement as auxiliary verification. The final velocity was obtained by fusing the millimeter-wave and ultrasonic results. Furthermore, all three sensors detected the pedestrian target, meeting the requirement of "at least two sensors," and the overall confidence level was improved to 0.95.
[0164] After obtaining the list of fused obstacles, the decision output layer 340 inputs the list of fused obstacles into the AEB decision module 341. The AEB decision module 341 performs graded braking control on the vehicle based on the list of fused obstacles and the minimum braking distance.
[0165] For example, the list of merged obstacles includes: the location of the merged obstacle (3 meters in front of the vehicle), its speed (lateral 0.5 m / s), its category: pedestrian, and a comprehensive confidence level of 0.95. Inputting this result into the AEB decision module, the system will prepare to issue a risk warning or decelerate, rather than immediately apply emergency braking, because the distance is greater than the minimum braking distance D_min and the obstacle is a moving pedestrian.
[0166] As can be seen from the above, the embodiments of this application achieve redundant and accurate positioning of obstacles outside the minimum braking distance through layered fusion and dynamic optimization, thus ensuring driving safety in low-speed scenarios.
[0167] This application provides a multi-sensor fusion sensing device, referring to... Figure 6 As shown, Figure 6 This is a schematic diagram of a multi-sensor fusion sensing device provided in an embodiment of this application. The multi-sensor fusion sensing device 6 includes: The module 601 is used to obtain the vehicle status during the vehicle's operation and the perception data collected by various types of sensors. The perception data collected by various types of sensors includes: first perception data collected by an ultrasonic radar array, second perception data collected by a millimeter-wave radar, and fisheye images collected by a fisheye camera. The determination module 602 is used to determine the target detection range when detecting obstacles around the vehicle based on the vehicle status; Processing module 603 is used to perform image processing and target detection on the fisheye image to obtain a fisheye depth map and image detection results. The image detection results include at least one of the following: the category, measurement distance, bounding box, and obstacle features of a first obstacle. Within the target detection range, data fusion processing is performed on the first perception data and the fisheye depth map to obtain a near-field occupancy grid map of near-field obstacles for the vehicle. The near-field occupancy grid map includes at least one of the following: the position and outline of near-field obstacles. Within the target detection range, feature fusion processing is performed on the second perception data and the image detection results to obtain a mid-to-far-field obstacle list for the vehicle. The mid-to-far-field obstacle list includes at least one of the following: the position, speed, and category of mid-to-far-field obstacles. The determination module 602 is also used to determine the list of target obstacles around the vehicle based on the first perception data, the second perception data, the image detection results, the near-field occupancy grid map, and the mid-far-field obstacle list; The control module 604 is used to control the braking of the vehicle based on a list of target obstacles.
[0168] This application provides a hardware entity diagram of a vehicle, such as... Figure 7As shown, the hardware entity of the vehicle 7 includes a processor 701 and a memory 702. The memory 702 stores a computer program that can run on the processor 701. When the processor 701 executes the computer program, it implements some or all of the steps in the multi-sensor fusion perception method as described in the above embodiments.
[0169] The memory 702 stores computer programs that can run on the processor. The memory 702 is configured to store instructions and applications that can be executed by the processor 701. It can also cache data to be processed or already processed by the processor 701 and various modules in the vehicle 7 (e.g., image data, audio data, voice communication data and video communication data). It can be implemented by flash memory or random access memory (RAM).
[0170] The processor 701 executes the program to implement the steps of the multi-sensor fusion perception execution method described above. The processor 701 typically controls the overall operation of the vehicle 7.
[0171] This application provides a computer-readable storage medium storing one or more computer programs, which can be executed by one or more processors to implement some or all of the steps in the above-described method. The storage medium can be transient or non-transient.
[0172] This application provides a computer program including computer-readable code, wherein when the computer-readable code is run in a vehicle, a processor in the vehicle executes some or all of the steps in the above-described method.
[0173] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0174] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0175] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.
[0176] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0177] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0178] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0179] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0180] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0181] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0182] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0183] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an in-vehicle terminal (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0184] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A multi-sensor fusion perception method, characterized in that, The method includes: The vehicle status during driving is obtained and the perception data collected by various types of sensors, including: first perception data collected by ultrasonic radar array, second perception data collected by millimeter-wave radar and fisheye images collected by fisheye camera. The target detection range for detecting obstacles around the vehicle is determined based on the vehicle's status. The fisheye image is processed and target detection is performed to obtain a fisheye depth map and image detection results, wherein the image detection results include at least one of the following: the category of the first obstacle, the measured distance, the bounding box, and the obstacle features; Within the target detection range, the first perception data and the fisheye depth map are fused to obtain a near-field occupancy grid map of the near-field obstacles of the vehicle; wherein, the near-field occupancy grid map includes at least one of the following: the position and outline of the near-field obstacles; Within the target detection range, feature fusion processing is performed on the second perception data and the image detection result to obtain a mid-to-far field obstacle list for the vehicle; wherein, the mid-to-far field obstacle list includes at least one of the following: the position, speed, and category of the mid-to-far field obstacle; Based on the first perception data, the second perception data, the image detection results, the near-field occupancy grid map, and the mid-to-far-field obstacle list, a list of target obstacles around the vehicle is determined, and braking control of the vehicle is performed based on the list of target obstacles.
2. The method according to claim 1, characterized in that, Within the target detection range, the first sensing data and the fisheye depth map are fused to obtain a near-field occupancy grid map of near-field obstacles for the vehicle, including: The first sensing data within the target detection range is projected and associated with the fisheye depth map to associate and match the ultrasonic measurement points of the same near-field obstacle in the first sensing data with the pixel regions in the fisheye depth map. Using the initial measurement distance in the first sensing data, the depth estimation of the corresponding near-field obstacle in the fisheye depth map is corrected for near-field error, and the corrected fisheye depth map is obtained. Based on the corrected fisheye depth map and the first sensing data, the near-field occupancy grid map is generated.
3. The method according to claim 1, characterized in that, Within the target detection range, feature fusion processing is performed on the second perception data and the image detection result to obtain a mid-to-far field obstacle list for the vehicle, including: The second perception data within the target detection range is projected onto the fisheye image plane by spatial transformation, and then matched with the bounding box of the first obstacle to obtain the position of the mid-far field obstacle. Using the second sensing data and the obstacle features of the first obstacle, a unified obstacle feature descriptor is constructed; the second sensing data includes at least one of the following: the measured distance, radial velocity, azimuth angle, and radar cross-section of the second obstacle; The obstacle feature descriptors are input into a lightweight fusion neural network or a traditional classifier for joint target recognition and state estimation to obtain the mid-to-far field obstacle list, wherein the mid-to-far field obstacle list includes the speed and category of the mid-to-far field obstacles.
4. The method according to any one of claims 1 to 3, characterized in that, The step of determining the list of target obstacles around the vehicle based on the first sensing data, the second sensing data, the image detection result, the near-field occupancy grid map, and the mid-to-far-field obstacle list includes: The first perception data is processed to obtain the first obstacle processing result around the vehicle, wherein the first obstacle processing result includes the position and estimated speed of the third obstacle; Clustering and target extraction are performed on the second perception data to obtain the second obstacle processing result around the vehicle. The second obstacle processing result includes at least one of the following: the position, speed, trajectory and category of the second obstacle. The second obstacle and the third obstacle are at least partially different. The target obstacle list is determined based on the first obstacle processing result, the second obstacle processing result, the image detection result, the near-field occupancy grid map, and the mid-far-field obstacle list.
5. The method according to claim 4, characterized in that, The step of determining the target obstacle list based on the first obstacle processing result, the second obstacle processing result, the image detection result, the near-field occupancy grid map, and the mid-to-far-field obstacle list includes: The first obstacle processing result, the second obstacle processing result, the image detection result, the near-field occupancy grid map, and the mid-far-field obstacle list are clustered and associated according to their spatial locations to obtain target obstacles in the same spatial location; The first obstacle processing result, the second obstacle processing result, and the image detection result are each configured with a corresponding confidence level; wherein, the confidence level is determined based on the characteristics of each type of sensor, the category matching degree of the target obstacle, the distance measurement stability, and the historical frame tracking continuity; For each target obstacle, select the target obstacle with the highest confidence level from multiple confidence levels, and determine the category of the target obstacle with the highest confidence level as the target category; Based on the multiple confidence levels, determine the position weight of each type of sensor for the target obstacle; Based on multiple position weights, the positions of the target obstacle in the first obstacle processing result, the second obstacle processing result, and the image detection result are weighted and fused to determine the target position of the target obstacle; The estimated velocity and the radial velocity in the second sensing data are weighted and fused to determine the target velocity of the target obstacle; wherein the target obstacle list includes the target category, the target position, the target velocity, and the target confidence level of the target obstacle, and the target confidence level includes the highest confidence level.
6. The method according to claim 5, characterized in that, The step of processing the first perceived data to obtain the first obstacle processing result around the vehicle includes: Based on the initial measurement distance in the first sensing data, determine the initial velocity of each of the third obstacles, and the position of the target ultrasonic radar closest to each of the third obstacles; Based on the position of the target ultrasonic radar and the initial measurement distance of the target ultrasonic radar to the corresponding third obstacle, a system of spherical equations is constructed. The initial positions of each of the third obstacles are obtained by calculating the spherical equations using the nonlinear least squares method. Based on the position and velocity of each of the third obstacles at the previous moment, the initial position and initial velocity of the third obstacles at the current moment are continuously smoothed using the unscented Kalman filter state estimation and tracking algorithm to obtain the smoothed position and velocity of each of the third obstacles at the current moment.
7. The method according to claim 6, characterized in that, Before determining the initial velocity of each of the third obstacles based on the initial measured distance in the first sensing data, the method includes: Obtain historical frame sensing data collected by the ultrasonic radar array; The random sampling consensus algorithm is used to fit and filter the static background in consecutive frames of the historical frame sensing data to obtain static background sensing data. Remove the static background perception data from the first perception data.
8. The method according to claim 6, characterized in that, The step of continuously smoothing the initial position and initial velocity of each of the third obstacles at the current moment using an unscented Kalman filter state estimation and tracking algorithm, based on the position and velocity of each of the third obstacles at the previous moment, to obtain the smoothed position and velocity of each of the third obstacles at the current moment, includes: Obtain historical frame sensing data collected by the ultrasonic radar array; An improved density-based spatial clustering algorithm is used to perform correlation clustering and tracking of spatially adjacent points in consecutive frames of the historical frame sensing data, generating dynamic obstacle clusters. For a cluster of dynamic obstacles, based on the position and velocity of each dynamic obstacle at the previous moment, an unscented Kalman filter state estimation and tracking algorithm is used to continuously smooth the initial position and initial velocity of the dynamic obstacle at the current moment, so as to obtain the smoothed position and velocity of the dynamic obstacle at the current moment; wherein, the third obstacle includes the dynamic obstacle.
9. A multi-sensor fusion sensing device, characterized in that, The device includes: The acquisition module is used to acquire the vehicle status during vehicle operation and the perception data collected by various types of sensors. The perception data collected by various types of sensors includes: first perception data collected by an ultrasonic radar array, second perception data collected by a millimeter-wave radar, and fisheye images collected by a fisheye camera. The determination module is used to determine the target detection range when detecting obstacles around the vehicle based on the vehicle status; A processing module is configured to perform image processing and target detection on the fisheye image to obtain a fisheye depth map and image detection results, wherein the image detection results include at least one of the following: the category, measurement distance, bounding box, and obstacle features of a first obstacle; within the target detection range, data fusion processing is performed on the first perception data and the fisheye depth map to obtain a near-field occupancy grid map of the vehicle's near-field obstacles; wherein the near-field occupancy grid map includes at least one of the following: the position and outline of the near-field obstacle; within the target detection range, feature fusion processing is performed on the second perception data and the image detection results to obtain a mid-to-far-field obstacle list of the vehicle's mid-to-far-field obstacles; wherein the mid-to-far-field obstacle list includes at least one of the following: the position, velocity, and category of the mid-to-far-field obstacle; The determining module is further configured to determine a list of target obstacles around the vehicle based on the first sensing data, the second sensing data, the image detection result, the near-field occupancy grid map, and the mid-far-field obstacle list; The control module is used to control the braking of the vehicle based on the list of target obstacles.
10. A vehicle, characterized in that, The vehicles include: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the method according to any one of claims 1 to 8.