Multi-sensor fusion method and system for environmental perception of autonomous vehicles
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本申请的目的是提供多传感器融合的无人车环境感知方法及系统,用以解决现有技术中存在由于环境变化传感器自身噪声或局部遮挡,导致各模态数据质量动态变化,进一步影响无人车行驶安全性的技术问题
[0015]通过基于自动驾驶运行设计域,确定感知需求指标,在无人车上配置传感器阵列;基于所述传感器阵列进行环境感知,获取多个异构传感器的环境原始数据集,并进行时间同步和空间对齐,得到同一时刻下且转换至统一坐标系的多模态数据;对所述多模态数据进行不确定性估计,得到各模态特征及对应的不确定性度量;融合各模态特征,并以所述不确定性度量作为动态调制系数作用于融合后的各模态特征,将经不确定性加权的融合特征解码作为感知输出,同时输出感知结果及对应的综合不确定性置信度;基于所述感知结果及对应的综合不确定性置信度,执行分层递进的路径规划策略。也就是说,通过对多模态数据进行不确定性估计并作为动态调制系数进行加权融合,同时输出感知结果与综合不确定性置信度,并基于该置信度执行分层递进路径规划,自适应抑制不可靠模态的贡献,提升恶劣环境下感知准确性,从而提高无人车行驶安全性。
Smart Images

Figure CN122561053A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving environmental perception technology, specifically to an unmanned vehicle environmental perception method and system based on multi-sensor fusion. Background Technology
[0002] Currently, autonomous vehicles primarily rely on sensors to acquire information about their surroundings, and then use this information for obstacle detection, lane recognition, and drivable area delineation. However, real-world driving environments are complex and variable. Factors such as rain, snow, fog, haze, backlighting, sudden changes in illumination at tunnel entrances and exits, sensor noise, and partial obstruction from vehicles ahead can all cause drastic and unpredictable degradation in sensor data quality. When a sensor experiences significant noise or data loss due to interference, it is highly susceptible to missed detections, false detections, or location drift in adverse weather conditions or when sensors experience temporary malfunctions, seriously threatening the driving safety of autonomous vehicles.
[0003] In summary, existing technologies suffer from technical problems where changes in the environment, sensor noise, or partial occlusion cause dynamic changes in the quality of data across different modes, further impacting the safety of autonomous vehicles. Summary of the Invention
[0004] The purpose of this application is to provide a multi-sensor fusion method and system for environmental perception of unmanned vehicles, in order to solve the technical problem in the prior art that the quality of data of each modality changes dynamically due to sensor noise or partial occlusion caused by environmental changes, which further affects the driving safety of unmanned vehicles.
[0005] To achieve the above objectives, this application provides a method and system for environmental perception of unmanned vehicles using multi-sensor fusion.
[0006] Firstly, this application provides a multi-sensor fusion-based environmental perception method for autonomous vehicles. This method is implemented through a multi-sensor fusion-based environmental perception system for autonomous vehicles. The method includes: determining perception requirement indicators based on an autonomous driving operation design domain, and configuring a sensor array on the autonomous vehicle; performing environmental perception based on the sensor array, acquiring raw environmental datasets from multiple heterogeneous sensors, and performing time synchronization and spatial alignment to obtain multimodal data at the same time and transformed to a unified coordinate system; performing uncertainty estimation on the multimodal data to obtain each modal feature and its corresponding uncertainty metric; fusing the modal features and using the uncertainty metric as a dynamic modulation coefficient applied to the fused modal features; decoding the uncertainty-weighted fused features as the perception output, and simultaneously outputting the perception result and its corresponding comprehensive uncertainty confidence level; and executing a hierarchical and progressive path planning strategy based on the perception result and its corresponding comprehensive uncertainty confidence level.
[0007] Optionally, feature extraction is performed on the multimodal data to obtain each modal feature; for each modal feature, the feature vector difference between each feature point and other feature points in its neighborhood is calculated, and the average value of the feature vector difference is used as a measure of random uncertainty; for the same modal feature, the feature point at the current time is matched with the feature point at the corresponding position at a historical time, and the cumulative change of the matching residual within the time window is calculated as a measure of cognitive uncertainty; the random uncertainty measure and the cognitive uncertainty measure of each modal feature are weighted and fused to obtain the comprehensive uncertainty measure corresponding to each modal feature.
[0008] Optionally, after aligning the modal features in the spatial dimension, the feature vectors at each spatial location are concatenated to form a joint feature tensor; the reciprocal of the uncertainty measure of each modality at the same spatial location is normalized to obtain the weight coefficient of each modality at the same spatial location, which is used as the dynamic modulation coefficient; the joint feature tensor and the dynamic modulation coefficient are weighted and summed to obtain the uncertainty-weighted fusion feature; the fusion feature is decoded to output the perception result, and the comprehensive uncertainty confidence obtained by weighting and summing the comprehensive uncertainty measures corresponding to each modal feature is also output.
[0009] Optionally, based on the static perception results in the perception results, a drivable area and static obstacles are obtained, and a global reference path from the current position to the target position is constructed; based on the dynamic perception results in the perception results, a baseline safety distance is determined, and the baseline safety distance is dynamically adjusted using the comprehensive uncertainty confidence level to obtain a dynamic safety distance; according to the global reference path and the dynamic safety distance, a constrained optimization method is used for local trajectory planning, and a path planning strategy is output.
[0010] Optionally, an optimization objective function is constructed, including the deviation cost, control smoothing cost, and collision cost of the global reference path; constraints are set, including vehicle kinematic constraints, drivable area boundary constraints, and obstacle avoidance hard constraints; the path state space is initialized based on the global reference path; numerical iterative optimization is performed in the path state space according to the optimization objective function and the constraints until convergence is obtained, resulting in a local trajectory that satisfies the constraints and minimizes the cost, and the path planning strategy is output.
[0011] Optionally, the uncertainty metric is used to construct an adaptive spatial alignment constraint, and the extrinsic parameters of the sensor array are corrected in real time online to obtain the corrected sensor extrinsic parameters. The spatially aligned multimodal data is then updated to form a sensing closed loop.
[0012] Optionally, when the uncertainty metric corresponding to a certain mode exceeds a preset threshold, it is determined that there is a deviation in the current spatial alignment, and the projection error of the high-confidence perception result among the sensor data is extracted; if the projection error deviates significantly from the expected error, it is determined that there is a shift in the current extrinsic parameters, and an adaptive spatial alignment constraint for online updating of extrinsic parameters is constructed with the optimization objective of minimizing the projection consistency error of the high-confidence perception target among the sensor data.
[0013] Secondly, this application also provides a multi-sensor fusion unmanned vehicle environment perception system for executing the multi-sensor fusion unmanned vehicle environment perception method as described in the first aspect. The multi-sensor fusion unmanned vehicle environment perception system includes: a sensor configuration module for determining perception requirement indicators based on the autonomous driving operation design domain and configuring a sensor array on the unmanned vehicle; a spatiotemporal synchronization module for performing environmental perception based on the sensor array, acquiring raw environmental datasets from multiple heterogeneous sensors, and performing time synchronization and spatial alignment to obtain multimodal data at the same time and transformed to a unified coordinate system; an uncertainty estimation module for performing uncertainty estimation on the multimodal data to obtain each modal feature and its corresponding uncertainty metric; a feature fusion decoding module for fusing each modal feature and using the uncertainty metric as a dynamic modulation coefficient applied to each fused modal feature, decoding the uncertainty-weighted fused features as the perception output, and simultaneously outputting the perception result and its corresponding comprehensive uncertainty confidence level; and a path planning module for executing a hierarchical and progressive path planning strategy based on the perception result and its corresponding comprehensive uncertainty confidence level.
[0014] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0015] By defining perception requirements based on the autonomous driving operation design domain, a sensor array is configured on the autonomous vehicle. Environmental perception is performed using this sensor array, acquiring raw environmental datasets from multiple heterogeneous sensors. These datasets are then synchronized in time and aligned spatially to obtain multimodal data at the same time and transformed to a unified coordinate system. Uncertainty estimation is performed on this multimodal data to obtain modal features and their corresponding uncertainty metrics. The modal features are then fused, and the uncertainty metrics are used as dynamic modulation coefficients applied to the fused modal features. The uncertainty-weighted fused features are decoded as the perception output, along with the perception result and its corresponding comprehensive uncertainty confidence level. Based on the perception result and the corresponding comprehensive uncertainty confidence level, a hierarchical progressive path planning strategy is executed. In other words, by performing uncertainty estimation on multimodal data and using it as dynamic modulation coefficients for weighted fusion, while simultaneously outputting the perception result and comprehensive uncertainty confidence level, and executing hierarchical progressive path planning based on this confidence level, the contribution of unreliable modes is adaptively suppressed, improving perception accuracy in harsh environments and thus enhancing the driving safety of the autonomous vehicle.
[0016] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the multi-sensor fusion method for environmental perception of unmanned vehicles proposed in this application.
[0019] Figure 2 This is a schematic diagram of the structure of the multi-sensor fusion unmanned vehicle environmental perception system of this application.
[0020] Figure labeling: Sensor configuration module 11, spatiotemporal synchronization module 12, uncertainty estimation module 13, feature fusion decoding module 14, path planning module 15. Detailed Implementation
[0021] This application provides a multi-sensor fusion-based environmental perception method and system for autonomous vehicles, addressing the technical problem in existing technologies where sensor noise or partial occlusion due to environmental changes cause dynamic variations in the quality of data across different modalities, further impacting the driving safety of autonomous vehicles. By estimating the uncertainty of multi-modal data and using it as dynamic modulation coefficients for weighted fusion, the method outputs both the perception result and the overall uncertainty confidence level. Based on this confidence level, hierarchical progressive path planning is performed, adaptively suppressing the contribution of unreliable modalities, improving perception accuracy in harsh environments, and thus enhancing the driving safety of autonomous vehicles.
[0022] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0023] Example 1, please refer to the appendix. Figure 1 This application provides a multi-sensor fusion method for autonomous vehicle environmental perception, wherein the multi-sensor fusion method for autonomous vehicle environmental perception is applied to a multi-sensor fusion method for autonomous vehicle environmental perception, and the multi-sensor fusion method for autonomous vehicle environmental perception specifically includes the following steps:
[0024] S100: Based on the autonomous driving operation design domain, determine the perception requirement indicators and configure sensor arrays on the driverless vehicle.
[0025] Specifically, based on the target application scenarios of autonomous driving, the specific boundary conditions of the operational design domain are defined, such as limiting operation to structured urban roads, good lighting, and no rain or snow, or extending to nighttime and light rain / fog environments. For the autonomous driving operational design domain, the perception requirements of the unmanned vehicle under different operating conditions are analyzed, and the emergency braking distance at the maximum driving speed is calculated to determine the minimum distance for forward long-range detection; the reaction time of lateral vehicles at intersections is considered to determine the lateral detection range; a system delay upper limit is set according to functional safety requirements; and the performance degradation of each sensor under adverse weather conditions is assessed to determine redundancy requirements. The autonomous driving operational design domain is the specific set of conditions under which the autonomous driving system is designed to operate safely, including road type, geographical area, traffic flow density, and whether high-precision maps are permitted. The autonomous driving operational design domain defines the boundaries within which the autonomous driving system can function normally; beyond these boundaries, the system should request manual intervention or implement the least-risk strategy.
[0026] Perception requirements are quantitative performance requirements for the environmental perception system of autonomous vehicles, based on the autonomous driving operation design domain. These requirements include forward detection range, lateral detection range, target detection accuracy, speed measurement error, detection latency, and performance retention under rain, fog, and low-light conditions at night. These requirements determine the minimum requirements for sensor selection and configuration. For example, if a forward detection range of 150m and all-weather capability are required, a 128-line main LiDAR with a detection range of no less than 200m and a forward-facing 4D millimeter-wave radar are needed. If 360° coverage is required, blind spot LiDARs and corner radars need to be deployed on the roof, four corners, and front and rear bumpers. If the detection of traffic signs and traffic lights is required, a high-resolution forward-looking camera and a surround-view camera are needed.
[0027] All selected sensors are installed on the autonomous vehicle in appropriate locations, ensuring that the overlapping areas of the fields of view between the sensors meet the fusion requirements. Interfaces for subsequent calibration and verification are reserved to form a complete sensor array. The model, installation location, electrical interface, and data output format of each sensor are clearly defined. By defining the operational design domain and determining the quantitative perception requirements accordingly, blind and excessive sensor configuration is avoided, thereby controlling hardware costs while meeting safety requirements. Within the specified autonomous driving operational design domain, the detection distance, accuracy, latency, and other performance indicators of each sensor meet preset thresholds, thus supporting the safe operation of the autonomous driving system within its design range.
[0028] S200: Based on the sensor array, perform environmental perception, acquire the raw environmental datasets of multiple heterogeneous sensors, and perform time synchronization and spatial alignment to obtain multimodal data at the same time and transformed to a unified coordinate system.
[0029] Specifically, after the autonomous vehicle is powered on, each sensor continuously collects raw environmental data according to its own operating frequency. The raw environmental dataset consists of unprocessed, low-level data directly collected by each sensor.
[0030] A unified time base is established through hardware synchronization. Upon receiving satellite signals, a precise pulse-per-second signal is output and distributed via a timing board to the clock modules of all sensors via a wired connection. Simultaneously, the receiver also outputs an information message containing the current absolute time. At the start of acquiring a frame of data, each sensor latches the rising edge of the pulse-per-second signal and appends its absolute timestamp to that frame. For sensors with different frame rates—for example, a lidar operating at 10Hz and a camera operating at 30Hz—after receiving data from each sensor, the trigger time of the lidar is used as the reference. For the camera's image data, linear interpolation is performed using the timestamps of adjacent frames to synthesize a virtual image frame aligned with the lidar's time.
[0031] After time synchronization is complete, the raw data is sent to the spatial alignment module, which reads the pre-calibrated extrinsic parameter matrix, including the rotation and translation parameters of the LiDAR to the vehicle coordinate system, the rotation and translation parameters of each camera relative to the LiDAR, and the extrinsic parameters of the millimeter-wave radar. For each point in the LiDAR point cloud, its 3D coordinates are multiplied left by the rotation matrix and then a translation vector is added to transform it to the vehicle coordinate system. For camera images, since the images themselves are 2D, distortion correction is performed first using calibrated camera intrinsic parameters, such as focal length, distortion coefficients, and optical center coordinates. Then, the extrinsic parameter matrix is used to establish a mapping relationship between pixel coordinates and 3D spatial points in the vehicle coordinate system, usually assisted by depth information, such as the distance from the corresponding point from the LiDAR. For target points of the millimeter-wave radar, the distance and angle in the polar coordinate system are also converted to planar coordinates in the vehicle coordinate system using the extrinsic parameter matrix. All data after spatiotemporal synchronization is packaged into a single data structure at the same time, containing the LiDAR point cloud array, the image matrix of multiple cameras, the target list of the millimeter-wave radar, and a timestamp.
[0032] By synchronizing time and aligning space, the spatiotemporal differences between different sensors are eliminated, so that the same physical target has consistent spatiotemporal coordinates in lidar point clouds, camera images and millimeter-wave radar data. Time synchronization ensures that moving objects will not produce false displacements due to time misalignment; spatial alignment ensures that information from different modes can be correlated and calculated point by point or region by region.
[0033] S300: Perform uncertainty estimation on the multimodal data to obtain the modal features and corresponding uncertainty measures.
[0034] Furthermore, S300 of this application includes: extracting features from the multimodal data to obtain each modal feature; for each modal feature, calculating the feature vector difference between each feature point and other feature points in its neighborhood, and using the average value of the feature vector difference as a measure of random uncertainty; for the same modal feature, performing temporal matching between the feature point at the current time and the feature point at the corresponding position at a historical time, calculating the cumulative change of the matching residual within the time window as a measure of cognitive uncertainty; and weightedly fusing the measures of random uncertainty and cognitive uncertainty of each modal feature to obtain a comprehensive uncertainty measure corresponding to each modal feature.
[0035] Specifically, multimodal data from the same time point and in a unified coordinate system are used, including LiDAR point clouds, camera images, and millimeter-wave radar target points. For each modality, a corresponding feature extraction network is invoked. For LiDAR point clouds, a voxel-based sparse convolutional network is used to divide the point cloud into a 0.2m cube grid, extracting a 64-dimensional feature vector for each non-empty grid. For camera images, multi-scale feature maps are extracted and then upsampled to the same spatial resolution as the point cloud using bilinear interpolation, with each pixel location corresponding to a 64-dimensional feature vector. For millimeter-wave radar target points, due to their sparse nature, a 32-dimensional feature vector composed of a position code and a velocity code is directly assigned to each target point. Features from all modalities are projected onto a unified bird's-eye view grid, with a grid size of 200 rows by 200 columns, each cell corresponding to a 0.4m x 0.4m area in actual space. Each modality obtains a feature map, with each feature point corresponding to a grid cell. Feature extraction extracts vectorized representations with semantic or geometric information from the multimodal data. Each modality of data is processed by feature extraction to form a feature map. The feature map consists of multiple feature points. Each feature point corresponds to a spatial location in the original data, such as a pixel in an image or a voxel in a point cloud, and contains a high-dimensional feature vector that describes the local structure, texture, or geometric properties of that location.
[0036] For image modalities, feature points typically correspond to a pixel region in the original image; for LiDAR point clouds, feature points correspond to the center of a small cube in space.
[0037] For each modality's feature map, a random uncertainty metric is calculated. For each feature point in the feature map, the system searches for other feature points within a spatial neighborhood radius of 0.8m (two grid units), centered on that point. The Euclidean distance between the feature vector of the current point and the feature vectors of every point in the neighborhood is calculated, and then the arithmetic mean of these distances is taken to obtain the random uncertainty metric for that feature point. If the surrounding environment around a feature point has abrupt changes in texture, such as vehicle edges or road curbs, its features differ significantly from those of its neighbors, resulting in higher random uncertainty. Conversely, on flat roads or in sky areas, the features of the neighbors are similar, leading to lower random uncertainty. The random uncertainty metric reflects the inherent noise or observational uncertainty of the data itself, and is quantified by calculating the average difference between the feature vectors of each feature point and other feature points in its spatial neighborhood. If a feature point differs greatly from its neighbors, it indicates that the point may be noise, an edge, or an occluded area, resulting in high random uncertainty; conversely, if the features are uniform and consistent, the uncertainty is low.
[0038] Cognitive uncertainty is measured for feature maps of the same modality. A fixed-length time window is preset, such as including the feature map history cache of the most recent 5 frames. For each feature point in the current frame, the coordinates of feature points at the same spatial grid position in previous frames are transformed to the coordinate system of the current frame using the vehicle motion measured by the IMU. The Euclidean distance between the feature vector of the current frame and the feature vector of the corresponding position in the historical frames is calculated, obtaining the matching residual for each historical frame. The residual values of these frames are grouped into a sequence, and the variance of the sequence is calculated as the cognitive uncertainty measure for that feature point. If the feature at a certain position changes drastically in the historical frames, such as the limbs of a pedestrian swinging or leaves swaying, the variance will be large, and the cognitive uncertainty will be high; if it is a stationary building, the variance will be small, and the cognitive uncertainty will be low.
[0039] The random uncertainty measure and the cognitive uncertainty measure for each feature point are weighted and fused. Two weight coefficients are preset, for example, random uncertainty weight is 0.6 and cognitive uncertainty weight is 0.4. For each feature point, these two weights are multiplied by its random uncertainty and cognitive uncertainty values respectively, and then the products are added together to obtain the comprehensive uncertainty measure for that feature point.
[0040] Random uncertainty measures can capture data unreliability caused by sensor noise and local edges and occlusions, while cognitive uncertainty measures can identify dynamic scenes or regions where the model has not been trained. The weighted fusion of the two measures allows the comprehensive uncertainty measure to reflect both the quality of the current observations and temporal stability, automatically reducing the negative impact of low-quality regions on the fusion results while retaining the accurate information of high-quality regions, thereby improving the robustness and accuracy of the overall perception system.
[0041] S400: Fuse the features of each modality, and use the uncertainty measure as a dynamic modulation coefficient to apply to each fused modality feature. Decode the fused features weighted by uncertainty as the perception output, and output the perception result and the corresponding comprehensive uncertainty confidence.
[0042] Furthermore, S400 of this application includes: aligning the modal features in spatial dimension, concatenating the feature vectors at each spatial location to form a joint feature tensor; normalizing the inverse of the uncertainty measure of each modality at the same spatial location to obtain the weight coefficient of each modality at the same spatial location, which is used as the dynamic modulation coefficient; performing a weighted summation of the joint feature tensor and the dynamic modulation coefficient to obtain an uncertainty-weighted fusion feature; decoding based on the fusion feature to output the perception result, and simultaneously outputting the comprehensive uncertainty confidence obtained by weighted summation of the comprehensive uncertainty measures corresponding to each modal feature.
[0043] Specifically, feature maps from different modalities are transformed and interpolated to achieve the same spatial resolution and origin within a unified bird's-eye view or perspective grid. Each feature point in each modal feature map corresponds to the same location in physical space. For the same spatial location, feature vectors from different modalities at that location are concatenated sequentially to form a longer joint vector, preserving all original feature information from each modality. The joint feature tensor is a three-dimensional array composed of the concatenated feature vectors from all spatial locations arranged in space, with dimensions of (number of height grids) × (number of width grids) × (sum of the dimensions of each modality's features).
[0044] For each grid location, the feature vector for that location is extracted from the feature maps of each modality, and these vectors are concatenated sequentially to form a 160-dimensional joint feature vector. Simultaneously, the comprehensive uncertainty metric for that location is extracted from the uncertainty metric maps of each modality. To prevent the denominator from being zero, a very small positive number, such as 0.001, is added to each of these uncertainty values, and then their reciprocals are calculated. These reciprocals are then summed to obtain a total. Each reciprocal is divided by the total to obtain the corresponding weight coefficient, and the sum of the weights is 1. These weight coefficients are the dynamic modulation coefficients, which assign confidence levels among different modalities at the same grid location, with modes having lower uncertainty receiving higher weights.
[0045] The joint feature vector and dynamic modulation coefficients are weighted and summed. Each modal component in the joint feature vector is multiplied by its corresponding weight, and these three weighted vectors are reassembled in their original order to form a new vector, which is the uncertainty-weighted fusion feature of that spatial location.
[0046] The fused feature tensor is input into a detection-decoding network. This network consists of three convolutional layers and a fully connected detection head. Each of the three convolutional layers has a 3×3 kernel size and a stride of 1, employing batch normalization and ReLU activation to extract higher-level semantic information. The refined feature map size remains unchanged, but the number of channels may be reduced to 64 dimensions. Two detection heads are input in parallel: one for target center probability prediction and the other for target attribute regression. The target center probability prediction head uses a 1×1 convolution, outputting a 200×200×1 heatmap, where each grid value represents the probability of a target center existing at that grid location. A probability threshold, such as 0.5, is set to filter all grids with probabilities higher than this threshold as candidate centers. For each candidate center, the target attribute regression head also uses a 1×1 convolution, outputting the target attribute vector corresponding to that grid location, including the center coordinate offset relative to the grid, target size, heading angle, velocity vector, and classification score for each category. By removing duplicate detections using a nonmaximum suppression algorithm, several perception results are obtained. Each result includes a category label, 3D position, size, heading angle, velocity, and classification confidence.
[0047] While outputting each sensing result, for each detected target, its comprehensive uncertainty confidence score is calculated. The comprehensive uncertainty measure of each modality corresponding to the grid position of the target is extracted, and a weighted average is performed using preset modal fusion weights to obtain the comprehensive uncertainty confidence score of the target. The comprehensive uncertainty confidence score, along with the target's position, size, category, and other information, is ultimately used as the sensing output.
[0048] By combining spatially aligned feature stitching with uncertainty-driven dynamic weighted fusion, adaptive integration of multimodal information is achieved. The dynamic modulation coefficients can automatically suppress the contribution of low-quality modes while amplifying the influence of high-quality modes based on the real-time uncertainty of each spatial location, thereby improving the robustness and accuracy of the fused features.
[0049] S500: Based on the perception results and the corresponding comprehensive uncertainty confidence level, execute a hierarchical and progressive path planning strategy.
[0050] Furthermore, S500 of this application includes: acquiring a drivable area and static obstacles based on the static perception results in the perception results, and constructing a global reference path from the current position to the target position; determining a baseline safety distance based on the dynamic perception results in the perception results, dynamically adjusting the baseline safety distance using the comprehensive uncertainty confidence level to obtain a dynamic safety distance; and performing local trajectory planning using a constrained optimization method according to the global reference path and the dynamic safety distance, and outputting a path planning strategy.
[0051] Furthermore, this application also includes the following steps: constructing an optimization objective function, including the deviation cost, control smoothing cost, and collision cost of the global reference path; setting constraints, including vehicle kinematic constraints, drivable area boundary constraints, and obstacle avoidance hard constraints; initializing the path state space based on the global reference path; performing numerical iterative optimization in the path state space according to the optimization objective function and the constraints until convergence, obtaining the local trajectory that satisfies the constraints and minimizes the cost, and outputting the path planning strategy.
[0052] Specifically, the perception results include static perception results and dynamic perception results. Static perception results extract the polygon of the drivable area and the outlines of static obstacles. Static perception results are environmental information that does not change rapidly over time, mainly including the polygon of the drivable area and the position and outline of static obstacles such as cones, parked vehicles, curbs, construction fences, etc.
[0053] The local space surrounding the vehicle is discretized into a grid map. Each grid cell is assigned a non-negative movement cost based on whether it is within a drivable area, whether it is occupied by a static obstacle, and its distance from the nearest obstacle. Starting from the grid cell corresponding to the vehicle's current position and ending at the grid cell corresponding to the global navigation target point, a heuristic search algorithm is used to search for the grid sequence with the minimum cumulative movement cost on the cost map. The resulting discrete grid sequence is then smoothed and filtered to eliminate sharp angles, generating a continuous global reference path composed of equally spaced reference points.
[0054] The dynamic perception result is target information that changes over time, including the instantaneous position, velocity, acceleration, size, category, and overall uncertainty confidence level of moving objects such as vehicles, pedestrians, and cyclists. For each dynamic target, the corresponding baseline safety distance value is read from a pre-defined baseline safety distance table based on its category, and the overall uncertainty confidence level output from the preceding steps for that target is obtained. A maximum safety distance magnification factor is set, and the dynamic safety distance is calculated as: Dynamic Safety Distance = Baseline Safety Distance × (1 + Maximum Magnification Factor) × (1 - Overall Uncertainty Confidence Level). The calculation result is then limited to between the baseline safety distance and the maximum safety distance to ensure that the dynamic safety distance is neither less than the baseline value nor too large.
[0055] An optimization objective function for local trajectory planning is constructed. The planning time domain is defined and discretized into multiple time points. The objective function comprises a weighted sum of three costs: the first is the deviation cost, calculated by multiplying the sum of squares of the lateral and heading deviations between the vehicle's position and the corresponding reference point on the global reference path at each time point by a preset weight; the second is the smoothing cost, calculated by multiplying the sum of squares of the rate of change of acceleration and the rate of change of steering angular velocity between adjacent time points by a preset weight; the third is the collision cost, calculated for each dynamic target by multiplying the reciprocal of the square of the distance between the vehicle and the target by a weight factor negatively correlated with the target's overall uncertainty confidence level (e.g., this factor equals 2 minus the overall uncertainty confidence level), summed over all targets and all time points, and finally multiplied by a preset collision coefficient.
[0056] Simultaneously, constraints that the optimization problem must satisfy are set. Kinematic constraints include upper and lower limits for acceleration, upper and lower limits for the rate of change of acceleration, upper and lower limits for steering angle, upper and lower limits for steering angular velocity, and upper and lower limits for velocity. The drivable area boundary constraint requires that the vehicle's position coordinates at each moment be located inside the drivable area polygon, which is achieved by calculating the signed distance from the point to the polygon boundary and constraining this distance to be non-negative. The obstacle avoidance hard constraint requires that the Euclidean distance between the vehicle and each dynamic target at each moment is strictly greater than the dynamic safety distance corresponding to that target.
[0057] The reference points on the global reference path are truncated according to the length of the planning time domain, resulting in a series of expected positions and heading angles from the current moment to the end of the future planning time domain. Based on the current actual state of the vehicle, including its current position, velocity, acceleration, and heading angle, spatiotemporal registration is performed on the truncated global reference path to generate an initial state sequence. Specifically, the current state of the vehicle is used as the initial state. For each discrete moment, it is assumed that the vehicle is traveling strictly along the global reference path at a constant current speed, and that the lateral position deviation is zero and the heading angle is equal to the heading angle of the reference point. Thus, the position, velocity, acceleration, and heading angle at each moment are calculated. This yields a complete initial trajectory, which lies on a curve in the path state space, serving as the starting point for numerical iterative optimization.
[0058] Set a maximum number of iterations and a convergence threshold. In each iteration, calculate the objective function value corresponding to the current trajectory, as well as the first and second derivatives of the objective function with respect to each state variable. For constraints, convert inequality constraints into penalty terms to ensure that constraints are gradually satisfied during optimization. Solve a quadratic programming subproblem to obtain the update direction of the state variables. Determine an appropriate step size through linear search so that the new trajectory updated along this direction can effectively reduce the objective function value without severely violating the constraints. After the update is complete, check the convergence condition. If the absolute value of the difference between the objective function value of the current trajectory and the previous iteration is less than a preset threshold, or the norm of the change in the state variables is less than a threshold, or the maximum number of iterations is reached, then stop the iteration; otherwise, continue to the next iteration.
[0059] After iteration, a trajectory that satisfies all constraints and minimizes the objective function value is obtained, i.e., a local trajectory. This trajectory contains the vehicle's state at each discrete moment in the planning time domain, such as position coordinates, velocity, acceleration, heading angle, and curvature. This local trajectory is extracted from the path state space, transformed into a format usable by the control module, and sent as the final output path planning strategy to the vehicle's underlying control actuators, such as the steering, throttle, and braking systems.
[0060] By initializing the path state space and employing numerical iterative optimization, the abstract global reference path and constraints are transformed into a feasible and optimal local trajectory. Based on the global reference path and the current vehicle state, a reasonable starting point for iteration is provided, accelerating the optimization convergence speed. The numerical iterative optimization algorithm can handle nonlinear objective functions and complex constraints, ensuring that the final trajectory meets kinematic limits, boundary safety, and dynamic obstacle avoidance requirements.
[0061] Furthermore, this application also includes the following steps: constructing an adaptive spatial alignment constraint using the uncertainty metric, performing real-time online correction on the extrinsic parameters of the sensor array to obtain the corrected sensor extrinsic parameters, updating the spatially aligned multimodal data, and forming a sensing closed loop.
[0062] Furthermore, this application also includes the following steps: when the uncertainty metric corresponding to a certain mode exceeds a preset threshold, it is determined that there is a deviation in the current spatial alignment, and the projection error of the high-confidence perception result among the sensor data is extracted; if the projection error deviates significantly from the expected error, it is determined that there is a shift in the current extrinsic parameters, and an adaptive spatial alignment constraint for online updating of extrinsic parameters is constructed with the optimization objective of minimizing the projection consistency error of the high-confidence perception target among the sensor data.
[0063] Specifically, all modalities are traversed to check if any modality's uncertainty metric exceeds a pre-set threshold in a specific spatial region. If it does, a preliminary determination is made that the region may have a spatial alignment deviation. High-confidence perception results with a comprehensive uncertainty confidence level below the lower threshold are extracted from the perception output. For each high-confidence target, its 3D position in the LiDAR coordinate system, the center pixel coordinates of the detection box in the camera image, and its position in the millimeter-wave radar coordinate system are obtained. Using the current extrinsic parameter matrix, the LiDAR points are projected onto the camera image plane, and the Euclidean distance between the projected points and the image detection center is calculated to obtain the LiDAR-camera projection error; similarly, the millimeter-wave radar target points are projected onto the camera image to obtain the radar-camera projection error. These projection errors are compared with the expected error obtained from pre-offline calibration. If the difference between a projection error and the expected error is greater than a preset tolerance range, it is determined that the extrinsic parameters have indeed shifted and need to be corrected online.
[0064] An adaptive spatial alignment constraint is constructed. Taking all current high-confidence targets as samples, a loss function is defined with respect to extrinsic variables. This loss function is the sum of squared projection consistency errors of each target across each sensor pair. Specifically, for each high-confidence target, the squared reprojection errors between the LiDAR and camera, the squared reprojection errors between the millimeter-wave radar and camera, and optionally the squared coordinate differences between the LiDAR and millimeter-wave radar are calculated, and then summed. Minimizing this loss function is used as the optimization objective, while constraining the range of variation of the extrinsic variables.
[0065] High-confidence targets with a comprehensive uncertainty confidence level below the lower threshold are extracted from the perception results of the current frame. For each high-confidence target, its 3D coordinates in the LiDAR coordinate system, the center pixel coordinates of the detection box in the camera image, and its position in the millimeter-wave radar coordinate system are obtained. Using the current extrinsic parameter matrix, the LiDAR coordinate points are projected onto the camera image plane, and the Euclidean distance between the projected points and the actual image detection center is calculated to obtain the LiDAR-camera projection error; the millimeter-wave radar-camera projection error is calculated similarly. These projection errors are compared with the expected error calculated offline beforehand. If the projection error deviates significantly from the expected error, it is determined that the current extrinsic parameters have drifted.
[0066] Using all high-confidence targets as samples, a loss function is constructed regarding the extrinsic parameters. This loss function is defined as the sum of squared reprojection errors of all targets across sensor pairs. Minimizing this loss function is the optimization objective, while a regularization term is added to prevent excessive changes in the extrinsic parameters, forming an adaptive spatial alignment constraint. A nonlinear least squares optimization algorithm is used to iteratively solve the problem with the current extrinsic parameters as initial values, obtaining the extrinsic parameter correction amount that minimizes the loss function. After iterative convergence, the correction amount is superimposed on the current extrinsic parameters to obtain the corrected sensor extrinsic parameters. Spatial alignment is then re-performed using the corrected extrinsic parameters, transforming all original sensor data to a unified vehicle coordinate system according to the new extrinsic parameters, generating updated multimodal data. This updated multimodal data will be used for uncertainty estimation and fusion perception in the next frame, thus forming a closed loop in the entire system. As the perception quality improves, the number of high-confidence targets increases, further improving calibration accuracy.
[0067] By using adaptive spatial alignment constraints, high-confidence targets that naturally appear in the perception results are used as calibration benchmarks without relying on external calibration references, and external parameter drift is detected and compensated in real time.
[0068] In summary, the multi-sensor fusion-based environmental perception method for autonomous vehicles provided in this application has the following technical advantages:
[0069] By defining perception requirements based on the autonomous driving operation design domain, a sensor array is configured on the autonomous vehicle. Environmental perception is performed using this sensor array, acquiring raw environmental datasets from multiple heterogeneous sensors. These datasets are then synchronized in time and aligned spatially to obtain multimodal data at the same time and transformed to a unified coordinate system. Uncertainty estimation is performed on this multimodal data to obtain modal features and their corresponding uncertainty metrics. The modal features are then fused, and the uncertainty metrics are used as dynamic modulation coefficients applied to the fused modal features. The uncertainty-weighted fused features are decoded as the perception output, along with the perception result and its corresponding comprehensive uncertainty confidence level. Based on the perception result and the corresponding comprehensive uncertainty confidence level, a hierarchical progressive path planning strategy is executed. In other words, by performing uncertainty estimation on multimodal data and using it as dynamic modulation coefficients for weighted fusion, while simultaneously outputting the perception result and comprehensive uncertainty confidence level, and executing hierarchical progressive path planning based on this confidence level, the contribution of unreliable modes is adaptively suppressed, improving perception accuracy in harsh environments and thus enhancing the driving safety of the autonomous vehicle.
[0070] Example 2: Based on the same inventive concept as the multi-sensor fusion unmanned vehicle environmental perception method in Example 1, this application also provides a multi-sensor fusion unmanned vehicle environmental perception system. Please refer to the appendix. Figure 2 The multi-sensor fusion-based autonomous vehicle environmental perception system includes:
[0071] The sensor configuration module 11 is used to determine the perception requirement indicators based on the autonomous driving operation design domain and configure a sensor array on the unmanned vehicle; the spatiotemporal synchronization module 12 is used to perform environmental perception based on the sensor array, acquire the raw environmental datasets of multiple heterogeneous sensors, and perform time synchronization and spatial alignment to obtain multimodal data at the same time and transformed to a unified coordinate system; the uncertainty estimation module 13 is used to perform uncertainty estimation on the multimodal data to obtain each modal feature and the corresponding uncertainty metric; the feature fusion decoding module 14 is used to fuse each modal feature and use the uncertainty metric as a dynamic modulation coefficient to act on each fused modal feature, decode the uncertainty-weighted fused feature as the perception output, and output the perception result and the corresponding comprehensive uncertainty confidence level; the path planning module 15 is used to execute a hierarchical and progressive path planning strategy based on the perception result and the corresponding comprehensive uncertainty confidence level.
[0072] Furthermore, the uncertainty estimation module 13 in the multi-sensor fusion unmanned vehicle environmental perception system is also used for: extracting features from the multimodal data to obtain each modal feature; for each modal feature, calculating the feature vector difference between each feature point and other feature points in its neighborhood, and using the average value of the feature vector difference as a measure of random uncertainty; for the same modal feature, performing temporal matching between the feature point at the current moment and the feature point at the corresponding position at a historical moment, calculating the cumulative change of the matching residual within the time window as a measure of cognitive uncertainty; and weightedly fusing the measures of random uncertainty and cognitive uncertainty of each modal feature to obtain a comprehensive uncertainty measure corresponding to each modal feature.
[0073] Furthermore, the feature fusion decoding module 14 in the multi-sensor fusion unmanned vehicle environmental perception system is also used for: aligning the modal features in the spatial dimension, concatenating the feature vectors at each spatial location to form a joint feature tensor; normalizing the inverse of the uncertainty measure of each modality at the same spatial location to obtain the weight coefficient of each modality at the same spatial location, which is used as the dynamic modulation coefficient; weighting and summing the joint feature tensor and the dynamic modulation coefficient to obtain the uncertainty-weighted fusion feature; decoding based on the fusion feature to output the perception result, and simultaneously outputting the comprehensive uncertainty confidence obtained by weighting and summing the comprehensive uncertainty measures corresponding to each modal feature.
[0074] Furthermore, the path planning module 15 in the multi-sensor fusion unmanned vehicle environment perception system is also used to: acquire drivable areas and static obstacles based on the static perception results in the perception results, and construct a global reference path from the current position to the target position; determine a baseline safety distance based on the dynamic perception results in the perception results, dynamically adjust the baseline safety distance using the comprehensive uncertainty confidence level, and obtain a dynamic safety distance; and perform local trajectory planning using a constrained optimization method based on the global reference path and the dynamic safety distance, and output a path planning strategy.
[0075] Furthermore, the path planning module 15 in the multi-sensor fusion autonomous vehicle environment perception system is also used to: construct an optimization objective function, including the deviation cost, control quantity smoothing cost, and collision cost of the global reference path; set constraints, including vehicle kinematic constraints, drivable area boundary constraints, and obstacle avoidance hard constraints; initialize the path state space based on the global reference path; and perform numerical iterative optimization in the path state space according to the optimization objective function and the constraints until convergence, obtaining a local trajectory that satisfies the constraints and has the minimum cost, and outputting a path planning strategy.
[0076] Furthermore, the multi-sensor fusion autonomous vehicle environment perception system also includes: constructing an adaptive spatial alignment constraint using the uncertainty metric, performing real-time online correction on the extrinsic parameters of the sensor array, obtaining the corrected sensor extrinsic parameters, updating the spatially aligned multimodal data, and forming a perception closed loop.
[0077] Furthermore, the multi-sensor fusion autonomous vehicle environment perception system also includes: when the uncertainty metric corresponding to a certain mode exceeds a preset threshold, it is determined that there is a deviation in the current spatial alignment, and the projection error of the high-confidence perception result among the sensor data is extracted; if the projection error deviates significantly from the expected error, it is determined that there is a shift in the current extrinsic parameters, and an adaptive spatial alignment constraint for online updating of extrinsic parameters is constructed with minimizing the projection consistency error of the high-confidence perception target among the sensor data as the optimization objective.
[0078] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The multi-sensor fusion autonomous vehicle environmental perception method and specific examples in the aforementioned Embodiment 1 are also applicable to the multi-sensor fusion autonomous vehicle environmental perception system of this embodiment. Through the foregoing detailed description of the multi-sensor fusion autonomous vehicle environmental perception method, those skilled in the art can clearly understand the multi-sensor fusion autonomous vehicle environmental perception system of this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.
[0079] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0080] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A multi-sensor fusion method for environmental perception in unmanned vehicles, characterized in that, include: Based on the autonomous driving operation design domain, the perception requirement indicators are determined, and sensor arrays are configured on the driverless vehicle. Based on the sensor array, environmental perception is performed to obtain the raw environmental datasets from multiple heterogeneous sensors, and time synchronization and spatial alignment are performed to obtain multimodal data at the same time and transformed into a unified coordinate system. Uncertainty estimation is performed on the multimodal data to obtain the modal features and corresponding uncertainty measures; The modal features are fused together, and the uncertainty measure is used as a dynamic modulation coefficient to apply to the fused modal features. The fused features weighted by uncertainty are decoded as the perception output, and the perception result and the corresponding comprehensive uncertainty confidence are output at the same time. Based on the perception results and the corresponding comprehensive uncertainty confidence level, a hierarchical and progressive path planning strategy is executed.
2. The multi-sensor fusion method for environmental perception of unmanned vehicles according to claim 1, characterized in that, An adaptive spatial alignment constraint is constructed using the uncertainty metric, and the extrinsic parameters of the sensor array are corrected in real time online to obtain the corrected sensor extrinsic parameters. The spatially aligned multimodal data is then updated to form a sensing closed loop.
3. The multi-sensor fusion method for environmental perception of unmanned vehicles according to claim 2, characterized in that, Constructing adaptive spatial alignment constraints using the aforementioned uncertainty metric includes: When the uncertainty measure corresponding to a certain mode exceeds the preset threshold, it is determined that there is a deviation in the current spatial alignment, and the projection error of the high-confidence perception result between the data of each sensor is extracted. If the projection error deviates significantly from the expected error, it is determined that the current extrinsic parameters have shifted. With minimizing the projection consistency error of the high-confidence sensing target among the sensor data as the optimization objective, an adaptive spatial alignment constraint for online updating of extrinsic parameters is constructed.
4. The multi-sensor fusion method for environmental perception of unmanned vehicles according to claim 1, characterized in that, Uncertainty estimation is performed on the multimodal data to obtain the modal features and corresponding uncertainty measures, including: Feature extraction is performed on the multimodal data to obtain the features of each modality; For each modal feature, the difference between the feature vectors of each feature point and other feature points in its neighborhood is calculated, and the average value of the feature vector differences is used as a measure of random uncertainty. For the same modal features, the feature points at the current time are matched with the feature points at the corresponding positions at historical times in a temporal sequence, and the cumulative change of the matching residual within the time window is calculated as a measure of cognitive uncertainty. By weighted and fused the random uncertainty measure and the cognitive uncertainty measure of each modality feature, a comprehensive uncertainty measure corresponding to each modality feature is obtained.
5. The multi-sensor fusion method for environmental perception of unmanned vehicles according to claim 1, characterized in that, The modal features are fused, and the aforementioned uncertainty measure is used as a dynamic modulation coefficient applied to the fused modal features. The uncertainty-weighted fused features are decoded as the sensing output, and the sensing result and the corresponding comprehensive uncertainty confidence level are output, including: After aligning the modal features in the spatial dimension, the feature vectors at each spatial location are concatenated to form a joint feature tensor. Normalize the reciprocal of the uncertainty measure of each mode at the same spatial location to obtain the weight coefficient of each mode at the same spatial location, which is used as the dynamic modulation coefficient; The joint feature tensor and the dynamic modulation coefficients are weighted and summed to obtain the fused features weighted by uncertainty. Decoding is performed based on the fused features, and the perception result is output. At the same time, the comprehensive uncertainty confidence level, which is obtained by weighted summation of the comprehensive uncertainty measures corresponding to each modality feature, is also output.
6. The multi-sensor fusion method for environmental perception of unmanned vehicles according to claim 1, characterized in that, Based on the perception results and the corresponding comprehensive uncertainty confidence level, a hierarchical and progressive path planning strategy is executed, including: Based on the static perception results, the drivable area and static obstacles are obtained, and a global reference path from the current position to the target position is constructed. Based on the dynamic perception results in the perception results, a baseline safety distance is determined, and the baseline safety distance is dynamically adjusted using the comprehensive uncertainty confidence level to obtain the dynamic safety distance; Based on the global reference path and the dynamic safety distance, a constrained optimization method is used for local trajectory planning, and a path planning strategy is output.
7. The multi-sensor fusion method for environmental perception of unmanned vehicles according to claim 6, characterized in that, Based on the global reference path and the dynamic safety distance, a constrained optimization method is used for local trajectory planning, outputting a path planning strategy, including: Construct an optimization objective function, including the deviation cost of the global reference path, the smoothing cost of the control quantity, and the collision cost; Set constraints, including vehicle kinematic constraints, drivable area boundary constraints, and obstacle avoidance hard constraints; Initialize the path state space based on the global reference path; Based on the objective function and the constraints, numerical iterative optimization is performed in the path state space until convergence is achieved, resulting in a local trajectory that satisfies the constraints and minimizes the cost, and the path planning strategy is output.
8. A multi-sensor fusion-based environmental perception system for unmanned vehicles, characterized in that: The steps for implementing the multi-sensor fusion unmanned vehicle environment perception method according to any one of claims 1 to 7, wherein the multi-sensor fusion unmanned vehicle environment perception system comprises: The sensor configuration module is used to determine perception requirement indicators based on the autonomous driving operation design domain and configure sensor arrays on the driverless vehicle. The spatiotemporal synchronization module is used to perform environmental perception based on the sensor array, acquire the raw environmental datasets of multiple heterogeneous sensors, and perform time synchronization and spatial alignment to obtain multimodal data at the same time and transformed to a unified coordinate system. An uncertainty estimation module is used to perform uncertainty estimation on the multimodal data to obtain the modal features and corresponding uncertainty measures. The feature fusion decoding module is used to fuse features of various modes and apply the uncertainty measure as a dynamic modulation coefficient to each fused modal feature. The fused features weighted by uncertainty are decoded as the perception output, and the perception result and the corresponding comprehensive uncertainty confidence level are output simultaneously. The path planning module is used to execute a hierarchical and progressive path planning strategy based on the perception results and the corresponding comprehensive uncertainty confidence level.