Automatic driving perception optimization method in uncertain scene based on multi-sensor fusion

CN122454345BActive Publication Date: 2026-09-22WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610905821.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-22
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

[0004]为了解决相关技术中,常采用多传感器静态加权融合,导致最终感知系统的融合结果准确度不高,严重影响自动驾驶系统的安全性的技术问题,本申请提供一种基于多传感器融合的不确定场景下自动驾驶感知优化方法

Benefits of technology

通过获取目标车辆在自动驾驶模式下多个传感器测得的环境识别数据,对不同传感器对应环境识别数据之间的特征点进行匹配处理,并根据匹配后的特征点对环境识别数据进行图像区域划分,得到多个局部区域,根据局部区域内的未匹配特征点数量以及特征匹配度,确定不同传感器对应数据之间的冲突程度,再基于传感器在局部区域内的特征点分布特征,计算数据可靠程度,进而,通过各个传感器的冲突程度对数据可靠程度进行修正,计算得到局部区域内的动态融合权重,进而,再通过动态融合权重对各传感器在局部区域内的初始置信度进行加权融合,输出整体感知结果,在本申请中,综合考虑不同传感器在不确定环境下的数据可靠程度,通过动态融合权重对各个传感器的初始置信度进行加权融合,以提高最终感知系统的融合结果准确度以及自动驾驶系统的安全性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454345B_ABST
    Figure CN122454345B_ABST
Patent Text Reader

Abstract

The application discloses an automatic driving perception optimization method based on multi-sensor fusion in an uncertain scene, and relates to the technical field of automatic driving obstacle identification. The method comprises the following steps: acquiring environment identification data measured by multiple sensors of a target vehicle in an automatic driving mode, matching feature points of different sensors, dividing an image region according to the matched feature points, and obtaining multiple local regions; determining a conflict degree based on the number of unmatched feature points and the feature matching degree in the local region; calculating a data reliability degree based on the feature point distribution characteristics of each sensor in the local region; correcting the data reliability degree by the conflict degree, and calculating a dynamic fusion weight; and performing weighted fusion on the initial confidence of each sensor based on the dynamic fusion weight, and outputting an overall perception result. The application improves the fusion result accuracy of the final perception system and the safety of the automatic driving system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of obstacle recognition technology for autonomous driving, and more specifically to an autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios. Background Technology

[0002] With the advancement of urbanization and the development of the automotive industry, the number of cars owned nationwide has gradually increased, making road safety an indispensable part of travel. Most traffic accidents are caused by driver error. With the development of computer vision and automation technology, autonomous driving has become a current research hotspot, and safety is the most primary and critical issue.

[0003] Current autonomous driving perception often employs static weighted fusion of multiple sensors. However, in uncertain scenarios, different sensors are affected by environmental interference to varying degrees. Relying solely on simple static weighted fusion can lead to low accuracy in the final perception system's fusion result, severely impacting the safety of the autonomous driving system. Summary of the Invention

[0004] To address the technical problem that the static weighted fusion of multiple sensors, often used in related technologies, results in low accuracy of the final perception system's fusion result, which seriously affects the safety of autonomous driving systems, this application provides an autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios.

[0005] The specific technical solution adopted is as follows: Acquire environmental identification data of the target vehicle in autonomous driving mode from multiple sensors, including cameras and LiDAR. Feature points are matched between environmental recognition data from different sensors, and the image regions of the environmental recognition data are divided based on the matched feature points to obtain multiple local regions. Based on the number of unmatched feature points and the feature matching degree within a local area, the degree of conflict between data from different sensors is determined; based on the feature point distribution characteristics of each sensor within a local area, the reliability of data from each sensor within the local area is calculated. The reliability of the data is corrected by the degree of conflict, and the dynamic fusion weight of each sensor in the local area is calculated. The initial confidence levels of each sensor in a local area are weighted and fused based on dynamic fusion weights to output the overall perception result.

[0006] In one possible implementation of this application, feature points are matched between environmental recognition data from different sensors, and the environmental recognition data is divided into image regions based on the matched feature points to obtain multiple local regions, including: Based on the distance differences between feature points in environmental identification data corresponding to different sensors, the feature matching degree between any feature points is calculated. Each feature point is matched according to its feature matching degree to obtain multiple feature matching pairs; Based on feature matching pairs, the environmental recognition data is divided into image regions to obtain multiple local regions.

[0007] In one possible implementation of this application, the feature matching degree between any feature points is calculated based on the distance difference between feature points in environmental identification data corresponding to different sensors, including: Determine the Euclidean distance between feature points of environmental identification data corresponding to different sensors, as well as the angle between gradient directions; The feature matching degree between any feature points is calculated based on the cosine value of the angle between gradient directions and the ratio between Euclidean distances.

[0008] In one possible implementation of this application, each feature point is matched according to its feature matching degree to obtain multiple feature matching pairs, including: Compare the feature matching degree with the preset matching threshold; Feature points with a feature matching degree greater than a preset matching threshold are grouped together and arranged in descending order of feature matching degree to obtain the first sequence; The combination of feature points in the first sequence is deduplicated and matched to obtain multiple feature matching pairs, so that any feature point can be matched with at most one other feature point.

[0009] In one possible implementation of this application, the environment recognition data is divided into image regions based on feature matching pairs to obtain multiple local regions, including: For any feature matching pair, the midpoint of the line segment connecting the feature points in the feature matching pair is taken as the representative position point of the current feature matching pair. Based on representative location points, the environmental recognition data is divided into image regions using a preset spatial segmentation algorithm to obtain multiple local regions, where each local region corresponds to any feature matching pair.

[0010] In one possible implementation of this application, the degree of conflict between data from different sensors is determined based on the number of unmatched feature points and the feature matching degree within a local region, including: For any feature matching pair, determine the feature matching degree of the current feature matching pair and the number of unmatched feature points in the corresponding local region; Extract the area of ​​a local region, and the average area among other local regions adjacent to the current local region; The degree of conflict between data from different sensors is calculated based on the number of unmatched feature points, the feature matching degree, and the ratio between the average area and the region area.

[0011] In one possible implementation of this application, the reliability of data from each sensor in a local area is calculated based on the feature point distribution characteristics of each sensor within a local area, including: Determine the first number of unmatched feature points in the image data corresponding to the camera within a local area, and the image contrast of the image data captured by different cameras; For any local area, traverse all feature points in the corresponding point cloud data of the lidar, and select the second number of feature points that belong to signal points and the third number that belong to noise points. The reliability of data from each sensor in a local area is calculated based on the ratio between the second and third quantities, and the product of the first quantity and the image contrast.

[0012] In one possible implementation of this application, the reliability of the data is corrected based on the degree of conflict, and the dynamic fusion weights of each sensor in a local area are calculated, including: The data reliability is normalized to obtain a normalized value; By using the degree of conflict as the exponent of the normalized value, and calculating the weight allocation of the normalized value, the dynamic fusion weight of each sensor in the local area is obtained.

[0013] In one possible implementation of this application, the initial confidence levels of each sensor within a local region are weighted and fused based on dynamic fusion weights to output an overall perception result, including: The first initial confidence level of the two-dimensional target candidate box output by the camera in the local area and the second initial confidence level of the three-dimensional target cluster block output by the lidar in the local area are obtained. Based on dynamic fusion weights, the first initial confidence level and the second initial confidence level are weighted and fused to output the overall perception result.

[0014] In one possible implementation of this application, based on dynamic fusion weights, a weighted fusion of a first initial confidence level and a second initial confidence level is performed to output an overall perception result, including: Based on dynamic fusion weights, the first initial confidence and the second initial confidence are weighted and fused to output the local fusion results for each local region; The local fusion results of each region are stitched together in a globally unified coordinate system to output the overall perception result.

[0015] This application has, but is not limited to, the following technical effects: By acquiring environmental recognition data from multiple sensors of the target vehicle in autonomous driving mode, feature points between environmental recognition data from different sensors are matched, and the environmental recognition data is divided into image regions based on the matched feature points to obtain multiple local regions. The degree of conflict between data from different sensors is determined based on the number of unmatched feature points and the feature matching degree within each local region. Then, the data reliability is calculated based on the feature point distribution characteristics of the sensors within the local regions. Subsequently, the data reliability is corrected by the degree of conflict of each sensor, and the dynamic fusion weight within the local region is calculated. Then, the initial confidence of each sensor within the local region is weighted and fused using the dynamic fusion weight to output the overall perception result. In this application, the reliability of data from different sensors in uncertain environments is comprehensively considered, and the initial confidence of each sensor is weighted and fused using dynamic fusion weight to improve the accuracy of the final perception system's fusion result and the safety of the autonomous driving system. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the first embodiment of the autonomous driving perception optimization method for uncertain scenarios based on multi-sensor fusion in this application; Figure 2 This is a schematic diagram of the overall implementation process of the autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios in this application; Figure 3 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of this application. Detailed Implementation

[0017] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0018] This application provides a method for optimizing perception in uncertain scenarios for autonomous driving based on multi-sensor fusion. In the first embodiment of this method, referring to... Figure 1 and Figure 2 The methods include: Step S10: Obtain environmental recognition data measured by multiple sensors of the target vehicle in autonomous driving mode. The sensors include cameras and LiDAR.

[0019] As an example, the scenario addressed in this application could be: during the autonomous driving process of a car, the environmental scenario is uncertain, and the impact of the uncertain scenario on each sensor is also unknown. Directly fusing sensor data according to predetermined weights may use unreliable data, resulting in low accuracy and poor safety of the fusion result.

[0020] As an example, the multi-sensor fusion perception system deployed on autonomous vehicles mainly includes sensors such as cameras and LiDAR. Their deployment locations and numbers vary. For instance, LiDAR sensors are fewer in number and mostly distributed on the roof and front of the vehicle, while cameras are relatively more numerous and distributed throughout the vehicle body, including the front and rear bumpers, rearview mirrors, and fenders. During autonomous driving, these sensors acquire environmental recognition data about the vehicle's surroundings in real time and upload it to the perception system. This environmental recognition data includes two-dimensional image data of the target vehicle's surrounding environment captured by cameras and three-dimensional point cloud data captured by LiDAR.

[0021] As an example, after acquiring environmental identification data, the data acquisition times of each sensor are synchronized via hardware devices (such as a GNSS receiver receiving pulse signals from GPS and outputting them to the synchronization interfaces of each sensor). Since the field of view of a camera is typically a fan-shaped area, while the field of view of a LiDAR is typically a 360° circular area, each sensor has its own individually covered area and an area covered by both. By simultaneously photographing the same checkerboard calibration board with both LiDAR and camera, and by continuously changing the position of the calibration board, enough data sets are captured where both LiDAR and camera can cover the location of the calibration board. Based on the area where the calibration board appears in the data sets, the local area jointly covered by the data acquired by both sensors is defined.

[0022] As an example, multiple corner points on a checkerboard pattern are extracted from images captured by a camera (2D coordinates) and point cloud data acquired by a LiDAR (3D coordinates). The initial extrinsic parameters are solved using multiple sets of 3D-2D corresponding point pairs, and adjustments and optimizations are made with the goal of reducing "reprojection error". This achieves preliminary spatiotemporal alignment of the data obtained by the LiDAR sensor and the camera sensor in the same shooting area, so as to evaluate the reliability of the data obtained by each sensor in each area later.

[0023] Step S20: Match feature points between environmental recognition data from different sensors, and divide the environmental recognition data into image regions based on the matched feature points to obtain multiple local regions.

[0024] As an example, the spatial alignment method described above is a calibration board method. Spatially aligning environmental recognition data from different sensors allows the vehicle to achieve good spatial alignment before driving, aided by a calibration board. However, during vehicle operation, the vehicle's driving state and surrounding environment constantly change dynamically. The environment captured by the sensors also changes relative to the vehicle. Without the reference of the calibration board, the spatial consistency of the scene areas captured by different sensors is affected by vehicle movement and changes in the surrounding environment. Data alignment between different sensors will exhibit a certain degree of deviation, affecting the relative positions of the sensors, while the environment itself within the same area remains unchanged. Therefore, based on the matching of feature points in the environmental recognition data acquired by different sensors, regions belonging to the same environment can be matched. Based on this, the environmental recognition data is divided into image regions according to the matched feature points, resulting in multiple local regions.

[0025] Among them, step S20, which optimizes autonomous driving perception in uncertain scenarios based on multi-sensor fusion, also includes steps S21 to S23, including: Step S21: Based on the distance difference between feature points in the environmental identification data corresponding to different sensors, calculate the feature matching degree between any feature points.

[0026] As an example, taking LiDAR and a camera as sensors, the environmental recognition data corresponding to LiDAR is point cloud data, and the environmental recognition data corresponding to the camera is image data. Both are used to reflect the position and distance of objects in the vehicle's surrounding environment. The objects and backgrounds differ significantly, and there are multiple feature points in the data corresponding to each sensor. First, it is necessary to unify the data dimensions between LiDAR and the camera. Specifically, the two-dimensional position of the point cloud on the image plane is calculated by using the camera focal length and the three-dimensional coordinate data of the point cloud. Then, the extrinsic and intrinsic parameter matrices of the camera are combined to project the 3D LiDAR point cloud data onto the 2D image plane (the specific implementation steps are well-known technology and will not be elaborated here). The three-dimensional point cloud data obtained by LiDAR is converted into a two-dimensional depth image, so that the discrete three-dimensional spatial coordinates are transformed into matrix data with a regular two-dimensional grid topology. Subsequently, based on the SIFT algorithm (existing technology, not elaborated here), feature points are extracted and identified from the two-dimensional depth image converted from the image obtained by the camera sensor and the point cloud environmental data obtained by the lidar, respectively, in order to match the feature points in the corresponding data of the two sensors. The point cloud feature points and the image feature points are aligned proportionally in the scale space. The alignment method is existing technology, which will not be elaborated here.

[0027] Step S21 includes: Determine the Euclidean distance between feature points of environmental identification data corresponding to different sensors, as well as the angle between gradient directions.

[0028] The feature matching degree between any feature points is calculated based on the cosine value of the angle between gradient directions and the ratio between Euclidean distances.

[0029] As an example, when extracting feature points from environmental recognition data using the SIFT (Scale Invariant Feature Transform) algorithm, a reference gradient direction is calculated and assigned to each feature point to ensure rotation invariance. This reference gradient direction is the gradient direction described here. Furthermore, the angle between the gradient directions of two feature points can be obtained. This angle essentially compares the main direction of the image texture or point cloud gradient observed by two sensors at the same physical location. The smaller the value (i.e., the more consistent the main directions), the higher the matching degree, indicating that the two sensors have captured the same geometric edges or structural features. The range of values ​​is Similarly, the smaller the Euclidean distance between two feature points, the closer they are, which means the higher the matching degree.

[0030] As an example, the feature matching degree between any two feature points The calculation method is as follows: in, This represents the cosine value corresponding to the angle between the gradient directions of two feature points. Represents Euclidean distance. To preset a minimum positive value and prevent the denominator from being zero, the Euclidean distance between two feature points is calculated. The smaller the value, the closer the two are; The smaller, Larger and closer to This indicates that the smaller the difference in the principal directions between two feature points, the higher the matching degree between the two feature points. The range of values ​​is , Avoid negative numerators.

[0031] Step S22: Perform matching processing on each feature point according to the feature matching degree to obtain multiple feature matching pairs.

[0032] As an example, when obtaining feature matching degree Afterwards, regarding the results The value is normalized to obtain (The normalization process in this scheme uses the maximum-minimum normalization method to map the calculated values ​​to...) The interval, where the maximum and minimum values ​​are obtained from boundary sample data during the system's historical normal operation (the normalization used thereafter is consistent with this, and will not be repeated here). This yields the normalized feature matching degree between any pair of feature points. The process of matching each feature point according to the feature matching degree is also based on this normalized feature matching degree. To be processed.

[0033] Step S22 includes: The feature matching degree is compared with the preset matching threshold.

[0034] Feature points with a matching degree greater than a preset matching threshold are grouped together and arranged in descending order of matching degree to obtain the first sequence.

[0035] The combination of feature points in the first sequence is deduplicated and matched to obtain multiple feature matching pairs, so that any feature point can be matched with at most one other feature point.

[0036] As an example, the preset matching threshold can be 0.6, 0.7, etc., and can be changed according to user needs. There is no specific limitation.

[0037] As an example, the normalized feature matching score is compared with a preset matching threshold. Value greater than Feature point pairs, with Arrange the values ​​in descending order to obtain the first sequence. The method for deduplicating and matching the combinations of feature points in the first sequence can be as follows: Starting with the first (highest matching) feature point pair in the sorted list, these two points are considered successfully paired. Subsequently, these two points are removed from their respective sensor's unmatched feature point sets (labeled as matched), and the process continues to check the next candidate feature point pair in the sorted list: If neither of the two points in a candidate feature point pair has been matched before, the pairing is successful and they are labeled; if either point in a candidate feature point pair has already been used in a previous high-scoring match, the candidate pair is discarded.

[0038] This process continues downwards until the entire sorted list has been traversed. Since points are removed once a match is found, this ensures that "each feature point is matched at most once," resolving many-to-one or many-to-many mismatch conflicts that occur when multiple feature points are clustered together, so that any feature point is matched with at most one other feature point.

[0039] Step S23: Based on feature matching pairs, the environment recognition data is divided into image regions to obtain multiple local regions.

[0040] As an example, consider two feature points in a feature matching pair: one from the image captured by the camera, and the other from the point cloud data acquired by the LiDAR. During vehicle movement, the vehicle's driving state and surrounding environment constantly change dynamically. The environment captured by the sensors also changes relative to the vehicle. The spatial consistency of the scene areas captured by different sensors is affected by vehicle movement and changes in the surrounding environment, influencing the relative positions of the sensors. However, the environment itself within the same area remains unchanged. Therefore, based on the matching of feature points in the environmental recognition data acquired by different sensors, regions belonging to the same environment can be matched to obtain multiple local regions. Each local region contains only one successfully matched feature pair and several unmatched feature points.

[0041] Step S23 includes: For any feature matching pair, the midpoint of the line segment connecting the feature points in the feature matching pair is taken as the representative position point of the current feature matching pair.

[0042] Based on representative location points, the environmental recognition data is divided into image regions using a preset spatial segmentation algorithm to obtain multiple local regions, where each local region corresponds to any feature matching pair.

[0043] As an example, the image region of the environment recognition data is divided based on any matched feature pairs. The midpoint of the line segment connecting two feature points of each feature pair is recorded as the representative position point of the current feature pair. The preset spatial segmentation algorithm can be the Voronoi spatial segmentation algorithm (existing technology, not described in detail here). After obtaining the representative position point, the image plane corresponding to the environment recognition data is divided into regions using this algorithm, resulting in multiple local regions. Before dividing the regions, the number of feature matching pairs in each region is determined. If there are fewer than 3, a uniform grid (e.g., 8x8) is used as the local region; if there are 3 or more, Voronoi spatial segmentation is performed. The specific segmentation method can be: The representative locations of each feature matching pair are used as the seed point set for generating the Voronoi diagram. All seed points are connected to form a series of non-intersecting triangles, ensuring that the circumcircle of any triangle does not contain other seed points. The perpendicular bisectors of each triangle's sides are drawn; the intersections of these perpendicular bisectors are the vertices of the Voronoi polygons, and the perpendicular bisector segments constitute the boundaries of the Voronoi polygons. Since the perpendicular bisectors of the outermost seed points extend infinitely outward, the physical boundary of the image (or the effective field of view boundary of the camera) is used as a global constraint truncation condition to close the outwardly diverging rays, forming the final closed region segmentation.

[0044] Through the above process, the two-dimensional plane is divided into multiple non-overlapping closed polygonal regions. Each region contains only one feature matching pair, and the distance from any point within that region to the seed point is less than the distance to any other seed point. This achieves adaptive local region segmentation of the environment based on the matching point density.

[0045] Step S30: Based on the number of unmatched feature points and the feature matching degree in the local area, determine the degree of conflict between the corresponding data of different sensors; based on the feature point distribution characteristics of each sensor in the local area, calculate the data reliability of each sensor in the local area.

[0046] As an example, the degree of conflict is used to represent the degree of inconsistency or contradiction in environmental data obtained by different sensors (LiDAR and camera) within the same local area. The greater the degree of conflict, the more serious the data contradiction (for example, LiDAR generates a lot of noise due to rain and snow, while the camera does not recognize the corresponding features).

[0047] As an example, the number of unmatched feature points and the feature matching degree are used to reflect the consistency between the data acquired by different sensors. When the number of unmatched feature points is greater, it means that there are more feature points that the two sensors do not match in the current local area, and the degree of conflict in the environmental identification data acquired by the two sensors in this area is higher. Similarly, the smaller the feature matching degree corresponding to the feature matching pair, the lower the degree of matching between the two matched feature points, and the higher the degree of difference in the environmental identification data acquired by the two sensors. Based on this correspondence, the degree of conflict is calculated.

[0048] As an example, taking LiDAR and cameras as examples, data reliability is used to represent the credibility of LiDAR relative to the camera within a current local area. During the operation of an autonomous vehicle, the uncertain environment it is in affects its sensors to varying degrees. For example, in dark scenes such as tunnels, the images obtained by the camera are dim and difficult to recognize, leading to the loss of some data; in heavy snow scenes, large snowflakes and water droplets reflect off the LiDAR, resulting in additional data. Therefore, the accuracy of data obtained by each sensor varies in different scenarios, and their reference value in the final fusion process also differs. Thus, it is necessary to evaluate the reliability of the vehicle's sensor data in the current scenario in real time for subsequent multi-sensor data fusion operations.

[0049] As an example, the reliability of data from different sensors in the same local area can be determined based on the distribution characteristics of feature points within that area. For instance, if a camera identifies more feature points than a LiDAR sensor, it indicates that the camera obtains more detailed data, and the reliability of the LiDAR sensor's data in that area is relatively lower than that of the camera's data.

[0050] The step of determining the degree of conflict between data from different sensors based on the number of unmatched feature points and the feature matching degree within a local area includes: For any feature matching pair, determine the feature matching degree of the current feature matching pair and the number of unmatched feature points in the corresponding local region.

[0051] As an example, take any set of feature matching pairs and obtain the normalized feature matching degree of that pair. The value is further denoted as the number of unmatched feature points existing in the feature matching pair and its surrounding area. .

[0052] Extract the area of ​​a local region, and the average area among other local regions adjacent to the current local region.

[0053] As an example, when a region has many feature points and is divided into many local regions, the environment within that region is more likely to be complex, easily affecting the sensor and thus impacting the data it obtains. Let the area of ​​the local region containing the feature matching pair be denoted as . Further calculations show that the average area of ​​each adjacent local region is... .

[0054] The degree of conflict between data from different sensors is calculated based on the number of unmatched feature points, the feature matching degree, and the ratio between the average area and the region area.

[0055] As an example, the degree of conflict between any feature match and the surrounding environmental data of its local region is denoted as . , The calculation formula is: in, This represents the number of unmatched feature points in the local region to which the current feature matching pair belongs. The larger this value is, the more feature points that the two sensors do not match in that region, and the higher the degree of conflict between the environmental data obtained by the two sensors in that region. This represents the normalized feature matching degree. This represents the ratio between the average area and the area of ​​the region. The larger the value, the more complex the current feature matching is to the local area, and the more likely the environmental data obtained by the two sensors in this local area will be affected and conflict will occur.

[0056] The step of calculating the data reliability of each sensor in a local area based on the feature point distribution characteristics of each sensor in the local area includes: Determine the first number of unmatched feature points in the image data corresponding to the camera within a local area, and the image contrast of the image data captured by different cameras; As an example, the images obtained by the camera are more realistic and intuitive than the 3D point clouds obtained by the LiDAR, but they are greatly affected by the ambient light. Let the number of unmatched feature points obtained by the camera within the selected local area be denoted as... (Record as the first quantity, or record as if it does not exist) Furthermore, the image contrast (local pixel grayscale standard deviation) of the image data captured by each camera in this local area is denoted as . .

[0057] For any local area, traverse all feature points in the corresponding point cloud data of the lidar, and select the second number of feature points that belong to signal points and the third number that belong to noise points.

[0058] As an example, within a selected local area, all feature points in the lidar point cloud data are traversed. If the reflection intensity of a feature point is less than... Or the number of neighboring points within a preset radius R is less than If the signal is positive, the point is considered a noise point; otherwise, it is considered a signal point. The total number of signal points in the current local area is recorded as follows: (Recorded as the second quantity), the total number of noise points is denoted as (Recorded as the third quantity; if it is 0, then it is recorded as 1). , R and The specific values ​​are obtained based on the sensor hardware characteristics (such as laser wavelength and beam divergence angle) and the noise distribution statistical curve under typical weather conditions; in one embodiment, The value is 30, and R is 0.3 meters. The value is 4.

[0059] The reliability of data from each sensor in a local area is calculated based on the ratio between the second and third quantities, and the product of the first quantity and the image contrast.

[0060] As an example, the reliability of the environmental identification data from the lidar sensor relative to the camera within the acquired local area is denoted as... , The calculation formula is: In the formula, The larger this value is, the more intuitive the camera can identify, the more feature points it can identify than the LiDAR, the more detailed data the camera obtains, and the lower the reliability of the data obtained by the LiDAR sensor relative to the data obtained by the camera in this area. Image contrast, its value range is: , Will Value normalized to The larger this value is, the greater the difference in pixel brightness in the uncertain scene where the vehicle is located, the clearer the image obtained by the camera is more likely to be, and the higher its reliability is compared with the data obtained by the lidar. This indicates the second quantity, which belongs to the valid feature points. The more features a lidar identifies, the more effective feature points it can identify around the outline of the area it captures. This results in more detailed data compared to the camera, and the higher the reliability of the data obtained. The third quantity represents the number of noise points, which are invalid feature points. The more features a lidar has, the higher the likelihood of it being affected by uncertain scenarios, and the lower the reliability of the data obtained. This indicates the effectiveness of the lidar's feature points. and The data from the LiDAR is directly proportional to the data from the camera, while the data from the camera is inversely proportional. This indicates the relative reliability of the LiDAR data compared to the camera data.

[0061] Furthermore, data reliability The larger the value, the higher the reliability of the data obtained by the lidar sensor relative to the camera sensor in that area, and the greater its reference value in the subsequent fusion process.

[0062] Step S40: Correct the data reliability based on the degree of conflict, and calculate the dynamic fusion weight of each sensor in the local area.

[0063] As an example, the degree of conflict (T) can be used as a "lever" to artificially amplify the advantage of data with higher reliability when data contradictions arise, achieving nonlinear reinforcement. This, in turn, amplifies the weight of sensors with higher data reliability in subsequent fusion processes. Dynamic fusion weights can be the weight values ​​assigned to different sensors during data fusion.

[0064] Step S40 includes: The reliability of the data is normalized to obtain a normalized value.

[0065] By using the degree of conflict as the exponent of the normalized value, and calculating the weight allocation of the normalized value, the dynamic fusion weight of each sensor in the local area is obtained.

[0066] As an example, regarding data reliability Normalization is performed to obtain the normalized value. After normalization, the reliability of the data is improved. A truncation constraint is applied so that the normalized value K' is strictly truncated within the interval [0, 1] to avoid erroneous calculations in extreme blizzard scenarios. Furthermore, the reliability of the environmental data obtained by the camera relative to the LiDAR within this local area is denoted as... .

[0067] As an example, the dynamic fusion weights of lidar The calculation method is as follows: in, During calculation, c is truncated to its upper limit (the truncation upper limit is...). The degree of conflict between the environmental data obtained by the two sensors in this area The larger the value, the higher the weight difference between the two should be. The parameters can be adjusted by the user; the default settings are as follows: , .like , The larger, and The closer it is to 1; similarly, if , The larger, and The closer to .

[0068] As an example, the dynamic fusion weights of the camera sensors are denoted as... .

[0069] Step S50: Based on the dynamic fusion weights, the initial confidence levels of each sensor in the local area are weighted and fused to output the overall perception result.

[0070] As an example, the dynamic fusion weights of the LiDAR sensor and the camera sensor in each region are finally obtained. and Based on the dynamic fusion weights, a multi-sensor fusion operation is performed on the local area to output the overall perception result, which may include the following data: 1. Target List: The location (3D bounding box / 2D bounding box) and category of all pedestrians, vehicles and obstacles identified in the global unified coordinate system (composed of stitched Voronoi local regions).

[0071] 2. Fused Confidence: Each target has a probability score that has been dynamically adjusted by fusion weights. Only targets with a weighted score that exceeds a threshold will be retained, thereby filtering out false alarms caused by environmental interference (such as false targets caused by rain or snow).

[0072] 3. Spatiotemporally aligned feature matrix: contains metadata such as position, size, and direction of motion, which can be directly used by the downstream "path planning" and "decision control" modules.

[0073] Step S50 includes steps S51 to S52: Step S51: Obtain the first initial confidence of the two-dimensional target candidate box output by the camera in the local area, and the second initial confidence of the three-dimensional target cluster block output by the lidar in the local area.

[0074] As an example, the first initial confidence level of the two-dimensional target candidate box output by the camera in the current area and the second initial confidence level of the three-dimensional target cluster block output by the lidar in the current area can be directly extracted from the autonomous driving system.

[0075] Step S52: Based on the dynamic fusion weight, the first initial confidence and the second initial confidence are weighted and fused to output the overall perception result.

[0076] Step S52 specifically includes: Based on dynamic fusion weights, the first initial confidence and the second initial confidence are weighted and fused to output the local fusion results for each local region; The local fusion results of each region are stitched together in a globally unified coordinate system to output the overall perception result.

[0077] As an example, the above-mentioned camera dynamic fusion weights are used. The initial confidence scores of the two-dimensional target candidate boxes are weighted and corrected, and the aforementioned dynamic fusion weights from the LiDAR are used. The initial confidence scores of the 3D target cluster blocks are weighted and corrected. Based on this, targets with a weighted and corrected confidence score greater than a preset threshold (which can be 0.7, determined empirically; if set too low (e.g., 0.3), false targets caused by rain and snow interference may be output, triggering emergency braking; if set too high (e.g., 0.9), real targets may be filtered out in harsh environments, causing collision risks) are selected. The corrected confidence scores of the two thresholds are then fused, and the output is the local fusion result for that local area.

[0078] As an example, the local fusion results of various regions are stitched together in a globally unified coordinate system to output the final overall perception result of the vehicle's surrounding environment. This provides basic data support for subsequent vehicle path planning, autonomous driving, and other operations.

[0079] This application provides a perception optimization method for autonomous driving in uncertain scenarios based on multi-sensor fusion. In this application, the reliability of data from different sensors in uncertain environments is comprehensively considered, and the initial confidence of each sensor is weighted and fused through dynamic fusion weights to improve the accuracy of the final perception system fusion result and the safety of the autonomous driving system.

[0080] Reference Figure 3 , Figure 3 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of this application.

[0081] like Figure 3 As shown, the autonomous driving perception optimization device based on multi-sensor fusion in uncertain scenarios may include: a processor 1001, a memory 1003, and a communication bus 1002. The communication bus 1002 is used to realize the connection and communication between the processor 1001 and the memory 1003.

[0082] Optionally, the autonomous driving perception optimization device based on multi-sensor fusion in uncertain scenarios may also include a user interface, a network interface, a camera, RF (Radio Frequency) circuitry, sensors, a WiFi module, etc. The user interface may include a display screen and an input submodule such as a keyboard; optionally, the user interface may also include standard wired or wireless interfaces. The network interface may include standard wired or wireless interfaces (such as a Wi-Fi interface).

[0083] Those skilled in the art will understand that Figure 3 The structure of the autonomous driving perception optimization device based on multi-sensor fusion in uncertain scenarios shown does not constitute a limitation on the autonomous driving perception optimization device based on multi-sensor fusion. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0084] like Figure 3 As shown, the memory 1003, serving as a storage medium, may include an operating system, a network communication module, and an autonomous driving perception optimization program for uncertain scenarios based on multi-sensor fusion. The operating system is a program that manages and controls the hardware and software resources of the autonomous driving perception optimization device for uncertain scenarios based on multi-sensor fusion, supporting the operation of the autonomous driving perception optimization program for uncertain scenarios based on multi-sensor fusion, as well as other software and / or programs. The network communication module is used to enable communication between the various components within the memory 1003, as well as communication with other hardware and software in the autonomous driving perception optimization system for uncertain scenarios based on multi-sensor fusion.

[0085] exist Figure 3In the autonomous driving perception optimization device based on multi-sensor fusion shown, the processor 1001 is used to execute the autonomous driving perception optimization program based on multi-sensor fusion in uncertain scenarios stored in the memory 1003, and implement the steps of the autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios described above.

[0086] The specific implementation of the autonomous driving perception optimization device based on multi-sensor fusion in this application is basically the same as the embodiments of the autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios described above, and will not be repeated here.

[0087] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0088] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0089] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0090] The above are merely preferred embodiments of this application and do not limit the scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of this application.

[0091] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0092] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A perception optimization method for autonomous driving in uncertain scenarios based on multi-sensor fusion, characterized in that, The method includes: Acquire environmental identification data of the target vehicle measured by multiple sensors in autonomous driving mode, including cameras and lidar; Feature points are matched between environmental recognition data from different sensors, and the environmental recognition data is divided into image regions based on the matched feature points to obtain multiple local regions. Based on the number of unmatched feature points and the feature matching degree within the local region, the degree of conflict between data from different sensors is determined, specifically including: For any feature matching pair, determine the feature matching degree of the current feature matching pair and the number of unmatched feature points in the corresponding local region; Extract the area of ​​the local region, and the average area among other local regions adjacent to the current local region; Based on the number of unmatched feature points, the feature matching degree, and the ratio between the average area and the region area, the degree of conflict between the data corresponding to different sensors is calculated. Based on the feature point distribution characteristics of each sensor within the local area, the reliability of the data from each sensor in the local area is calculated, specifically including: Determine the first number of unmatched feature points in the image data corresponding to the camera within the local area, and the image contrast of the image data captured by different cameras; For any local area, traverse all feature points in the point cloud data corresponding to the lidar, and select the second number of feature points that belong to signal points and the third number that belong to noise points. Based on the ratio between the second quantity and the third quantity, and the product between the first quantity and the image contrast, the data reliability of each sensor in the local area is calculated. The reliability of the data is corrected by the degree of conflict, and the dynamic fusion weight of each sensor in the local area is calculated. The initial confidence levels of each sensor within the local area are weighted and fused based on dynamic fusion weights to output the overall perception result.

2. The autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios as described in claim 1, characterized in that, The process involves matching feature points between environmental recognition data from different sensors, and then dividing the environmental recognition data into image regions based on the matched feature points to obtain multiple local regions, including: Based on the distance differences between feature points in environmental recognition data corresponding to different sensors, the feature matching degree between any feature points is calculated. Each feature point is matched according to the feature matching degree to obtain multiple feature matching pairs; Based on the feature matching pairs, the environmental recognition data is divided into image regions to obtain multiple local regions.

3. The autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios as described in claim 2, characterized in that, The calculation of the feature matching degree between any feature points based on the distance difference between feature points in environmental identification data corresponding to different sensors includes: Determine the Euclidean distance between feature points of environmental identification data corresponding to different sensors, as well as the angle between gradient directions; Based on the cosine value corresponding to the angle between the gradient directions and the ratio between the Euclidean distances, the feature matching degree between any feature points is calculated.

4. The autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios as described in claim 2, characterized in that, The matching process performed on each feature point according to the feature matching degree yields multiple feature matching pairs, including: The feature matching degree is compared with a preset matching threshold; The feature points with a feature matching degree greater than a preset matching threshold are combined and arranged in descending order of feature matching degree to obtain the first sequence; The combination of feature points in the first sequence is deduplicated and matched to obtain multiple feature matching pairs, so that any feature point can be matched with at most one other feature point.

5. The autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios as described in claim 2, characterized in that, The step of dividing the environment recognition data into image regions based on the feature matching pairs yields multiple local regions, including: For any feature matching pair, the midpoint of the line segment connecting the feature points in the feature matching pair is taken as the representative position point of the current feature matching pair. Based on the representative location points, the environmental recognition data is divided into image regions using a preset spatial segmentation algorithm to obtain multiple local regions, wherein each local region corresponds to any feature matching pair.

6. The autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios as described in claim 1, characterized in that, The step of correcting the data reliability based on the degree of conflict and calculating the dynamic fusion weights of each sensor within the local area includes: The reliability of the data is normalized to obtain a normalized value; The degree of conflict is used as the exponent of the normalized value, and the normalized value is weighted and calculated to obtain the dynamic fusion weight of each sensor in the local area.

7. The autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios as described in claim 1, characterized in that, The initial confidence levels of each sensor within the local region are weighted and fused based on dynamic fusion weights to output the overall perception result, including: The first initial confidence level of the two-dimensional target candidate box output by the camera in the local area and the second initial confidence level of the three-dimensional target cluster block output by the lidar in the local area are obtained. Based on the dynamic fusion weights, the first initial confidence level and the second initial confidence level are weighted and fused to output the overall perception result.

8. The autonomous driving perception optimization method based on multi-sensor fusion in uncertain scenarios as described in claim 7, characterized in that, The step of weighted fusion of the first initial confidence and the second initial confidence based on the dynamic fusion weights, and outputting the overall perception result, includes: Based on the dynamic fusion weights, the first initial confidence and the second initial confidence are weighted and fused to output the local fusion results of each local region; The local fusion results of each local region are stitched together in a globally unified coordinate system to output the overall perception result.

Citation Information

Patent Citations

  • Wharf unmanned vehicle obstacle detection and identification method based on multi-sensor fusion

    CN120766241A

  • Three-dimensional model and geographic coordinate dynamic matching method based on multi-source data fusion

    CN121213810A