Multi-mode assisted analysis and re-calibration method for vehicle detection failure sample
Through multimodal-assisted vehicle detection failure sample analysis and recalibration methods, the problem of low recognition accuracy in complex environments of the autonomous driving system is solved, and the effect of improving recognition accuracy and saving computing power is achieved.
Patent Information
- Application Number
- CN202411859920.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-13
AI Technical Summary
In complex environments such as overexposure, haze and object occlusion, the recognition accuracy of the on-board camera is low, and existing methods increase the computing power burden.
Multimodal-assisted vehicle detection failure sample analysis and recalibration methods are used to collect failure samples, custom interpretability interpreters for Shapley value analysis, statistical change laws, analyze the causes of failure and select appropriate sensors for assisted identification.
It improves the recognition accuracy and robustness of the system in complex environments, reduces convolutional calculations for unimportant areas, reduces computing power requirements, and maintains high recognition accuracy.
Smart Images

Figure CN120147683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-modal assistance to improve the recognition accuracy algorithm, and particularly relates to a method for analyzing and recalibrating vehicle detection failure samples with multi-modal assistance. Background Art
[0002] With the continuous development of artificial intelligence, its combination with the vehicle industry is also in progress. In recent years, the development of autonomous driving has been in full swing. Using autonomous driving to replace manual driving, no driver intervention is required during the driving process. The emergence of autonomous driving makes people's travel safer and their lives more convenient. When the driver is fatigued or inattentive, autonomous driving can help humans make decisions, improve the safety and stability of driving, and avoid accidents. Moreover, autonomous driving can optimize parameters such as vehicle driving routes and speeds, reduce unnecessary energy consumption, thereby reducing environmental pollution and waste of resources. For complex and congested road conditions, autonomous driving can also better control the direction and avoid pedestrians and vehicles, giving people a better driving experience.
[0003] Autonomous driving mainly consists of three parts: a perception module, a decision-making module, and a control module. Among them, the perception module collects surrounding environmental information through different sensors, and then uses a convolutional neural network model to identify the collected information, simulating the human brain. The decision-making module mainly judges the next action of the vehicle based on the information collected by the perception module through an algorithm and tells it to the control module. The control module is responsible for transmitting the decision-making instructions to the vehicle hardware and driving it. In summary, it can be seen that the perception module is the most critical step in the entire autonomous driving and is the soul of autonomous driving. Everything is based on perception.
[0004] However, the current maturity of the perception module is not very high, and mistakes occur frequently. The accuracy and applicability of the algorithm still need to be improved. For autonomous driving, the camera is equivalent to the vehicle's eyes. The vehicle takes pictures of the surrounding scenery, and the obtained pictures are calculated and recognized through the algorithm of the autonomous driving system. With the rapid development of deep neural network technology and big data, autonomous driving recognition algorithms mainly rely on these deep learning technologies. By training the neural network with a large amount of data and continuously increasing the number of network parameters, the accuracy is improved. However, models based on image recognition are prone to incorrect predictions in some special scenarios, which may lead to some fatal traffic accidents. Moreover, visible light cameras have limitations. Only two-dimensional detection is performed through convolutional neural networks. Not only will it detect errors in some complex situations, but too many convolutional layers and parameters will overload autonomous driving, consume too much computing power, and incur immeasurable costs. Therefore, solving the recognition problem is crucial, and reducing the recognition computing power of autonomous driving while ensuring recognition accuracy is also a great challenge. Summary of the Invention
[0005] At present, although autonomous driving technology has developed rapidly, there are still many problems. Among them, the main concern is that the on-vehicle cameras of autonomous driving have relatively low recognition accuracy in some cases, such as overexposed scenes, haze scenes, and object occlusion scenes. However, currently, the problem of improving the recognition accuracy of on-vehicle cameras usually focuses on strengthening the training of the internal convolutional neural network, increasing the number of parameters or the number of convolutional layers. This method will increase the computing power required for autonomous driving. Therefore, based on these two problems, this paper proposes a research that can not only improve the recognition accuracy in overexposed scenes, haze scenes, and object occlusion scenes, but also save the computing power used for autonomous driving. The main content is as follows:
[0006] A method for analyzing and recalibrating vehicle detection failure samples assisted by multi-modalities, comprising the following steps:
[0007] S1: Collect the failure samples captured by the on-vehicle camera during daily driving and organize and classify them;
[0008] S2: Customize an interpretable interpreter, use the Shapley value for interpretation, and obtain an interpretable map of the failure samples with distributed Shapley values;
[0009] S3: Use statistical sampling to statistically analyze the Shapley values in the interpretable map and obtain the corresponding change rules of the Shapley values;
[0010] S4: Compare the statistical data of different scenes and analyze the reasons for sample failure;
[0011] S5: Collect the basic information of the sensors recognized by the on-vehicle camera and analyze which type of failure sample each sensor is suitable for solving;
[0012] S6: Project the point cloud data recognized by the corresponding sensor onto a two-dimensional image through reprojection, obtain the corresponding detection boxes according to the minimum adjacency matrix, match these detection boxes with the detection boxes recognized by the visible light camera through the convolutional neural network model, and calculate the corresponding detection accuracy.
[0013] Preferably, the failure samples in S1 are specifically classified into overexposed scenes, haze scenes, and object occlusion scenes. Here, each scene will cause a decrease in the detection accuracy of the camera, and their reasons and manifestations are also different. After classification, different optimization methods can be adopted according to the characteristics of different scenes.
[0014] Preferably, the overexposure scenario includes a sunlight illumination scenario or a reflection scenario with an illumination intensity greater than 2000 lux, the haze scenario includes weather conditions with a visibility lower than 1000 meters, and the occlusion scenario includes a view partially or completely occluded by an object. Here, by precisely defining these scenarios, the system can more accurately collect and classify failure samples, and at the same time, in the subsequent optimization process, it can accurately match different sensors to make up for the deficiencies of the camera under these specific conditions.
[0015] Preferably, the interpretability interpreter interprets the output information of the seventh layer of the convolutional neural network, and uses the Shapley value to indicate the distribution weight of the features of each pixel point on the image. Here, using the Shapley value for interpretation makes the decision-making process of the network more transparent. As an interpretation tool widely used in game theory, the Shapley value can accurately quantify the contribution degree of each pixel point to the final model prediction. In this way, the system can accurately reveal which image regions have an important impact on the decision-making of the convolutional neural network.
[0016] Preferably, the distribution weight is input into the convolutional neural network. After passing through the interpreter, an interpretable map of failure samples with Shapley values distributed is obtained. Here, using the Shapley value to generate an interpretable map makes the output result of the convolutional neural network no longer a "black box", but can intuitively show through an image which regions have a greater impact on the judgment of the convolutional neural network under specific scenarios. In the application of autonomous driving, it is very crucial to clearly understand why the camera fails to recognize in overexposure scenarios, haze scenarios, and object occlusion scenarios.
[0017] Preferably, the statistical sampling in S3 includes using the method of concentric circle sampling in overexposure scenarios and object occlusion scenarios, with the center at the point where the Shapley value is 0, and concentric circles spreading outward with the same radius. In haze weather, concentric circles are spread outward from the center of the picture with the same radius to calculate the average Shapley value of each concentric circle, and the changing rules of these Shapley values are statistically analyzed. Here, by starting from the point where the Shapley value is 0 and gradually expanding the concentric circles, it is possible to well capture the contribution of the features of different regions in the image to the network decision-making. This refined sampling method can reveal the specific impact of important feature regions on the recognition result better than simple random sampling, thereby helping to better understand the reasons for recognition failure.
[0018] Preferably, the changing rules of the Shapley value in S3 include that in the case of overexposure, the Shapley value changes slowly, in the occlusion scenario, the Shapley value mutates at a certain moment, and in the haze scenario, there is no obvious change in the Shapley value. Here, the precise analysis of the feature changes in different environmental scenarios provides very targeted guidance for the autonomous driving system when dealing with overexposure, occlusion, and haze environments.
[0019] Preferably, for the reason of the analysis sample failure in S4, in the overexposure scenario, it is due to the generation of halos; in the occlusion scenario, it is due to the occluder affecting the outline of the sample; in the haze scenario, it is because every part of the sample is covered by gray pixels. Here, by clarifying the specific reasons for the failure scenarios, the system can optimize the multi-layer structure of the convolutional network, avoiding the multi-layer convolution from calculating each image area one by one. Especially in the overexposure scenario, haze scenario, and object occlusion scenario, the convolutional layer can optimize the calculation amount and reduce the over-convolution of unimportant areas.
[0020] Preferably, the sensors include lidar, millimeter-wave radar, and ultrasonic radar. By analyzing the working principles, advantages, and disadvantages of these sensors, different sensors are selected for auxiliary recognition of the camera for the overexposure scenario, haze scenario, and object occlusion scenario through causal relationships. Here, selecting sensors targeted to assist the camera in recognition, rather than relying on the continuous input of all sensors, enables the system to reduce the calculation amount, improve the processing speed, and save computing resources at the same time.
[0021] Preferably, the detection box is matched with the detection box obtained by the convolutional neural network recognition through the Hungarian algorithm to optimize the recognition accuracy and compare the detection effects under different IOU values. Here, the Hungarian algorithm is a classic optimal matching algorithm widely used in multi-object tracking and object detection. The IOU value is a standard for measuring the accuracy of detecting corresponding objects in a specific dataset. Using the Hungarian algorithm to match the detection box with the box recognized by the convolutional neural network can effectively optimize the pairing between the two detection boxes and ensure the optimal matching result.
[0022] The beneficial effects of the present invention are as follows:
[0023] (1) The present invention uses multi-modal sensor fusion to assist the camera in recognition, selects the most suitable sensors for auxiliary recognition for overexposure, haze, and occlusion environmental conditions, thereby improving the recognition accuracy and robustness of the system in complex environments.
[0024] (2) The present invention provides targeted optimization for light, occlusion, and haze scenarios through refined analysis of failure samples, reducing the convolutional calculation of unimportant areas. This method enables the convolutional neural network to run more efficiently by reducing redundant calculations, reduces the computing power requirements, and maintains a high recognition accuracy at the same time. Description of the Drawings
[0025] Figure 1 Flowchart of an interpretability interpreter for an analysis and recalibration method of multi-modal assisted vehicle detection failure samples;
[0026] Figure 2The output of the interpretability interpreter in an overexposed environment;
[0027] Figure 3 The output of the interpretability interpreter in a haze environment;
[0028] Figure 4 The output of the interpretability explainer in an occluded environment. DETAILED DESCRIPTION
[0029] In order to better understand the above method, the following will combine the drawings of this specification with specific implementation examples to illustrate the detailed experimental results:
[0030] Step 1: Collect and classify failure samples
[0031] The perception module of the autonomous driving system collects a large number of failure samples in different driving environments through integrated cameras and multiple sensors such as lidar, millimeter-wave radar and ultrasonic radar. Failure samples refer to images or data in which the system fails to correctly identify the target in a specific scenario.
[0032] By driving the vehicle in different environments such as cities, villages, mountains, tunnels, etc., sensor data in various environments can be obtained, especially for some extreme weather or complex scenes such as overexposure, haze and occlusion.
[0033] The collected image data will undergo image preprocessing, including denoising, cropping, and formatting, to ensure data quality. Then the failed samples will be classified by the algorithm:
[0034] Overexposed scene: The light in the image is too strong, resulting in loss of details of the target object. This usually occurs in scenes with light intensity greater than 2000 lux.
[0035] Haze scene: The image is blurred or dark, affecting visibility. The typical feature is blurred details in the image. It usually occurs in an environment with visibility less than 1000 meters.
[0036] Object occlusion scenario: The target object is partially or completely blocked by other objects, causing the camera to be unable to recognize the blocked part.
[0037] The collected failure samples are annotated to mark the overexposed, blurred or blocked areas in the image, providing basic data for subsequent analysis.
[0038] Step 2: Customize the interpretability interpreter for Shapley value analysis
[0039] like Figure 1As shown in the figure, a kind of interpretability interpreter is designed, which is specifically used to analyze the output of the seventh layer of a convolutional neural network (CNN). The seventh layer of the CNN usually involves higher-level abstract features, which can reflect important target regions in the image.
[0040] This interpreter will use the Shapley value method to evaluate the contribution of each pixel to the final decision of the network. The Shapley value is a weighted method from game theory. By considering the impact of the addition of each feature on the model result, the Shapley value assigns a fair degree of contribution. The Shapley value measures the role of a pixel in different combinations of input features and gives the degree of its impact on the result. The interpreter will input the failed samples into the CNN, calculate the contribution degree of each pixel in the seventh layer of the convolutional neural network, and generate a Shapley value map. The Shapley value map intuitively shows the importance of different regions in the image to the recognition result.
[0041] Using the Shapley value map, it is possible to analyze which parts of the features in the image have a key impact on the target recognition result, especially in overexposure, haze, and occlusion scenarios, helping the system better understand the model error.
[0042] Step 3: Shapley value statistics and analysis of variation laws
[0043] Perform statistical sampling on the Shapley value map of the failed samples and analyze the variation laws of the Shapley values in different scenarios:
[0044] Overexposure scenario: As Figure 2 shown, use the concentric circle sampling method. Taking the point where the Shapley value is 0 as the center, and the concentric circles spreading outwards to statistically analyze the variation law of the Shapley value. In the overexposure scenario, the Shapley value increases slowly, indicating that the impact of overexposure on recognition is progressive.
[0045] Haze scenario: As Figure 3 shown, take the center of the image as the center and perform concentric circle sampling to analyze the change of the Shapley value. In the haze scenario, the Shapley value has no obvious change, indicating that the entire image is uniformly affected in the haze environment, and the overall image needs to be optimized.
[0046] Object occlusion scenario: As Figure 4 shown, also use the concentric circle sampling method. Taking the point where the Shapley value is 0 as the center, and the concentric circles spreading outwards to statistically analyze the variation law of the Shapley value. The variation law of the Shapley value shows a mutation, indicating that the occluder will suddenly affect the features of the sample, resulting in a decrease in recognition accuracy.
[0047] Step 4: Analyze the failure reasons and select appropriate sensors
[0048] According to the results of the Shapley value statistical analysis, analyze the reasons for failure in different scenarios, and select appropriate sensors to assist recognition according to the analysis results:
[0049] Overexposed scene: As Figure 2 shown, the Shapley value graph indicates that the influence of the overexposed area on the recognition result gradually increases. Due to the halo effect caused by overexposure, the camera cannot effectively recognize the target. At this time, the lidar can be used as an auxiliary sensor to provide accurate depth information and supplement the spatial information lost by the camera due to insufficient light.
[0050] Haze scene: As Figure 3 shown, the Shapley value is evenly distributed, indicating that the blur caused by haze affects the entire image. Using millimeter-wave radar can penetrate haze to provide accurate distance information, thereby improving the accuracy of image recognition.
[0051] Object occlusion scene: As Figure 4 shown, the mutation of the Shapley value indicates that object occlusion leads to the loss of target features. At this time, using ultrasonic radar can effectively detect nearby occlusions and make up for the lack of perception of the occlusion area by the camera.
[0052] Select a suitable sensor for auxiliary recognition according to the specific failure reasons of each scene.
[0053] Step Five: Point Cloud Processing and Reprojection of Sensor Data
[0054] Process the point cloud data collected by lidar, millimeter-wave radar, and ultrasonic radar sensors, and reproject the three-dimensional point cloud data into a two-dimensional image through geometric transformation.
[0055] Point cloud data is three-dimensional spatial data generated by sensors such as lidar, millimeter-wave radar, and ultrasonic radar, representing the positions and shapes of objects in the environment scanned by the sensors. The point cloud consists of a series of three-dimensional coordinates, and each coordinate represents a point on the surface of an object. For an autonomous driving system, point cloud data can provide depth information about the surrounding environment, that is, the distance of an object from the sensor.
[0056] Reprojection is the process of converting three-dimensional point cloud data into a two-dimensional image. In three-dimensional space, the data obtained by the sensor is based on depth information, but convolutional neural networks (CNNs) usually need to process two-dimensional image data. Therefore, the three-dimensional data must be "mapped" onto a two-dimensional plane.
[0057] Geometric transformation refers to the process of converting point cloud data in three-dimensional space to a two-dimensional image plane through a mathematical model. This usually includes operations such as rotation, translation, and scaling. Specifically, geometric transformation will convert the three-dimensional coordinates of the point cloud into a two-dimensional image coordinate system through the internal and external parameters of the camera. After reprojection, the system can fuse the depth information of different sensors with the camera image data.
[0058] The minimum adjacency matrix algorithm is used to project 3D point cloud data into 2D detection boxes. Through this process, the spatial information of the object can be more accurately mapped onto the image.
[0059] The minimum adjacency matrix algorithm is a commonly used graph theory algorithm, usually used to process and optimize multi-object matching problems. In the present invention, it is used to optimize the matching between point cloud data and the object bounding boxes in the image. The algorithm constructs an adjacency matrix, where each element of the matrix represents the distance or similarity between two detection boxes, and then selects the best match by calculating the minimum value.
[0060] The detection box is a key concept in the object detection task, used to identify the position of the detected object in the image. The detection box is usually represented by a rectangular box, containing two coordinates: the coordinates of the upper left corner and the coordinates of the lower right corner, enclosing the region of interest in the image. In a multi-sensor fusion autonomous driving system, sensors such as lidar, millimeter-wave radar, and cameras will generate corresponding detection boxes.
[0061] Step 6: Use the Hungarian algorithm for detection box matching and recalibration
[0062] The Hungarian algorithm is used to perform optimal matching on the detection boxes from different sensors. The algorithm ensures that each detection box is paired with the most relevant object by minimizing the matching cost, optimizing the matching accuracy. The Hungarian algorithm automatically selects the optimal matching path based on factors such as the similarity and distance between detection boxes, thus avoiding false matches and missed matches.
[0063] The accuracy is evaluated according to the IOU (Intersection over Union) value of each detection box, and the matching accuracy is analyzed. The IOU value is a commonly used metric to evaluate the overlapping degree of two detection boxes. The IOU value is the ratio of the area of the overlapping region of the two boxes to the area of their union region. The higher the IOU, the greater the overlapping degree of the two boxes. The matching results are adjusted through recalibration to ensure the accuracy of the final recognition result. The recognition accuracy is further optimized according to the IOU value evaluation in different scenarios.
[0064] The above-described embodiments only represent the implementation manners of the present invention, but should not be construed as limiting the scope of the invention patent. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A multi-modal assisted vehicle detection failure sample analysis and recalibration method, characterized in that: The following steps are involved: S1: Collect failure samples captured by the vehicle camera during daily driving and sort them out; S2: Customize an interpretability interpreter, use Shapley values for interpretation, and obtain an interpretability graph of failure samples with Shapley values distributed; S3: Use statistical sampling to count the Shapley values in the interpretability graph and derive the corresponding change pattern of the Shapley values; S4: Compare the statistical data of different scenarios and analyze the reasons for sample failure; S5: Collect basic information of sensors identified by vehicle cameras and analyze which failure samples each sensor is suitable for solving; S6: The point cloud data obtained by the corresponding sensor is projected onto the two-dimensional image through reprojection, and the corresponding detection frame is obtained according to the minimum adjacency matrix. These detection frames are matched with the detection frames obtained by the visible light camera through the convolutional neural network model, and the corresponding detection accuracy is calculated.
2. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 1, characterized in that: The failed samples described in S1 are sorted and classified into overexposure scenes, haze scenes, and object occlusion scenes.
3. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 2, characterized in that: The overexposed scene includes a sunlight scene or a reflection scene with a light intensity greater than 2000 lux, the haze scene includes a weather condition with visibility less than 1000 meters, and the blocked scene includes a view that is partially or completely blocked by an object.
4. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 1, characterized in that: The interpretability interpreter interprets the output information of the seventh layer of the convolutional neural network, and uses the Shapley value to indicate the distribution weight of the convolutional neural network for the features of each pixel point on the image.
5. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 4, characterized in that: The distribution weights are input into a convolutional neural network, and after passing through an interpreter, an interpretability graph of failure samples with distributed Shapley values is obtained.
6. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 1, characterized in that: The statistical sampling described in S3 includes the use of concentric circle sampling in overexposure and object occlusion scenarios, with the Shapley value of 0 as the center and concentric circles spreading outward with the same radius. In haze weather, the center of the picture is used as the center and the concentric circles spread outward with the same radius to calculate the average Shapley value of each concentric circle, and the change pattern of these Shapley values is statistically analyzed.
7. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 1, characterized in that: The variation law of the Shapley value described in S3 includes that in the case of overexposure, the variation law of the Shapley value is a slow increase, in the case of occlusion, the Shapley value suddenly changes at a certain moment, and in the case of haze, the Shapley value has no detailed change.
8. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 1, characterized in that: The reasons for the failure of the analysis samples described in S4 are: in the overexposure scene, it is due to the generation of halos; in the occlusion scene, it is due to the occlusion affecting the outline of the sample; in the haze scene, it is due to every part of the sample being covered by gray pixels.
9. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 1, characterized in that: The sensors include laser radar, millimeter wave radar and ultrasonic radar. The working principles and advantages and disadvantages of these sensors are analyzed, and different sensors are selected for overexposure scenes, haze scenes, and object occlusion scenes through cause and effect relationships to assist the camera in identification.
10. The method for analyzing and recalibrating a vehicle detection failure sample assisted by multi-modality according to claim 1, characterized in that: The detection frame is matched with the detection frame obtained by convolutional neural network recognition through the Hungarian algorithm to optimize the recognition accuracy and compare the detection effects under different IOU values.