An evaluation method and system for simulation confidence of a high-level automatic driving system
By decomposing the simulation confidence level of an autonomous driving system into multiple levels and establishing a cross-level interactive verification mechanism, the problems of multimodal uncertainty and insufficient simulation-real-world calibration in the evaluation of simulation confidence level of high-level autonomous driving systems are solved. This achieves interpretability and reproducibility of simulation results and improves the efficiency of system development and regulatory compliance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA AUTOMOTIVE ENG RES INST
- Filing Date
- 2025-10-27
- Publication Date
- 2026-06-02
AI Technical Summary
Current confidence evaluation of high-level autonomous driving system simulations suffers from problems such as a single confidence expression, lack of quantification of multimodal uncertainty, insufficient calibration between simulation and reality, and lack of interpretable evidence chains. These issues make it difficult to directly translate simulation results into credibility in real-world deployment scenarios.
A hierarchical confidence evaluation model is constructed, which decomposes the simulation confidence of autonomous driving system into scene layer, sensor layer, actuator/dynamics layer and algorithm decision layer. Through hierarchical data collection and feature extraction, the sub-confidence of each layer is calculated, and a cross-layer interactive verification mechanism is established to locate the root cause level of simulation confidence loss and output interpretable simulation confidence evaluation results.
It improves the interpretability and reproducibility of simulation confidence, enhances the efficiency of high-level autonomous driving systems in the development, testing and regulatory compliance process, and achieves closed-loop calibration between simulation results and real-world scenarios by accurately identifying the source of problems as distortions in perception input or defects in decision-making logic.
Smart Images

Figure CN121388467B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and specifically to a method and system for evaluating the confidence level of a high-level autonomous driving system simulation. Background Technology
[0002] As high-level autonomous driving systems (SAE Level 3 / 4) gradually move towards practical application, simulation platforms play a crucial role in system evaluation and safety verification. However, the industry still has shortcomings in the evaluation and expression of "simulation confidence." Specific issues include:
[0003] (1) The confidence level is expressed in a single way and lacks multimodal uncertainty quantification: Existing methods mostly use a single indicator (such as trajectory deviation, collision rate, etc.) or a simple threshold judgment to evaluate system performance, which is difficult to reflect the comprehensive performance and risk of perception, decision-making and control links in different scenarios; there is a lack of unified modeling for the uncertainty of different modal outputs (trajectory prediction, path planning, etc.), which makes it difficult to quantify the transmission of uncertainty across modalities and cannot give a comparable and reproducible confidence score.
[0004] (2) Insufficient confidence calibration between simulation and reality: Existing platforms often evaluate based on a single simulation environment, lacking a systematic simulation-reality alignment mechanism. In particular, deviations in complex urban scenarios, extreme weather, sensor noise characteristics, and sensor configuration differences have not been effectively resolved. The confidence output lacks calibration in the real domain, making it difficult to directly convert the credibility of simulation results into credibility in real deployment scenarios. It is also difficult to establish the correspondence between "simulation credibility" and "real system performance" during the review stage.
[0005] (3) Lack of an explanatory chain of evidence and causal output: Most systems only provide a comprehensive score or statistical indicators, lacking traceability and explanatory explanation of the source of uncertainty. It is difficult for the reviewing agency to trace the cause of the confidence level and the chain of evidence layer by layer, which affects the compliance assessment and the verifiability of the evidence.
[0006] The aforementioned technical issues collectively constitute the technical barriers that urgently need to be overcome in the field of confidence evaluation for high-level autonomous driving system simulations. Summary of the Invention
[0007] The present invention aims to provide a method and system for evaluating the confidence level of simulations of high-level autonomous driving systems, which can improve the interpretability and reproducibility of simulation confidence, thereby improving the efficiency of high-level autonomous driving systems in the development, testing and regulatory compliance process.
[0008] To achieve the above objectives, the present invention provides the following basic solution.
[0009] Option 1
[0010] A method for evaluating the confidence level of a high-level autonomous driving system simulation includes the following steps:
[0011] S1. Construct a confidence-level hierarchical evaluation model, decomposing the simulation confidence of the autonomous driving system into four levels of sub-confidence: scene layer, sensor layer, actuator / dynamics layer, and algorithm decision layer.
[0012] S2, Based on the hierarchical evaluation model, hierarchical data collection and feature extraction are performed on the operation process of the autonomous driving system in a specific simulation scenario;
[0013] S3, based on the extracted features, calculate the sub-confidence of each level;
[0014] S4. Establish a cross-layer interactive verification mechanism to locate the root cause level of the overall simulation confidence loss based on the correlation between sub-confidences at each level.
[0015] S5, output an interpretable simulation confidence evaluation result containing the sub-confidence of each level and the root cause level location information.
[0016] Option 2
[0017] An evaluation system for the confidence level of a high-level autonomous driving system simulation includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the evaluation method for the confidence level of a high-level autonomous driving system simulation as described in Scheme 1.
[0018] The working principle and advantages of this invention are as follows:
[0019] This invention provides a method and system for evaluating the confidence level of high-level autonomous driving system simulations, which can improve the interpretability and reproducibility of simulation confidence levels, thereby enhancing the efficiency of high-level autonomous driving systems in development, testing, and regulatory compliance processes. The key points are:
[0020] This approach constructs a novel framework for systematically solving the problem of simulation confidence evaluation, rather than a partial improvement on existing methods.
[0021] Existing technologies typically rely on a single metric to evaluate the final output of an autonomous driving system (such as trajectory deviation), which fails to reveal the underlying causes of performance defects. This solution creatively decomposes simulation confidence into four closely related yet independently quantifiable levels: scene, sensors, actuators / dynamics, and algorithm decisions, and establishes a cross-level interactive verification mechanism.
[0022] Notably, this solution breaks away from the traditional mindset of "black box" or "grey box" evaluation, delving into the evaluation perspective from terminal performance to the internal information flow and causal links of the system. For example, when the system detects an anomaly at the algorithm decision-making layer, its interactive verification mechanism can automatically trigger a collaborative analysis of the degree of data distortion at the sensor layer and the consistency of the dynamic model response, thereby accurately determining whether the problem stems from distortion of the perceived input, inaccuracy of the vehicle model, or defects in the decision-making logic itself. This dynamic, causal reasoning-based root cause localization capability cannot be achieved by simply weighting and averaging the confidence levels of each layer. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the method flow of the method and system embodiment one for evaluating the confidence level of a high-level autonomous driving system simulation according to the present invention. Detailed Implementation
[0024] The following detailed explanation illustrates the specific implementation methods:
[0025] Example 1
[0026] The basic implementation examples are as follows: Figure 1 As shown: A method for evaluating the confidence level of a high-level autonomous driving system simulation includes the following steps:
[0027] S0.1, Edge Scene Filtering Steps:
[0028] Based on real driving data, we construct a realism model to describe the rationality of scenario parameters and a hazard model to describe the potential risks of the scenario.
[0029] Based on testing requirements, the contribution weights of realism and risk in the joint scoring are dynamically adjusted; and based on the realism model and risk model, the joint scoring results of various scenarios in the scenario library are calculated.
[0030] Based on the joint scoring results, edge scenarios with high realism and high risk are selected from the scenario library as specific simulation scenarios for simulation testing.
[0031] Specifically, the real driving data on which the realism model is based includes extreme value distribution data of vehicle motion; the input parameters for constructing the hazard model include at least collision type, relative speed, and collision angle.
[0032] In this embodiment, the steps for constructing the realism model specifically include the following operations:
[0033] Extract extreme value distribution data of key driving parameters from real driving databases;
[0034] For the simulation scenario to be evaluated, the realism of its specific parameters ( The parameter value is quantified by calculating the probability density function value corresponding to the extreme value distribution.
[0035] The steps for constructing the hazard model specifically include the following operations:
[0036] Based on preset collision type weights and the relative speeds of traffic participants in the scene ( The basic hazard level (BN) is calculated using the collision angle (θ) as the core input parameters. );
[0037] The calculation steps for the joint scoring result specifically include the following operations:
[0038] The joint score (S) is composed of the truthiness ( ) and risk level ( The weighted sum of the results is calculated as follows: ,in, and The contribution weight is determined by the data collected based on different testing objectives and verification scenarios. and The numerical value can be used to prioritize the selection of high-realism, high-risk scenarios for simulation testing, so that testing resources can be focused on high-value edge cases.
[0039] S0.2, Data preprocessing steps:
[0040] The system automatically accesses raw data collected from real road tests via standardized data interfaces (such as CAN bus interface, Ethernet interface, and video data stream interface). In this embodiment, the raw data specifically includes: vehicle's own bus data (such as vehicle speed, acceleration, yaw rate, steering angle, braking pressure, throttle and brake pedal opening, etc.), pose data output by a high-precision integrated navigation system (such as GNSS / IMU integrated navigation system) (such as latitude, longitude, elevation, attitude angle, etc.), and environmental perception data synchronously collected by onboard sensors (cameras, lidar, millimeter-wave radar).
[0041] The raw data is parsed, time-synchronized, coordinate-system unified, and format-standardized, and a reference time-series signal file (containing vehicle trajectory, speed curve, acceleration curve, and other data) is automatically generated for use in the bidirectional differential tracing step.
[0042] The time synchronization operation includes unifying the timestamps of all data streams to a high-precision master clock, ensuring that vehicle attitude, sensor data, and control commands are aligned with millisecond-level precision. The coordinate system operation includes converting all sensing data, positioning data, and vehicle data to a unified coordinate system (e.g., a vehicle coordinate system or a global geodetic coordinate system). The format standardization process includes removing obviously abnormal and invalid data, standardizing and encoding the data, and finally outputting a structured reference time-series signal file containing time sequences.
[0043] S1. Construct a confidence-based hierarchical evaluation model, decomposing the simulation confidence of the autonomous driving system into four levels of sub-confidence: scene layer, sensor layer, actuator / dynamics layer, and algorithm decision layer.
[0044] S2, Based on the hierarchical evaluation model, hierarchical data collection and feature extraction are performed on the operation process of the autonomous driving system in a specific simulation scenario.
[0045] Specifically, the autonomous driving system to be evaluated is placed on a simulation platform to reproduce the simulation scenario corresponding to the baseline timing signal file. The simulation is run, and data from four levels are collected simultaneously.
[0046] S3, based on the extracted features, calculate the sub-confidence of each level.
[0047] The calculation of sub-confidence scores at each level includes the following sub-steps:
[0048] The scene layer confidence is obtained by calculating the fidelity of static elements (such as the size, position, and color of traffic signs) and dynamic elements (such as the behavior model of surrounding traffic participants) of the simulated scene relative to the real-world scene database in terms of geometric properties (such as positional deviation) and physical properties (such as material reflectivity and kinematic model parameters).
[0049] For example, structural similarity indices can be used to compare texture details between simulated and real video images, or to calculate differences in traffic flow density statistics.
[0050] The confidence level of the sensor layer is obtained by calculating the similarity and distortion of the feature distribution between simulated sensor data and real sensor data in the same scenario.
[0051] Specifically, the simulated sensor data refers to the raw data (such as image point clouds and radar signals) generated by virtual sensors (such as virtual cameras and LiDAR) of the simulation platform in the simulated scenario. The real sensor data refers to the data of real sensors in the same pose extracted from the reference time-series signal file.
[0052] In this embodiment, simulated and real data are processed separately by a feature extraction network (such as a pre-trained convolutional neural network), and their distribution similarity (such as using Frèchet distance) or distortion measure is calculated in the feature space to quantify the confidence of the sensor layer.
[0053] The confidence level of the actuator / dynamics layer is obtained by calculating the consistency between the simulated vehicle's response to control commands and the dynamic response of the real vehicle in the time and frequency domains.
[0054] Specifically, in the actuator / dynamics layer confidence calculation process, control commands (such as steering, throttle, and braking commands) of the same sequence as those recorded in real road tests are first sent to the simulated vehicle model. Then, the dynamic response output by the simulated vehicle (such as longitudinal / lateral acceleration and yaw rate) is compared with the response data recorded on the real vehicle's bus. This comparison specifically involves: analyzing response delay and overshoot in the time domain; analyzing the consistency of frequency response characteristics through Fourier transform in the frequency domain; and finally, comprehensively calculating the actuator / dynamics layer confidence.
[0055] The confidence level of the algorithm's decision layer is obtained by comparing the consistency of the decision logic output of the autonomous driving system in both simulated and real environments when facing the same driving scenario.
[0056] Specifically, the same scenario-driven signals are injected into the autonomous driving system under test in both simulation and real-world environments. The logical outputs at key decision points are recorded and compared.
[0057] For example, for the same vehicle cutting in, do the behavioral predictions made in both the simulation and real-vehicle systems classify it as a "cutting in"? Do the path planning systems generate similar avoidance trajectories? Do the collision assessment systems output consistent time-to-collision (TTC) warnings? By calculating the consistency ratio of these decision logic outputs, the confidence level of the algorithm's decision layer can be obtained.
[0058] S4. Establish a cross-layer interactive verification mechanism to locate the root cause level of the overall simulation confidence loss based on the correlation between sub-confidences at each level.
[0059] The cross-layer interactive verification mechanism specifically involves: when the confidence level of the algorithm decision layer is lower than the threshold, the confidence levels of the sensor layer and the actuator / dynamics layer are analyzed to determine whether the problem is caused by perceptual data distortion or vehicle model inaccuracy.
[0060] Furthermore, the source tracing analysis specifically includes: a simulation-real-scene bidirectional differential source tracing step:
[0061] The key timing signals output from the simulation test are compared frame by frame with the reference timing signals from the actual road test to calculate the forward difference; for example, the final vehicle trajectory, speed and other timing signals obtained from the simulation are compared frame by frame with the reference timing signals to calculate the forward difference.
[0062] When the forward difference exceeds a preset threshold, the reverse tracing mechanism is activated; for example, assuming the threshold for trajectory error is set to 0.2 meters, when the average error over a certain period of time is found to reach 0.21 meters, the reverse tracing mechanism is triggered.
[0063] The reverse tracing mechanism calls the hierarchical evaluation model from the bottom up, analyzing the data associated with the key time-series signals layer by layer to locate the root cause data distributions leading to the forward discrepancies. The key time-series signals include one or more of vehicle trajectory, speed, and acceleration. Specifically, it includes the following operations:
[0064] First, check whether the decision logic output of the algorithm's decision layer is consistent with the output in the real environment during different time periods;
[0065] If there is a discrepancy, the degree of distortion between the output data of the sensor layer and the real sensor data during the time period of the difference is traced back to analyze; for example, whether the difference between the simulated lidar point cloud and the real point cloud in terms of target contour and density during that time period led to the misjudgment of the perceived target attributes.
[0066] Simultaneously, the consistency between the actuator / dynamics layer's response to control commands and the real vehicle's dynamic response during the time difference is analyzed in parallel. For example, it analyzes whether the simulated vehicle's deceleration response to braking commands is slower than that of the real vehicle, thus causing trajectory deviation. Through this correlation analysis, this approach can accurately pinpoint the root cause to "approximately xx% distortion in target contour extraction by the sensor layer," rather than simply reporting "trajectory inconsistency."
[0067] S5, output an interpretable simulation confidence evaluation result containing the sub-confidence scores of each level (e.g., expressed as numerical values of 0-1 or percentages) and root-level location information.
[0068] The interpretable simulation confidence evaluation result is an evidence chain report, which includes at least: the calculation process of each level of sub-confidence, the data samples used for calculation, the analysis path of cross-level interactive verification (e.g., the specific differences that trigger tracing, the corresponding bidirectional differential tracing steps), and the final root cause level localization conclusion (e.g., "the sensor layer has about xx% distortion in target contour extraction").
[0069] The data samples used for calculation specifically refer to representative key data samples used for comparison. For example, to illustrate the confidence level of the sensor layer, a comparison chart of the simulated lidar point cloud screenshot and the real lidar point cloud screenshot at the moment of maximum difference is attached; to illustrate the confidence level of the dynamics layer, a comparison chart of the simulated and real yaw rate response curves is attached.
[0070] This embodiment also provides an evaluation system for the confidence level of a high-level autonomous driving system simulation, including a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the evaluation method for the confidence level of a high-level autonomous driving system simulation as described above.
[0071] This embodiment provides a method and system for evaluating the simulation confidence of a high-level autonomous driving system. Through cross-modal uncertainty modeling and fusion, simulation-reality adaptive domain alignment and confidence calibration, and interpretable scenario-level output, it can improve the interpretability and reproducibility of simulation confidence, thereby enhancing the efficiency of high-level autonomous driving systems in development, testing, and regulatory compliance processes. The key points are:
[0072] First, this solution constructs a novel framework for systematically addressing the challenge of simulation confidence evaluation, rather than a partial improvement on existing methods. This solution creatively decomposes simulation confidence into four closely related yet independently quantifiable levels: scene, sensors, actuators / dynamics, and algorithmic decision-making. It also establishes a cross-level interactive verification mechanism, capable of accurately determining whether the problem stems from distortions in perceptual input, inaccuracies in the vehicle model, or defects in the decision-making logic itself.
[0073] Secondly, addressing the subjectivity and inefficiency of corner case construction in autonomous driving simulations, this solution provides an objective, quantitative, and efficient screening method. Currently, the industry generally relies on expert experience or random generation to construct corner cases, resulting in high subjectivity and low efficiency. This solution proposes a joint "realism-hazard" scoring mechanism, transforming the scene selection process from relying on qualitative judgment to mathematical calculations based on extreme value distributions of real driving data and multi-factor collision risk models. Furthermore, it introduces a dynamically adjustable contribution ratio mechanism, enabling test resources to intelligently focus on "high-value" corner cases that both conform to real-world physical laws (high realism) and contain high potential risks (high hazard), rather than scenarios that are dangerous but almost impossible in reality, or common but with extremely low hazard, effectively resolving the contradiction between testing efficiency and test coverage.
[0074] Third, by constructing a two-way differential tracing mechanism based on real road test data, a closed-loop calibration and continuous improvement of simulation evaluation and real-world performance is achieved. Traditional methods often remain at the unidirectional, result-level comparison stage, that is, using real data to verify simulation results. However, once discrepancies occur, there is a lack of systematic tools to trace the root cause of the discrepancies. The non-obviousness of this solution lies in the organic combination of "forward differential" and "reverse tracing," forming a complete diagnostic closed loop. The forward differential mechanism is responsible for performing high-precision frame-by-frame comparisons on key time-series signals such as vehicle trajectory and speed, quantifying inconsistencies. When inconsistencies exceed the tolerance limit, the reverse tracing mechanism is intelligently activated, using the constructed hierarchical evaluation model to trace the original data distribution or model deviation that led to the final difference from the bottom up. Combined with the automated data processing flow realized by standardized data interfaces, the simulation model can be continuously iterated and accurately optimized based on real road test evidence, greatly enhancing the credibility of simulation results and their industrial practical value.
[0075] Example 2
[0076] An evaluation method for the simulation confidence of a high-level autonomous driving system is proposed, with the following adjustments made based on Example 1.
[0077] In S0.1, the edge scene screening step, the joint scoring mechanism incorporates: regional traffic rules and driving habit factors.
[0078] Specifically, in constructing the hazard model, firstly, a preset collision type weight and the relative speeds of traffic participants in the scenario are used (…). The basic hazard level is calculated using the collision angle (θ) and the collision angle (θ) as the core input parameters.
[0079] Then, a scenario-based regional factor (L) is introduced. This factor is obtained by weighting regional traffic rules (such as right-turn permission at red lights), driving habit factors (such as the frequency of cutting in line), and road infrastructure conditions. This factor is used to correct the basic risk level (e.g., by multiplying by L as a correction coefficient) to obtain the final risk level. ).
[0080] In this embodiment, L = f(R, B, I).
[0081] Here, R is a traffic rule factor determined based on regional traffic rules. Its value is determined by the clarity and specificity of traffic rules in the target region; the more ambiguous the rules or the existence of special right-of-way regulations, the higher the R value. B is a driving habit factor determined based on driving behavior habits. Its value is obtained through statistical analysis of real driving data in the target region (such as calculating the cutting-in rate through a cutting-in frequency model), used to quantify the aggressiveness of driving behavior and the rate of rule compliance in that region. I is an infrastructure factor determined based on road infrastructure. Its value is based on a comprehensive evaluation of road lane width, shoulder conditions, and traffic sign density parameters in the target region. A higher I value indicates safer or more complete infrastructure.
[0082] This embodiment provides a method and system for evaluating the confidence level of a high-level autonomous driving system simulation. By embedding regional factors in the construction of the hazard model to reflect local traffic characteristics, the simulation scenario can be made closer to the actual operation scenario of the vehicle.
[0083] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent.
Claims
1. A method for evaluating the confidence level of a high-level autonomous driving system simulation, characterized in that, Includes the following steps: S1. Construct a confidence-level hierarchical evaluation model, decomposing the simulation confidence of the autonomous driving system into four levels of sub-confidence: scene layer, sensor layer, actuator / dynamics layer, and algorithm decision layer. S2, Based on the hierarchical evaluation model, hierarchical data collection and feature extraction are performed on the operation process of the autonomous driving system in a specific simulation scenario; S3, based on the extracted features, calculate the sub-confidence of each level, namely the sub-confidence of the scene layer, the sub-confidence of the sensor layer, the sub-confidence of the actuator / dynamics layer, and the sub-confidence of the algorithm decision layer; S4. Establish a cross-layer interactive verification mechanism to locate the root cause level of the overall simulation confidence loss based on the correlation between sub-confidences at each level. S5, output an interpretable simulation confidence evaluation result containing the sub-confidence of each level and the root cause level location information.
2. The method for evaluating the confidence level of a high-level autonomous driving system simulation according to claim 1, characterized in that, The calculation of sub-confidence scores at each level includes the following sub-steps: The sub-confidence of the scene layer is obtained by calculating the fidelity of the static and dynamic elements of the simulation scene relative to the real-world scene database in terms of geometric and physical properties. The sub-confidence of the sensor layer is obtained by calculating the similarity and distortion of the feature distribution between simulated sensor data and real sensor data in the same scenario; The sub-confidence of the actuator / dynamics layer is obtained by calculating the consistency between the simulated vehicle's response to control commands and the real vehicle's dynamic response in the time and frequency domains. The sub-confidence of the algorithm's decision layer is obtained by comparing the consistency of the decision logic output of the autonomous driving system in both simulated and real environments when facing the same driving scenario.
3. The method for evaluating the confidence level of a high-level autonomous driving system simulation according to claim 1, characterized in that, Before constructing the confidence level hierarchical evaluation model, an edge scene screening step is also included: Based on real driving data, we construct a realism model to describe the rationality of scenario parameters and a hazard model to describe the potential risks of the scenario. Based on testing requirements, the contribution weights of realism and risk in the joint scoring are dynamically adjusted; and based on the realism model and risk model, the joint scoring results of various scenarios in the scenario library are calculated. Based on the joint scoring results, edge scenarios with high realism and high risk are selected from the scenario library as specific simulation scenarios for simulation testing.
4. The method for evaluating the confidence level of a high-level autonomous driving system simulation according to claim 3, characterized in that, The real driving data on which the realism model is constructed includes extreme value distribution data of vehicle motion; the input parameters for constructing the hazard model include at least collision type, relative speed and collision angle.
5. The method for evaluating the confidence level of a high-level autonomous driving system simulation according to claim 1, characterized in that, The cross-layer interactive verification mechanism specifically involves: when the sub-confidence of the algorithm decision layer is lower than the threshold, the sub-confidence of the sensor layer and the sub-confidence of the actuator / dynamics layer are analyzed to determine whether the problem is caused by perceptual data distortion or vehicle model inaccuracy.
6. The method for evaluating the confidence level of a high-level autonomous driving system simulation according to claim 5, characterized in that, The source tracing analysis specifically includes: a simulation-real-scene two-way differential source tracing step: The key timing signals output from the simulation test are compared frame by frame with the reference timing signals from the actual road test to calculate the forward difference; When the forward difference exceeds a preset threshold, the reverse tracing mechanism is activated; The reverse tracing mechanism calls the hierarchical evaluation model from the bottom up, analyzes the data associated with the key time series signal layer by layer, and locates the root data distribution that causes the forward difference.
7. The method for evaluating the confidence level of a high-level autonomous driving system simulation according to claim 6, characterized in that, The key timing signals include one or more of the following: vehicle trajectory, speed, and acceleration.
8. The method for evaluating the confidence level of a high-level autonomous driving system simulation according to claim 6, characterized in that, It also includes data preprocessing steps: Through standardized data interfaces, it automatically accesses raw data collected from real road tests; The raw data is parsed and formatted to automatically generate a reference timing signal file for the actual road test, which can be used in the bidirectional differential tracing step.
9. The method for evaluating the confidence level of a high-level autonomous driving system simulation according to claim 1, characterized in that, The interpretable simulation confidence evaluation result is an evidence chain report, which includes at least: the calculation process of sub-confidence at each level, the data samples used for calculation, the analysis path of cross-level interactive verification, and the final root cause level localization conclusion.
10. An evaluation system for the confidence level of a high-level autonomous driving system simulation, characterized in that, It includes a processor and a memory, the memory storing a computer program that, when executed by the processor, implements a method for evaluating the confidence level of a high-level autonomous driving system simulation as described in any one of claims 1-9.