Method for evaluating an autonomous driving algorithm and related device
Patent Information
- Application Number
- CN202610722517.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]然而,经发明人的长期研究发现,相关技术中对自动驾驶算法的评测准确性不佳
[0010]本实施例通过将场景感知的指标特征与表现差距特征进行拼接,得到第一拼接特征矩阵;对第一拼接特征矩阵施加自注意力,以根据待测指标与相应的人类驾驶行为之间的差距,确定待测指标对应的特征的注意力权重,得到差距感知的待测评指标特征,避免因单一短板而影响模型对算法整体性能的公平评价。
Smart Images

Figure CN122839010A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent evaluation technology, and in particular to an evaluation method and related apparatus for autonomous driving algorithms. Background Technology
[0002] In related technologies, it is necessary to evaluate autonomous driving algorithms in order to assess their capabilities.
[0003] However, the inventors' long-term research revealed that the accuracy of evaluation of autonomous driving algorithms in related technologies is poor. Summary of the Invention
[0004] This application provides a method and related apparatus for evaluating autonomous driving algorithms, aiming to improve the accuracy of autonomous driving algorithm evaluation.
[0005] In a first aspect, embodiments of this application provide an evaluation method for an autonomous driving algorithm, comprising: acquiring scene features, evaluation index features, and performance gap features, wherein the scene features are features corresponding to the current driving scenario, the evaluation index features include features corresponding to multiple evaluation indicators, the multiple evaluation indicators are indicators corresponding to the autonomous driving algorithm in the current driving scenario, and the performance gap features are used to represent the gap between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior; fusing the scene features, evaluation index features, and performance gap features to obtain fused features, wherein the weights of the evaluation index features in the fused features are dynamically correlated with the scene features and the performance gap features; and based on the fused features, determining key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds, wherein the key evaluation indicators and the corresponding indicator thresholds are used to evaluate the capabilities of the autonomous driving algorithm in the current driving scenario.
[0006] In this embodiment, scene features, evaluation index features, and performance gap features are obtained. Scene features are the features corresponding to the current driving scenario. Evaluation index features include features corresponding to multiple evaluation indicators, which are the indicators corresponding to the autonomous driving algorithm in the current driving scenario. Performance gap features are used to represent the gap between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior. Scene features, evaluation index features, and performance gap features are fused to obtain fused features. The weights of the evaluation index features in the fused features are dynamically correlated with the scene features and performance gap features. Based on the fused features, key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds are determined. The key evaluation indicators and the corresponding indicator thresholds are used to evaluate the capabilities of the autonomous driving algorithm in the current driving scenario. By combining the gap between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior to obtain fused features to determine key evaluation indicators and the corresponding indicator thresholds, the capabilities of the autonomous driving algorithm in the current driving scenario can be evaluated, which is beneficial to improving the accuracy of autonomous driving algorithm evaluation.
[0007] In one possible implementation, the scene features include multiple elements. The process involves fusing scene features, evaluation metric features, and performance gap features to obtain fused features. This includes: determining the similarity between the feature corresponding to each evaluation metric and multiple scene features, and using the similarity as the weight of the evaluation metric feature; fusing the feature corresponding to each metric with multiple scene features based on the similarity between the feature corresponding to each evaluation metric and multiple scene features to obtain the fused metric feature for each metric; concatenating the fused metric features corresponding to multiple metrics to obtain scene-aware metric features; fusing the scene-aware metric features with performance gap features to obtain gap-aware evaluation metric features; and determining the fused features based on the gap-aware evaluation metric features.
[0008] This embodiment determines the similarity between the feature corresponding to each test indicator and multiple scene features, and uses the similarity of the feature corresponding to each test indicator as the weight of the test indicator feature. Based on the similarity between the feature corresponding to each test indicator and multiple scene features, the feature corresponding to each indicator is fused with multiple scene features to obtain the fused indicator feature corresponding to each indicator. The fused indicator features corresponding to multiple indicators are concatenated to obtain the scene-aware indicator features. The scene-aware indicator features are fused with the performance gap features to obtain the gap-aware test indicator features. Determining the fused features based on the gap-aware test indicator features can enhance the test indicator features, thereby enhancing the fused features and improving the accuracy of autonomous driving algorithm evaluation.
[0009] In one possible implementation, the scene-aware indicator features and performance gap features are fused to obtain gap-aware evaluation indicator features, including: concatenating the scene-aware indicator features and performance gap features to obtain a first concatenated feature matrix; applying self-attention to the first concatenated feature matrix to determine the attention weight of the features corresponding to the test indicator based on the gap between the test indicator and the corresponding human driving behavior, thereby obtaining gap-aware evaluation indicator features.
[0010] This embodiment concatenates the scene-aware indicator features with the performance gap features to obtain a first concatenated feature matrix. Self-attention is applied to the first concatenated feature matrix to determine the attention weight of the features corresponding to the indicator to be tested based on the gap between the indicator to be tested and the corresponding human driving behavior, thereby obtaining the gap-aware evaluation indicator features and avoiding the impact of a single weakness on the fair evaluation of the overall performance of the algorithm.
[0011] In one possible implementation, determining the fusion feature based on the gap-aware evaluation indicator features includes: concatenating scene features with performance gap features to obtain a second concatenated feature matrix; obtaining preset weights and bias parameters, calculating the product between the preset weights and the second concatenated feature matrix, and superimposing the product with the bias parameters to obtain a superimposed result; calculating a gating value by applying an activation function to the superimposed result; fusing the scene features with the performance gap features using the gating value to obtain gap-aware scene features; and concatenating the gap-aware scene features with the gap-aware evaluation indicator features to obtain the fusion feature.
[0012] This embodiment obtains a second concatenated feature matrix by concatenating scene features and performance gap features; it acquires preset weights and bias parameters, calculates the product between the preset weights and the second concatenated feature matrix, and superimposes the product with the bias parameters to obtain a superimposed result; it then calculates a gating value by applying an activation function to the superimposed result; it uses the gating value to fuse scene features and performance gap features to obtain gap-aware scene features; and it concatenates the gap-aware scene features with gap-aware evaluation indicator features to obtain fused features. This effectively avoids interference from individual abnormal gap data on scene understanding.
[0013] In one possible implementation, the indicator determination model, based on fusion features, determines key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds. This includes: determining a first probability distribution of multiple indicators based on fusion features, where the first probability distribution represents the probability value of each indicator in the current driving scenario; determining key evaluation indicators among the multiple indicators based on the first probability distribution, where the key evaluation indicators are a subset of the multiple indicators, and the probability value of the key evaluation indicators is not lower than the probability value of any other indicator; and determining the indicator thresholds corresponding to the key evaluation indicators.
[0014] This embodiment uses an indicator determination model based on fusion features to determine the first probability distribution of multiple indicators. The first probability distribution represents the probability value of each indicator in the current driving scenario. Based on the first probability distribution of multiple indicators, key evaluation indicators are determined among the multiple indicators. Key evaluation indicators are some of the multiple indicators, and the probability value of the key evaluation indicator is not lower than the probability value of any other indicator. The indicator threshold corresponding to the key evaluation indicator is determined, which can improve the interpretability of the selection of key evaluation indicators.
[0015] In one possible implementation, determining the threshold corresponding to the key evaluation indicator includes: for each key evaluation indicator, determining a second probability distribution of multiple thresholds corresponding to the key evaluation indicator, the second probability distribution representing the probability value corresponding to each of the multiple thresholds in the current driving scenario; and determining the threshold corresponding to the key evaluation indicator based on the second probability distribution of the multiple thresholds corresponding to the key evaluation indicator, wherein the threshold is a subset of the multiple thresholds, and the probability value corresponding to the threshold is not lower than the probability value corresponding to any threshold.
[0016] This embodiment determines a second probability distribution of multiple thresholds corresponding to each key evaluation indicator. The second probability distribution represents the probability value of each threshold in the current driving scenario. Based on the second probability distribution of multiple thresholds corresponding to the key evaluation indicator, the indicator threshold corresponding to the key evaluation indicator is determined. The indicator threshold is a partial threshold among multiple thresholds, and the probability value corresponding to the indicator threshold is not lower than the probability value corresponding to any threshold. This can improve the interpretability of the indicator threshold selection.
[0017] Secondly, embodiments of this application provide an evaluation device for an autonomous driving algorithm, comprising: an acquisition module, configured to acquire scene features, evaluation index features, and performance gap features, wherein the scene features are features corresponding to the current driving scenario, the evaluation index features include features corresponding to multiple evaluation indicators, the multiple evaluation indicators are indicators corresponding to the autonomous driving algorithm in the current driving scenario, and the performance gap features are used to represent the gap between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior; a fusion module, configured to fuse the scene features, evaluation index features, and performance gap features to acquire fused features, wherein the weights of the evaluation index features in the fused features are dynamically correlated with the scene features and the performance gap features; and an evaluation module, configured to determine key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds based on the fused features, wherein the key evaluation indicators and the corresponding indicator thresholds are used to evaluate the capabilities of the autonomous driving algorithm in the current driving scenario.
[0018] Thirdly, embodiments of this application propose an electronic device, including a processor and a memory, wherein: the memory is used to store computer programs; and the processor is used to execute the programs stored in the memory to implement the method of the first aspect.
[0019] Fourthly, embodiments of this application propose a vehicle that includes the electronic equipment of the third aspect.
[0020] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method of the first aspect. Attached Figure Description
[0021] Figure 1 This is a comparison diagram illustrating an evaluation scenario according to an embodiment of this application.
[0022] Figure 2 This is a flowchart illustrating an evaluation method for an autonomous driving algorithm according to an embodiment of this application.
[0023] Figure 3 This is a flowchart illustrating an evaluation method for an autonomous driving algorithm according to another embodiment of this application.
[0024] Figure 4 This is a schematic flowchart illustrating a model training method according to an embodiment of this application.
[0025] Figure 5 This is a structural block diagram of an evaluation device for an autonomous driving algorithm according to an embodiment of this application.
[0026] Figure 6 This is a structural diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0027] To make the technical problems, technical solutions, and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0028] Terminology Explanation: Geographic bounding box (BBox): Defines a geographic region by latitude and longitude range, used for map queries, spatial indexing, and data visualization.
[0029] Time to Collision (TTC): A parameter used in autonomous driving to predict the risk of vehicle collisions; the lower the value, the higher the risk.
[0030] In related technologies, it is necessary to evaluate autonomous driving algorithms in order to assess their capabilities.
[0031] However, through long-term research, the inventors discovered that in the evaluation of autonomous driving algorithms in related technologies, the evaluation indicators and thresholds are fixed, but the indicators and thresholds are different in different evaluation scenarios. If fixed indicators and thresholds are used for evaluation, the accuracy of the evaluation will be poor.
[0032] To facilitate understanding, a comparative diagram of the evaluation scenarios is provided below.
[0033] Please see Figure 1 , Figure 1 This is a comparison diagram illustrating an evaluation scenario according to an embodiment of this application. Figure 1 The following example illustrates the scenario where a self-driving car 110 makes an unprotected left turn using an autonomous driving algorithm.
[0034] Please see Figure 1 (a) in the example Figure 1 As shown in (a), when oncoming traffic is dense and speeds are high, it is difficult to find a left-turn opportunity in a short time, and the opportunity window is also short. In this case, traffic efficiency should not be required to be too high in order to meet users' requirements for safety and comfort. Then the indicator may focus more on safety. For example, the indicator could be travel time or travel safety. If the indicator is travel time, then the threshold for travel time can be set higher to constrain the autonomous driving algorithm to control the vehicle 110 to complete the left turn within the threshold of travel time.
[0035] Please see Figure 1 (b) in the example Figure 1As shown in (b), if there are few vehicles in the opposite lane, there will be more windows and longer time. The algorithm should be able to achieve safety, comfort and efficiency at the same time. Therefore, this indicator can be the passage time, and the threshold of the passage time can be set lower to constrain the autonomous driving algorithm to control the vehicle 110 to complete the left turn within the threshold of the passage time.
[0036] In view of this, this application proposes an evaluation method and related apparatus for autonomous driving algorithms. The method involves acquiring scene features, evaluation index features, and performance gap features. Scene features are those corresponding to the current driving scenario. Evaluation index features include features corresponding to multiple evaluation indices, which are the indicators of the autonomous driving algorithm in the current driving scenario. Performance gap features represent the difference between the driving behavior of the autonomous driving algorithm and the corresponding human driving behavior. The method then fuses the scene features, evaluation index features, and performance gap features to obtain fused features. The weights of the evaluation index features in the fused features are dynamically correlated with the scene features and performance gap features. Based on the fused features, key evaluation indices associated with the current driving scenario and their corresponding threshold values are determined. These key evaluation indices and their corresponding threshold values are used to evaluate the capabilities of the autonomous driving algorithm in the current driving scenario. This improves the accuracy of autonomous driving algorithm evaluation.
[0037] The evaluation methods for autonomous driving algorithms will be explained in detail below.
[0038] Please see Figure 2 , Figure 2 This is a flowchart illustrating an evaluation method for an autonomous driving algorithm according to an embodiment of this application. Figure 2 The method shown can be executed by an electronic device, which may include a terminal or a server (also known as the cloud). The terminal can be a smartphone, tablet, laptop, desktop computer, smart TV, smart home device, etc., without specific limitations. The server can be a standalone physical server, a server cluster consisting of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Figure 2 The methods shown may include: S210. Obtain scene features, evaluation index features, and performance gap features. Scene features are the features corresponding to the current driving scene. Evaluation index features include the features corresponding to multiple evaluation indexes. The multiple evaluation indexes are the indicators corresponding to the autonomous driving algorithm in the current driving scene. Performance gap features are used to represent the gap between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior.
[0039] The scene features (which can also be referred to as the first scene features in this embodiment) can be extracted from the first scene data corresponding to the current driving scene (e.g., a preset scene). The first scene data is used to indicate the scene to be evaluated. Optionally, the first scene data in this embodiment includes the trajectory data of the first traffic participant, the first map data of the area where the first traffic participant is located, and the first environmental data of the area where the first traffic participant is located, which are not limited here. The trajectory data may include, but is not limited to, location, heading angle, etc.; the map data may include, but is not limited to, lane line type, traffic light phase, etc.; and the first environmental data may include, but is not limited to, weather, such as illumination, precipitation intensity, etc., which are not limited here. For example, the current driving scene may be, for example, as follows: Figure 1 The scenarios shown are either left turns or straight-ahead driving, and there are no restrictions on which scenario is considered. The first traffic participant can be any other vehicle or person in the current driving scenario, and there are no restrictions on which scenario is considered.
[0040] Optionally, the first scene data includes the trajectory data of the first traffic participant, the first map data of the area where the first traffic participant is located, and the first environmental data of the area where the first traffic participant is located. Therefore, when determining scene features, the following methods may be used: Feature extraction is performed on the trajectory data of the first traffic participant to obtain the trajectory features of the first traffic participant, feature extraction is performed on the first map data to obtain the first map features, and feature extraction is performed on the first environment data to obtain the first environment features; the trajectory features of the first traffic participant, the first map features, and the first environment features are fused to obtain the scene features.
[0041] In this embodiment, the fusion method is weighted calculation. Optionally, during weighted calculation, the weights of the trajectory features of the first traffic participant, the first map features, and the first environmental features can be set as needed, for example, giving greater weight to more important features. Optionally, features can be extracted from the first scene data using a scene data encoder.
[0042] In one possible implementation, the first scene data includes scene data within a first time period. The first scene data is input into a pre-trained metric determination model to determine a first evaluation metric and its corresponding threshold based on the first scene data. This includes: Obtain the first sequence of each of the multiple indicators corresponding to the first scene data. The first sequence of the indicators is used to represent the changes of the indicators in the first time period. Input the first scene data and the first sequences of the multiple indicators into a pre-trained indicator determination model so that the indicator determination model can determine the scene features based on the first scene data and the first sequences of the multiple indicators.
[0043] The feature to be evaluated (which can also be referred to as the first feature in this embodiment) can be a feature representing the indicator to be evaluated (also referred to as a preset indicator or indicator). In this embodiment, the feature to be evaluated includes features corresponding to multiple indicators. For example, the multiple indicators to be evaluated may include, but are not limited to, passage time, passage safety, time-to-collision (TTC), acceleration or deceleration, etc., and are not limited here. Optionally, features can be extracted from multiple indicators separately, and the features of the indicators can be represented by a deeper level of expression.
[0044] The performance gap feature is used to represent the difference between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior. In this embodiment, the performance gap feature (hereinafter also referred to as the gap feature) can be the gap feature corresponding to each of multiple test indicators. In this embodiment, it can be that in the same scenario, the autonomous driving algorithm controls the vehicle to obtain the indicator value, and the driver controls the vehicle to obtain the indicator value. Then, the difference between the indicator value corresponding to the indicator obtained by the autonomous driving algorithm controlling the vehicle and the indicator value obtained by the driver controlling the vehicle is used to determine the gap feature corresponding to the indicator.
[0045] S220: Integrate scene features, evaluation indicator features, and performance gap features to obtain integrated features. The weights of the evaluation indicator features in the integrated features are dynamically correlated with the scene features and performance gap features.
[0046] In this embodiment, the fusion feature integrates information from the scene feature, the evaluation indicator feature, and the performance gap feature, respectively. Furthermore, the weight of the evaluation indicator feature in the fusion feature is dynamically correlated with the scene feature and the performance gap feature. This allows the weight of the evaluation indicator feature to be adjusted according to the actual situation, which is beneficial to improving the fusion effect and thus improving the evaluation accuracy of the autonomous driving algorithm.
[0047] S230. Based on the fusion features, determine the key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds. The key evaluation indicators and the corresponding indicator thresholds are used to evaluate the capabilities of the autonomous driving algorithm in the current driving scenario.
[0048] In this embodiment, when evaluating the autonomous driving algorithm for the current driving scenario using key evaluation indicators (which can also be referred to as the first evaluation indicator in this embodiment) and the corresponding indicator thresholds, the evaluation can be based on a comparison between the indicator values of the key evaluation indicators under the control of the autonomous driving algorithm and the indicator thresholds of the key evaluation indicators, with the autonomous driving algorithm controlling the vehicle. The evaluation result of the autonomous driving algorithm is then obtained based on this comparison result. The autonomous driving algorithm in this embodiment can refer to an autonomous driving algorithm installed in the vehicle. Taking travel time as an example, the travel time value under the control of the autonomous driving algorithm can be compared with the corresponding threshold. If the travel time value under the control of the autonomous driving algorithm is less than the corresponding threshold, the autonomous driving algorithm is considered to be reasonably controlled; otherwise, the autonomous driving algorithm is considered to be unreasonably controlled. Taking traffic safety as an example, the distance between the vehicle and the first traffic participant under the control of the autonomous driving algorithm can be compared with a distance threshold. If the distance between the vehicle and the first traffic participant under the control of the autonomous driving algorithm is less than the distance threshold, the autonomous driving algorithm is considered to be reasonably controlled; otherwise, the autonomous driving algorithm is considered to be unreasonably controlled. Optionally, the threshold values for the indicators in this embodiment may be one or more, depending on the actual indicators being evaluated, and are not limited here.
[0049] Optionally, the scene features include multiple elements. By fusing scene features, features of the metrics to be evaluated, and performance gap features, the fused features are obtained, including: The similarity between the feature corresponding to each target indicator and multiple scene features is determined, and the similarity between the feature corresponding to each target indicator is used as the weight of the target indicator feature. Based on the similarity between the feature corresponding to each target indicator and multiple scene features, the feature corresponding to each indicator is fused with multiple scene features to obtain the fused indicator feature corresponding to each indicator. The fused indicator features corresponding to multiple indicators are concatenated to obtain the scene-aware indicator features. The scene-aware indicator features are fused with the performance gap features to obtain the gap-aware target indicator features. The fusion features are determined based on the gap-aware target indicator features.
[0050] For example, the fusion in this embodiment can be performed using a cross-attention paradigm, that is... .in: ① The query Q vector comes from the indicator features ,Right now, .
[0051] ② The key K vector comes from scene features ,Right now: .
[0052] ③ The value vector V comes from scene features. :Right now: .
[0053] ④ It is the dimension of vector K. This is a scaling factor used to prevent the gradient from vanishing during training due to an excessively large dot product result.
[0054] Functions and parameters: This operation calculates the similarity between each element (Query) in the metric features and all elements (Keys) in the scene features, and then uses this similarity as a weight to perform a weighted summation of the scene features (Values). This ensures that the feature representation of each evaluation metric (such as TTC) incorporates the most relevant scene information (such as oncoming traffic density). For example, in an unprotected left-turn scenario, the query vector (Query) of the TTC metric will be calculated with the Keys representing scene information such as oncoming vehicles and traffic signals to obtain a high attention weight, thus affecting the output. In this study, the TTC feature was significantly enhanced.
[0055] The following section explains how to integrate the features of scene perception indicators with the features of performance gaps to obtain the features of the indicators to be evaluated for gap perception.
[0056] In one possible implementation, the scene-aware indicator features are fused with the performance gap features to obtain the gap-aware evaluation indicator features, including: The scene-aware indicator features and performance gap features are concatenated to obtain the first concatenated feature matrix. Self-attention is applied to the first concatenated feature matrix to determine the attention weight of the features corresponding to the indicator to be tested based on the gap between the indicator to be tested and the corresponding human driving behavior, thereby obtaining the feature of the indicator to be tested for gap perception.
[0057] In this embodiment, a self-attention mechanism can be used to incorporate the differences in performance between the algorithm and experienced drivers. Metrics and features related to scene perception The features are concatenated to form a concatenated feature matrix (i.e., the first concatenated feature matrix). Then, self-attention is applied to the matrix: .in, Indicator characteristics representing perceived gaps.
[0058] Self-attention mechanisms allow Each element interacts with all other elements, allowing the model to dynamically learn the impact of gap information on the importance of indicators. For example: if If the current algorithm being tested shows a significant gap compared to experienced drivers in terms of comfort (longitudinal acceleration), then the self-attention mechanism will automatically reduce its sensitivity. The weights of comfort-related characteristics are adjusted to avoid affecting the fair evaluation of the overall algorithm performance due to a single weakness.
[0059] The following section explains how to determine the fusion features based on the characteristics of the evaluation indicators based on gap perception.
[0060] In one possible implementation, the fusion features are determined based on the gap-aware characteristics of the indicators to be evaluated, including: The scene features and performance gap features are concatenated to obtain a second concatenated feature matrix. Preset weights and bias parameters are obtained, and the product between the preset weights and the second concatenated feature matrix is calculated. The product is then superimposed with the bias parameters to obtain a superimposed result. An activation function is used to calculate the superimposed result to obtain a gating value. The scene features and performance gap features are fused using the gating value to obtain gap-perceived scene features. The gap-perceived scene features are then concatenated with the gap-perceived evaluation indicator features to obtain fused features.
[0061] To more precisely control algorithm gap information Features of the original scene To correct the intensity of the signal and filter out potential noise, we introduce a gating mechanism. This gating mechanism is typically a learnable gated signal G with an output range of [0, 1], which is then processed by... and Joint decision: .in: This represents a vector concatenation operation; and These are the preset weights and bias parameters for the science department; It is the Sigmoid activation function, which controls the gate value G between 0 and 1.
[0062] Functions and parameters: The gate signal G is used for the difference feature Perform weighting and compare it with the original scene features By integrating the data, we can obtain scene features that reflect the perceived gaps. .in: Represents the Hadamard product; It is a linear projection matrix, which will Dimension mapping to same. The scene features that represent the perception of gaps.
[0063] The value of the gate G reflects the reliability and importance of the algorithm's gap information in the current scene. When G->1, it indicates that the gap information is highly correlated with and reliable in the current scene, and its correction is most effective; when G->0, the gap information is almost ignored, and the algorithm mainly relies on the original scene features. This effectively avoids interference with scene understanding caused by individual abnormal data discrepancies.
[0064] For example, scene features can be established through a three-level attention mechanism. Indicator characteristics Features of the gap between algorithm performance The core of this approach lies in utilizing cross-attention and self-attention mechanisms to selectively enhance and degrade information, ultimately outputting a reinforced feature modulated by both the scene and algorithmic weaknesses. This provides input for subsequent adaptive decision-making.
[0065] 1. Scenario-metric interaction ( and ): The goal of this level is to make the evaluation metrics focus on the current scene context, generating scene-aware metric features through a cross-attention layer. .
[0066] Calculation principles and formulas This level follows the cross-attention paradigm, that is... .in: ① The query Q vector comes from the indicator features ,Right now, .
[0067] ② The key K vector comes from scene features ,Right now: .
[0068] ③ The value vector V comes from scene features. :Right now: .
[0069] ④ It is the dimension of vector K. This is a scaling factor used to prevent the gradient from vanishing during training due to an excessively large dot product result.
[0070] Functions and parameters: This operation calculates the similarity between each element (Query) in the metric features and all elements (Keys) in the scene features, and then uses this similarity as a weight to perform a weighted summation of the scene features (Values). This ensures that the feature representation of each evaluation metric (such as TTC) incorporates the most relevant scene information (such as oncoming traffic density). For example, in an unprotected left-turn scenario, the query vector (Query) of the TTC metric will be calculated with the Keys representing scene information such as oncoming vehicles and traffic signals to obtain a high attention weight, thus affecting the output. In this study, the TTC feature was significantly enhanced.
[0071] Output: Scene-aware metrics Dimensions and Maintain consistency.
[0072] 2. Performance gap between algorithms and experienced drivers - indicator interaction ( and ): This level only considers the performance difference between the algorithm under test and that of experienced drivers (by...). Representation), the indicator features after scene perception. A second adjustment is made to generate indicator features for gap perception. .
[0073] Calculation principles and formulas: This level employs a self-attention mechanism to identify the performance gap between the algorithm and experienced users. Metrics and features related to scene perception The features are concatenated to form a concatenated feature matrix. Then, self-attention is applied to the matrix: .in, Indicator characteristics representing perceived gaps.
[0074] Functions and parameters: Self-attention mechanisms allow Each element interacts with all other elements, allowing the model to dynamically learn the impact of gap information on the importance of indicators. For example: if If the current algorithm being tested shows a significant gap compared to experienced drivers in terms of comfort (longitudinal acceleration), then the self-attention mechanism will automatically reduce its sensitivity. The weights of comfort-related characteristics are adjusted to avoid affecting the fair evaluation of the overall algorithm performance due to a single weakness.
[0075] 3. Scene-Algorithm and Experienced Driver Performance Gap Interaction ( and Gating mechanism: Calculation principles and formulas To more precisely control algorithm gap information Features of the original scene To correct the intensity of the signal and filter out potential noise, we introduce a gating mechanism. This gating mechanism is typically a learnable gated signal G with an output range of [0, 1], which is then processed by... and Joint decision: .in: This represents a vector concatenation operation; and These are the weights and bias parameters of the science department; It is the Sigmoid activation function, which controls the gate value G between 0 and 1.
[0076] Functions and parameters: The gate signal G is used for the difference feature Perform weighting and compare it with the original scene features By integrating the data, we can obtain scene features that reflect the perceived gaps. .in: Represents the Hadamard product; It is a linear projection matrix, which will Dimension mapping to same. The scene features that represent the perception of gaps.
[0077] The value of the gate G reflects the reliability and importance of the algorithm's gap information in the current scene. When G->1, it indicates that the gap information is highly correlated with and reliable in the current scene, and its correction is most effective; when G->0, the gap information is almost ignored, and the algorithm mainly relies on the original scene features. This effectively avoids interference with scene understanding caused by individual abnormal data discrepancies.
[0078] Output: Then, the features obtained through three levels of interaction and It will be further integrated (through splicing) to form the final integrated feature. The process then moves on to subsequent steps, such as outputting key evaluation metrics and their corresponding thresholds based on the fusion features.
[0079] The following section explains how to determine key evaluation indicators and the corresponding threshold values for those indicators.
[0080] In one possible implementation, the indicator determination model, based on fused features, determines key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds, including: The model determines the first probability distribution of multiple indicators based on fusion features. The first probability distribution represents the probability value of each indicator in the current driving scenario. Based on the first probability distribution of multiple indicators, the key evaluation indicators among the multiple indicators are determined. The key evaluation indicators are some of the multiple indicators, and the probability value of the key evaluation indicators is not lower than the probability value of any other indicator. The indicator threshold corresponding to the key evaluation indicators is then determined.
[0081] In this embodiment, the probability distribution may include, but is not limited to, Gaussian distribution, Poisson distribution, etc., and is not restricted here. In this embodiment, the first probability distribution of multiple indicators may represent the probability density corresponding to each of the multiple indicators, or represent the probability corresponding to each of the multiple indicators, and is not restricted here.
[0082] In this embodiment, based on the probability values of multiple indicators in the current driving scenario, the indicator corresponding to one or more of the highest probability values can be selected as the key evaluation indicator, and the indicator threshold corresponding to the key evaluation indicator can be determined.
[0083] The following explains how to determine the threshold values corresponding to key evaluation indicators.
[0084] In one possible implementation, for each key evaluation indicator, a second probability distribution of multiple thresholds corresponding to the key evaluation indicator is determined. The second probability distribution is used to represent the probability value of each threshold in the current driving scenario. Based on the second probability distribution of multiple thresholds corresponding to key evaluation indicators, the indicator thresholds corresponding to the key evaluation indicators are determined. The indicator thresholds are a subset of the multiple thresholds, and the probability value corresponding to each indicator threshold is not lower than the probability value corresponding to any of the thresholds.
[0085] In this embodiment, based on the probability values corresponding to multiple thresholds in the current driving scenario, the threshold with the highest probability value or one or more corresponding thresholds can be selected as the indicator threshold corresponding to the key evaluation indicator.
[0086] The training of the indicator determination model will be explained below.
[0087] In one possible implementation, the metric determination model is trained in the following way: Acquire training data sets, each set including second scenario data, sample index data, and sample index threshold data corresponding to the current driving scenario; input the second scenario data into the index determination model, so that the index determination model outputs predicted index data and predicted index threshold data based on the second scenario data; determine the target training loss based on the sample index data, predicted index data, sample index threshold data, and predicted index threshold data; update the model parameters of the index determination model from the first model parameters to the second model parameters based on the target training loss.
[0088] In this embodiment, the sample indicator data may include the evaluation indicator corresponding to the second scenario data (e.g., the second evaluation indicator), and the sample indicator threshold data may include the threshold corresponding to the evaluation indicator corresponding to the second scenario data; or, the sample indicator data may include a third probability distribution of multiple indicators, and the sample indicator threshold data may include a fourth probability distribution of multiple thresholds corresponding to the second evaluation indicator, wherein the second evaluation indicator is an indicator among multiple indicators, and the probability corresponding to the second evaluation indicator is not less than the probability corresponding to the indicators among multiple indicators other than the second evaluation indicator, which is not limited here. The predicted indicator data may include the evaluation indicator predicted by the model based on the second scenario data, and the predicted indicator threshold data may include the threshold corresponding to the predicted evaluation indicator output by the model; or, the predicted indicator data may include a third probability distribution of multiple indicators, and the predicted indicator threshold data may include a third probability distribution of multiple thresholds corresponding to a third evaluation indicator, wherein the third evaluation indicator is an indicator among multiple indicators, and the probability corresponding to the third evaluation indicator is not less than the probability corresponding to the indicators among multiple indicators other than the third evaluation indicator. When determining the target training loss based on sample indicator data, predicted indicator data, sample indicator threshold data, and predicted indicator threshold data, the loss can be determined based on the loss between sample indicator data and predicted indicator data, as well as the loss between sample indicator threshold data and predicted indicator threshold data.
[0089] Optionally, in this embodiment, the sample indicator data can be predetermined (for example, the probability corresponding to the second evaluation indicator is 1, and the probability corresponding to the indicators other than the second evaluation indicator among multiple indicators is 0). The sample indicator threshold data can be obtained by simulating the second scenario data and statistically analyzing the driver's reaction in the simulated second scenario data, and there is no limitation here.
[0090] Another possible implementation is to train the model through unsupervised training, which can improve the robustness of the model in determining evaluation metrics and their corresponding thresholds.
[0091] In one possible implementation, based on the second scenario data, predictive indicator data and predictive indicator threshold data are output, including: The first scenario data includes scenario data within the second time period. Based on the second scenario data, predictive indicator data and predictive indicator threshold data are output, including: Obtain the second sequence of each of the multiple indicators corresponding to the second scenario data. The second sequence of the indicators is used to represent the changes of the indicators in the second time period. Based on the second scenario data and the second sequences of the multiple indicators, output the predicted indicator data and the predicted indicator threshold data.
[0092] The second time period can be set as needed. The second sequence can include the values of indicators arranged chronologically within the second time period. The descriptions of the first time period and the second sequence are similar and will not be repeated here.
[0093] In another possible implementation, it is not necessary to train the model using the second sequence corresponding to each of the multiple indicators, which can improve the training efficiency of the model.
[0094] In one possible implementation, based on the second scenario data and the second sequences of multiple indicators, predictive indicator data and predictive indicator threshold data are output, including: Feature extraction is performed on the second scene data to obtain second scene features; feature extraction is performed on the second sequence of each of the multiple indicators to obtain a scene-aware indicator feature matrix, which includes the scene-aware indicator features corresponding to each of the multiple indicators; the gap features corresponding to each of the multiple indicators are obtained, which are used to represent the gap between the indicator under the control of the autonomous driving algorithm and under the control of the driver; based on the second scene features, the scene-aware indicator feature matrix and the gap features corresponding to each of the multiple indicators, predicted indicator data and predicted indicator threshold data are output.
[0095] In this embodiment, feature extraction is performed on the second scene data to obtain second scene features, which can be represented by a deeper level of expression. Feature extraction is performed on multiple indicators separately, allowing for a deeper representation of these indicators. In this embodiment, feature extraction is performed on multiple indicators separately to obtain scene-aware indicator features corresponding to each indicator. These scene-aware indicator features are then combined to obtain a scene-aware indicator feature matrix. For example, these multiple indicators may include, but are not limited to, TTC, minimum following distance, longitudinal acceleration, passage time, and passage safety. The difference feature corresponding to the indicator is used to represent the difference between the indicator under the control of the autonomous driving algorithm and under the control of the driver. In this embodiment, under the same scenario, the autonomous driving algorithm controls the vehicle to obtain the indicator value, and the driver controls the vehicle to obtain the indicator value. Then, the difference between the indicator value obtained by the autonomous driving algorithm and the indicator value obtained by the driver is used to determine the difference feature corresponding to the indicator.
[0096] Optionally, features can be extracted from the second scene data using a scene data encoder. The data encoder may include, but is not limited to, spatiotemporal graph convolutional networks, graph attention networks, and multilayer perceptrons.
[0097] Optionally, the second scene data includes trajectory data of the second traffic participant, second map data of the area where the second traffic participant is located, and second environmental data of the area where the second traffic participant is located. Feature extraction is performed on the second scene data to obtain second scene features, including: Feature extraction is performed on the trajectory data of the second traffic participant to obtain the trajectory features of the second traffic participant, feature extraction is performed on the second map data to obtain the second map features, and feature extraction is performed on the second environment data to obtain the second environment features. The trajectory features of the second traffic participant, the second map features, and the second environment features are fused to obtain the second scene features.
[0098] For example, the input in this embodiment includes second scene data, including but not limited to: traffic participant trajectory data (coordinates, bounding boxes, heading angles, types), map data (lane line types, traffic light phases), and environmental data (illuminance, precipitation intensity). Processing: Trajectory data is input into a spatiotemporal graph convolutional network (ST-GCN) to extract motion patterns. Map data is input into a graph attention network (GAT) to encode topological relationships, embedding lane line types (solid / dashed lines) and traffic light phases into semantic vectors. Environmental data is input into a multilayer perceptron (MLP) to encode environmental vectors.
[0099] Output: Second scene features (features can also be called feature vectors) .in, Indicates the characteristics of the second scene. These represent the trajectory features, map features, and environmental features of the second traffic participant, respectively. The second scene features can be referenced from the explanation of the first scene features, and will not be repeated here. In other words, the first scene features are the scene features during inference, and the second scene features are the scene features during training.
[0100] Optionally, in this embodiment, the second sequence of indicators can be encoded by a scene evaluation data encoder to obtain a scene-aware indicator feature matrix.
[0101] The input consists of time-series values (i.e., the second sequence) of metrics such as TTC, minimum distance, longitudinal acceleration, and passage time output from the simulation platform. Processing involves extracting local temporal patterns using a 1D convolutional layer and capturing long-range dependencies using a Transformer encoder. Each metric is then independently encoded and concatenated. Output: Scene-aware feature matrix Where k represents the number of indicators, which is the number of rows in the matrix, and dm represents the feature dimension of a single indicator, which is the number of columns in the matrix.
[0102] Optionally, this embodiment can use a test data encoder to measure the difference between the autonomous driving algorithm and the driver's (also known as a human driver or experienced driver) performance metrics.
[0103] For example, the input is the difference between the performance metrics of the autonomous driving algorithm under test and those of a human driver (e.g., a 30% difference in TTC). The processing involves mapping the difference vector to an MLP. This reflects the algorithm's weaknesses in each metric. Output: Gap Features .
[0104] In another possible implementation, it is also possible to eliminate the need to input gap features into the model. This is beneficial to improving the efficiency of the model outputting evaluation metrics and their corresponding thresholds, thereby improving the efficiency of model training.
[0105] In one possible implementation, based on the second scene features, the scene-aware indicator feature matrix, and the gap features corresponding to each of the multiple indicators, the predicted indicator data and predicted indicator threshold data are output, including: For each of the multiple indicators, the scene-aware indicator features corresponding to the indicator are fused with the second scene features to obtain the scene-aware indicator features corresponding to the indicator. For each indicator, the scene-aware indicator features corresponding to the indicator and the gap features corresponding to the indicator are concatenated to obtain the concatenated feature matrix corresponding to the indicator. The concatenated feature matrix corresponding to the indicator is then fused using an attention mechanism to obtain the gap-aware indicator features. For each indicator, the second scene features are corrected based on the gap features corresponding to the indicator to obtain the gap-aware scene features corresponding to the indicator. For each indicator, the gap-aware indicator features corresponding to the indicator and the gap-aware scene features corresponding to the indicator are fused to obtain the target fusion features corresponding to the indicator. Based on the target fusion features corresponding to each of the multiple indicators, the predicted indicator data and the predicted indicator threshold data are output.
[0106] In this embodiment, the second scene features, the scene-aware indicator feature matrix, and the gap features corresponding to each of the multiple indicators can be fused through the cross attention of the cross attention layer. Then, the fused features are used to output the predicted indicator data and the predicted indicator threshold data.
[0107] In one possible implementation, the sample indicator data includes a third probability distribution of multiple indicators; the sample indicator threshold data includes a fourth probability distribution of multiple thresholds corresponding to a second evaluation indicator, wherein the second evaluation indicator is one of the multiple indicators, and the probability corresponding to the second evaluation indicator is not less than the probability corresponding to the indicators other than the second evaluation indicator; the prediction indicator data includes a third probability distribution of multiple indicators; the prediction indicator threshold data includes a third probability distribution of multiple thresholds corresponding to a third evaluation indicator, wherein the third evaluation indicator is one of the multiple indicators, and the probability corresponding to the third evaluation indicator is not less than the probability corresponding to the indicators other than the third evaluation indicator; based on the sample indicator data, prediction indicator data, sample indicator threshold data, and prediction indicator threshold data, the target training loss is determined, including: A first training loss is determined based on the third probability distribution of multiple indicators and the loss between the third probability distributions of multiple indicators; a second training loss is determined based on the fourth probability distribution of multiple thresholds and the loss between the third probability distributions of multiple thresholds; a first confidence level is determined based on the fourth probability distribution of multiple thresholds, and a second confidence level is determined based on the third probability distribution of multiple thresholds, and a third training loss is determined based on the loss between the first confidence level and the second confidence level; the first training loss, the second training loss and the third training loss are fused to determine the target training loss.
[0108] The explanations for the second and third probability distributions can be found in the description of the first probability distribution, and will not be repeated here.
[0109] In this embodiment, the processing can be performed through an adaptive decision output layer.
[0110] This layer receives fusion features from the cross-attention fusion module. Based on this, the model makes two decisions: dynamically selecting the most critical evaluation metric in the current scenario, and generating adaptive thresholds and their confidence levels that conform to the scenario characteristics for the selected metric. Its structure and working principle are as follows: 1) Indicator Selector: The goal of the metric selector is to simulate the decision-making process of human testing experts and dynamically determine a set of the most relevant evaluation metrics based on the characteristics of the current scenario (e.g., in some scenarios, safety should be prioritized over comfort).
[0111] Input and preprocessing: Input is fused features First, a linear transformation is performed using a two-layer MLP to map it to a hidden space that is related to the number of candidate indicators k.
[0112] .
[0113] in, , , These are the learnable parameters of an MLP; ReLU is the activation function. Output Each dimension corresponds to an initial score for a candidate metric.
[0114] Gumbel-Softmax relaxation: In order to obtain a differentiable, approximately discrete probability distribution for index selection from the initial score H We employ the Gumbel-Softmax technique, which transforms the sampling process into a differentiable operation by introducing Gumbel noise.
[0115] .
[0116] in: It is random noise that follows a Gumbel(0,1) distribution; It's a temperature parameter. (In the early stages of training) Larger outputs with a smoother distribution encourage exploration; as training progresses... As the distribution gradually decreases, it approaches the one-hot vector, achieving an approximately discrete distribution.
[0117] Output: It is a probability distribution (that is, the third probability distribution of the index in this embodiment), and each element This represents the probability of selecting the i-th evaluation metric. Ultimately, the top n highest-ranking metrics can be selected for evaluation in the current scenario.
[0118] 2) Threshold generator: The threshold generator is responsible for generating dynamic and reasonable thresholds for the metrics selected by the metric selector. It employs a conditional variational autoencoder (CVAE) structure, designed to capture the distribution of threshold choices made by experienced drivers in similar scenarios, rather than outputting a single fixed value.
[0119] CVAE structure and working principle: CVAE is a generative model that incorporates conditional information (scene features in this example) into both the encoder and decoder. (and selected indicator features). Its goal is to learn a conditional probability distribution. This allows for the generation of diverse thresholds that conform to the characteristics of the scene.
[0120] ① Encoder: Encoder receive condition information ( The encoder learns the posterior distribution of the latent variable z, given the target value (a reference threshold y obtained from experienced drivers' driving data) and the target value. Assuming this posterior distribution follows a Gaussian distribution, the encoder outputs its mean and variance.
[0121] .
[0122] .
[0123] in, These are the encoded parameters.
[0124] ② Latent space sampling: A latent vector z is sampled from the Gaussian distribution obtained from the encoder: .
[0125] This latent vector z contains, given conditions Under this condition, the threshold y is compressed and probabilistically represented.
[0126] ③ Decoder: The decoder uses conditional information ( It takes the latent variable z obtained from sampling as input to reconstruct or generate a threshold. Its output is the parameters of the Gaussian distribution that generates the threshold.
[0127] .
[0128] .
[0129] in, These are the parameters of the decoder. The final generated threshold. It can be obtained from The mean was obtained by sampling from the middle and used directly. As the threshold for generation.
[0130] Output and confidence level: ① Threshold distribution parameters: The decoder directly outputs the distribution parameters for generating the threshold, i.e., the mean. and variance .
[0131] ②Threshold sampling: The specific threshold is obtained through sampling using reparameterization techniques: .
[0132] ③ Threshold confidence level: Confidence reflects the model's certainty regarding the generated threshold. Variance A larger variance indicates higher model uncertainty and lower confidence. We normalize the variance using the Sigmoid activation function to obtain the confidence score: A fallback mechanism is triggered when the confidence level is low, using a preset conservative threshold (based on developer experience).
[0133] In this embodiment, CVAE is trained by optimizing the lower bound of the data (ELBO). The loss function includes two parts: reconstruction loss and regularization term. The reconstruction loss is used to encourage the threshold distribution generated by the decoder to be close to the distribution of real experienced drivers, and is implemented using the negative log-likelihood function. The regularization term is implemented using KL divergence to constrain the distribution of the learned latent variables. Approximate prior distribution This ensures that the potential space is regular, facilitating sampling generation.
[0134] For example, multi-task joint loss is used here: .in: Indicator selection loss (i.e., first training loss) Weighted cross-entropy enhances the selection of safety metrics for critical scenarios (such as unprotected left turns).
[0135] Threshold fitting loss (also known as the second training loss) The negative log-likelihood of human driving index values constrains the generation of thresholds to conform to the behavior distribution of experienced drivers. In this embodiment, the mean can be determined based on a second distribution of multiple thresholds, and the mean can also be determined based on a third distribution of multiple thresholds. The difference between the mean determined by the second distribution of multiple thresholds and the mean determined by the third distribution of multiple thresholds is used as the second training loss. Alternatively, the second training loss can be obtained by comparing the similarity between the curves formed by the second distribution of multiple thresholds and the curves formed by the third distribution of multiple thresholds.
[0136] Confidence calibration loss (also known as the third loss) Use the Brier Score to align confidence with prediction accuracy.
[0137] Optionally, the model may perform inconsistently in different scenarios. However, this embodiment aims for the model to run stably and output consistently, providing definite and reliable outputs in different scenarios. Monte Carlo Dropout: Multiple sampling is performed during inference to calculate the threshold accuracy of the model output after multiple samplings in different scenarios. The accuracy is simulated, and this variance is used to quantify the threshold uncertainty. In this embodiment, low confidence ( When a rule rollback is triggered, such as when using a preset threshold for scene classification, dynamic confidence fusion is employed: final decision weights. That is, the final indicator threshold is calculated by using both model output and expert opinion, and then multiplying the two by the confidence level and 1-confidence level respectively, which can balance data-driven and rule-based fallback.
[0138] In one possible implementation, after updating the model parameters of the indicator determination model from the first model parameters to the second model parameters based on the target training loss, the method further includes: acquiring third scene data corresponding to the current driving scenario; inputting the third scene data into the indicator determination model to determine a fourth evaluation indicator and the corresponding indicator threshold based on the third scene data; generating motion control parameters using an autonomous driving algorithm based on the third scene data and the indicator threshold corresponding to the fourth evaluation indicator, the motion control parameters being used to control the movement of the vehicle; determining a reward value based on the motion control parameters; and updating the model parameters of the indicator determination model from the second model parameters to the third model parameters with the goal of maximizing the reward value.
[0139] In this embodiment, since metric selection is a discrete decision-making process, the Policy Gradient Approach (PPO algorithm) is used for optimization. The reward function R is defined as the long-term evaluation gain. For example, in simulation testing, it represents the overall improvement in the performance of the algorithm under test in subsequent scenarios after selecting a certain set of metrics. The training objective is to maximize the expected reward. This allows us to learn a strategy: to select the combination of metrics that most effectively drives algorithm optimization in a specific scenario.
[0140] The third scenario data can be referenced from the description of the second scenario data. In this embodiment, the autonomous driving algorithm generates motion control parameters based on the third scenario data and the threshold values corresponding to the fourth evaluation indicator. These parameters can instruct the autonomous driving algorithm to control the vehicle within the threshold values corresponding to the fourth evaluation indicator. For example, in a left-turn scenario, these parameters could instruct the autonomous driving algorithm to control the vehicle to complete the left turn within a time limit corresponding to the traffic flow time threshold.
[0141] In this embodiment, the autonomous driving algorithm generates motion control parameters based on the third scene data and the corresponding threshold of the fourth evaluation index. The reward value is determined based on the motion control parameters, and the model's ability to determine the evaluation index and its corresponding threshold is further evaluated through the reward value.
[0142] In one possible implementation, the reward value is determined based on motion control parameters, including: Based on motion control parameters, the safety parameter value, ride comfort parameter value, traffic efficiency parameter value, and number of false alarms during the vehicle's movement are determined. Based on the safety parameter value, ride comfort parameter value, traffic efficiency parameter value, and number of false alarms, a reward value is determined. The reward value is positively correlated with the safety parameter value, ride comfort parameter value, and traffic efficiency parameter value, respectively, and negatively correlated with the number of false alarms.
[0143] An example, the reward function: By optimizing decision-making strategies through the PPO algorithm, long-term evaluation benefits can be maximized.
[0144] in, Indicates the reward value. Indicates the value of the security parameter. Indicates the value of ride comfort parameters, This represents the traffic efficiency parameter value. Indicates the number of false alarms. , , and Each represents its corresponding weight.
[0145] In this embodiment, calculating reward values from different dimensions enables the model to determine evaluation metrics from different perspectives.
[0146] In another possible implementation, it may be unnecessary to calculate the reward value, which can improve the training efficiency of the model.
[0147] Optionally, this embodiment can progressively train from simple scenarios (straight ahead, few cars) to complex scenarios (left turn, congestion). The scenarios in the data are classified by labeling, and then divided into scenario sets of different difficulty levels with reference to expert opinions. The scenario sets are fed into the model for training from easy to difficult. When each model is trained to 85% accuracy, the next level of scenario set is introduced into the training, until the output of the model on all scenario sets is stable and converged.
[0148] To verify the effectiveness of the solution, the following test can be conducted.
[0149] Experimental verification plan: Test scenario design: Evaluation metrics: Evaluation metrics are used to measure whether the output of this model's metric selection model and threshold output model is consistent with the data performance of experienced drivers in the real world, thereby measuring the overall level of the model.
[0150] Indicator selection accuracy: The result of indicator selection is categorical data (Nominal Variable), and its accuracy is measured by the F1 score.
[0151] ① Data preparation: By running the model in a large number of simulation tests, the set of indicators selected by the model (i.e., predicted values) and the set of indicators to be tested labeled by human experts (true values) are recorded to support subsequent comparisons.
[0152] ② Binary classification comparison: In the data prepared above, for each indicator, a confusion matrix can be constructed based on whether the predicted value equals the true value: True Case (TP): The model predicts that this indicator should be used, and human experts also believe that this indicator should be used.
[0153] False positives (FP): The model predicts that the indicator should be used, but human experts believe that the indicator is unnecessary.
[0154] False negatives (FN): This metric is not required by the model prediction, but human experts believe it is necessary.
[0155] True negative (TN): The model predicts that the indicator is not needed, and human experts also believe that the indicator is not needed.
[0156] ③ Calculate precision and recall: Based on the above confusion matrix, the following two indicators can be calculated: Accuracy (P) = TP / (TP + FP). It measures how many of the metrics selected by the model are also recognized by experts.
[0157] Recall (R) = TP / (TP + FN). It measures how many of the expert-approved metrics were selected by the model.
[0158] ④ Calculate the F1 score: Then, the accuracy and recall can be reconciled by calculating the F1 score: .
[0159] ⑤ Macroeconomic average: Since we have multiple indicators, each with its own F1 score, we use the macro-F1 average as the final indicator selection accuracy. This involves calculating each F1 score individually and then taking the arithmetic mean. The value ranges from 0 to 1; the closer it is to 1, the better the model's output matches human expert decisions.
[0160] Threshold rationality: The accuracy of the threshold compared with the driving level of an experienced driver is a quantitative data (scale variable), which is measured by KL divergence.
[0161] ① Data preparation: Generation distribution (P): In a large number of test scenarios, record all the thresholds generated by the model for a specific metric and normalize them into a probability distribution.
[0162] Reference distribution (Q): Extract values of the same indicators from the data of experienced drivers and normalize them into a probability distribution as the ground truth for rationality.
[0163] ② Calculate the KL divergence: KL divergence is used to calculate the difference between two distributions. The formula is: , where i iterates through all possible data.
[0164] The KL divergence value is always >= 0. The smaller the value, the more similar the threshold distribution generated by the model is to the threshold distribution in the experienced driver data, meaning the model's output is more reasonable.
[0165] Baseline comparison: Tested on the nuPlan-Cross dataset: Fixed threshold method: PDMS=72.1, high frequency of emergency braking (3.2 times / km).
[0166] Rule-adaptive method: PDMS=79.4, misjudgment rate of 15% in complex scenarios.
[0167] This model has the following parameters: target PDMS > 85, emergency braking frequency < 1.5 times / km, and confidence error < 0.1.
[0168] In summary, this embodiment achieves a closed loop of "scene understanding → indicator decision → threshold generation" through deep learning, driving the evolution of autonomous driving testing from static rules to cognitive intelligence. The model code can be implemented using PyTorch, with a focus on optimizing the real-time performance (<50ms) of the cross-attention and CVAE modules to support large-scale simulation testing. Furthermore, an implicit mapping between scene semantics and indicators can be established through cross-attention, replacing manual threshold rules. Test gap data is introduced to strengthen the weights of key indicators, accelerating algorithm iteration and optimization. Finally, decision confidence scores can be jointly output to support human-machine collaborative evaluation and review.
[0169] In other words, this embodiment significantly improves the accuracy and efficiency of autonomous driving algorithm testing through a dynamic adaptive evaluation mechanism, specifically as follows: (1) Improved evaluation accuracy: Dynamic threshold adjustment makes the evaluation indicators adaptable to various scenarios, especially in complex scenarios such as unprotected left turns, which can simultaneously optimize the balance between safety, comfort and traffic efficiency. (2) Enhanced coverage of long-tail scenarios: Based on the generalization capability of Transformer, it supports zero-shot learning and scenario adaptation, improves the deployment efficiency of new scenarios, and effectively reduces the cost of manual intervention; (3) Evaluation consistency optimization: Effectively reduce cross-scenario evaluation bias. By setting a bias level close to a certain threshold of human drivers, the credibility of test results can be significantly improved. (4) Accelerated R&D iteration: Reduce the coding of rule-based evaluation scripts, train the evaluation system dynamically with data, decouple the evaluation system training and algorithm testing, greatly reduce the algorithm iteration cycle, and support rapid verification of large-scale simulation tests; (5) Enhanced security verification: By filtering with probabilistic confidence, the false positive rate of abnormal scenarios is reduced and the detection rate of key risk scenarios that meet the current algorithm level is improved.
[0170] In general, the model training process can be as follows: Figure 3 As shown, Figure 3 This is a flowchart illustrating an evaluation method for an autonomous driving algorithm according to another embodiment of this application.
[0171] like Figure 3 The process shown may include: S310, Design Model.
[0172] S320, Collect data.
[0173] The data collected in this embodiment may include, but is not limited to, data sample groups.
[0174] S330, Complete model training.
[0175] In this embodiment, the conditions for model training may include the target loss stabilizing and being less than a loss threshold.
[0176] S340, Select indicators and thresholds.
[0177] S350, access closed-loop simulation test, output test results.
[0178] S340 and S350 can be referred to in the description of S210-S230, and will not be repeated here.
[0179] The model training process will now be explained in detail with reference to the above examples.
[0180] Please see Figure 4 , Figure 4 This is a schematic flowchart illustrating a model training method according to an embodiment of this application. Figure 4 The method shown can be performed by an electronic device, such as Figure 4 The methods shown may include: S401, Input test data.
[0181] S402, Input evaluation data.
[0182] The evaluation data includes the second sequence of each of the multiple indicators.
[0183] S403, Input the second scene data.
[0184] S404. The test data is encoded by the test data encoder to obtain the gap features.
[0185] S405. The evaluation data is encoded by the evaluation data encoder to obtain the second indicator feature matrix.
[0186] S406. The second scene data is encoded by the second scene data encoder to obtain the second scene features.
[0187] S407, the second indicator feature matrix, and the second scene feature interact to obtain the second indicator feature for scene perception.
[0188] S408. The gap feature interacts with the second indicator feature of scene perception to obtain the second indicator feature of gap perception.
[0189] S409. The gap feature interacts with the second scene feature to obtain the second scene feature for gap perception.
[0190] S410, the second scene feature of gap perception and the second indicator feature of gap perception are fused to obtain the target fusion feature.
[0191] S411. Output the probability distribution of indicators based on the target fusion features through the indicator selector.
[0192] S412. The threshold generator outputs the threshold and the confidence level of the threshold based on the target fusion features.
[0193] S413, Adaptive evaluation results.
[0194] In this embodiment, the adaptive evaluation result can be calculated based on the probability distribution of the index, the output threshold, and the confidence level of the threshold to determine the target loss.
[0195] In this embodiment, the scenario, vehicle performance, and test metrics are encoded separately, then cross-fused to form an understanding of the scenario and the metrics. The model will learn "which metrics should be focused on in such a scenario, and what the thresholds for these metrics should be set at." Working principle: Scenario encoding: The scenario is encoded, recording key scenario features related to the evaluation in the feature vector. Evaluation encoding: The evaluation metrics are encoded, recording the characteristics of the evaluation algorithm in the feature vector, i.e., the level that human drivers can achieve in each dimension. Cross-fusion: After the test data and evaluation features are cross-fused, the model learns "what thresholds should be focused on for this metric." After the test data and scenario data are cross-fused, the model learns "what these thresholds are generally for this scenario."
[0196] Please see Figure 5 , Figure 5 This is a structural block diagram of an evaluation device for an autonomous driving algorithm according to an embodiment of this application. Figure 5 The device can be applied to electronic devices, such as Figure 5 The apparatus may include an acquisition module 510, a fusion module 520, and an evaluation module 530, wherein: The acquisition module 510 is used to acquire scene features, evaluation indicator features, and performance gap features. Scene features are the features corresponding to the current driving scenario. Evaluation indicator features include features corresponding to multiple evaluation indicators, which are the indicators corresponding to the autonomous driving algorithm in the current driving scenario. Performance gap features are used to represent the gap between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior. The fusion module 520 is used to fuse scene features, evaluation indicator features, and performance gap features to acquire fused features. The weights of the evaluation indicator features in the fused features are dynamically associated with the scene features and performance gap features. The evaluation module 530 is used to determine the key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds based on the fused features. The key evaluation indicators and the corresponding indicator thresholds are used to evaluate the capabilities of the autonomous driving algorithm in the current driving scenario.
[0197] The apparatus described in this embodiment can be referred to the description of the method embodiment, and will not be repeated here.
[0198] The apparatus in this embodiment can be described with reference to the above method, and will not be repeated here.
[0199] This application also provides an electronic device, please refer to... Figure 6 , Figure 6 The electronic device 600 shown includes a processor 610 and a memory 620, wherein the memory 610 is used to store computer programs; and the processor 620 is used to execute the programs stored in the memory 610 to implement the methods described in any embodiment of this application.
[0200] This application also provides a vehicle that includes the electronic equipment described in the above embodiments.
[0201] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any embodiment of this application.
[0202] In this application, "multiple" refers to two or more.
[0203] In this application, unless otherwise expressly defined, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0204] The terms “first,” “second,” “third,” “fourth,” etc., in this application (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0205] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0206] Unless otherwise specified, all steps in this application may be performed sequentially or randomly. For example, if a method includes steps A and B, it means that the method may include steps A and B performed sequentially, or it may include steps B and A performed sequentially. For example, if a method may also include step C, it means that step C may be added to the method in any order. For example, the method may include steps A, B, and C, or it may include steps A, C, and B, or it may include steps C, A, and B, etc.
[0207] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for evaluating an autonomous driving algorithm, characterized in that, include: The system acquires scene features, evaluation index features, and performance gap features. The scene features are the features corresponding to the current driving scene. The evaluation index features include features corresponding to multiple evaluation indicators. The multiple evaluation indicators are the indicators of the autonomous driving algorithm in the current driving scene. The performance gap features are used to represent the gap between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior. The scene features, the features of the indicators to be evaluated, and the performance gap features are integrated to obtain the integrated features. The weights of the features of the indicators to be evaluated in the integrated features are dynamically associated with the scene features and the performance gap features. Based on the fusion features, key evaluation indicators associated with the current driving scenario and corresponding indicator thresholds are determined. The key evaluation indicators and corresponding indicator thresholds are used to evaluate the capabilities of the autonomous driving algorithm in the current driving scenario.
2. The method according to claim 1, characterized in that, The scene features include multiple features. The fusion of these scene features, the features of the metrics to be evaluated, and the performance gap features to obtain the fused features includes: Determine the similarity between the feature corresponding to each target indicator and multiple scene features, and use the similarity between the feature corresponding to each target indicator as the weight of the target indicator feature. Based on the similarity between the feature corresponding to each target indicator and multiple scene features, the feature corresponding to each indicator is fused with multiple scene features to obtain the fused indicator feature corresponding to each indicator. The fused indicator features corresponding to each of the multiple indicators are spliced together to obtain scene-aware indicator features. The scene-aware indicator features are fused with the performance gap features to obtain the gap-aware evaluation indicator features. The fusion features are determined based on the characteristics of the evaluation indicators to be evaluated based on the aforementioned gap perception.
3. The method according to claim 2, characterized in that, The process of fusing the scene-aware indicator features with the performance gap features to obtain the gap-aware evaluation indicator features includes: The scene-aware indicator features are concatenated with the performance gap features to obtain a first concatenated feature matrix; Self-attention is applied to the first spliced feature matrix to determine the attention weight of the feature corresponding to the test indicator based on the gap between the test indicator and the corresponding human driving behavior, thereby obtaining the gap-aware test indicator feature.
4. The method according to claim 2, characterized in that, The determination of fusion features based on the features of the evaluation index under test based on the gap perception includes: The scene features and the performance gap features are concatenated to obtain a second concatenated feature matrix; Obtain preset weights and bias parameters, calculate the product between the preset weights and the second concatenated feature matrix, and superimpose the product with the bias parameters to obtain a superimposed result. Then, calculate the gating value by using an activation function on the superimposed result. The scene features and the performance gap features are fused using the gating value to obtain the scene features for gap perception. The scene features of the gap perception are concatenated with the evaluation index features of the gap perception to obtain the fused features.
5. The method according to claim 1, characterized in that, The indicator determination model, based on the fusion features, determines key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds for the key evaluation indicators, including: Based on the fusion features, the indicator determination model determines the first probability distribution of multiple indicators, which represents the probability value of each of the multiple indicators in the current driving scenario. Based on the first probability distribution of the multiple indicators, key evaluation indicators are determined among the multiple indicators. The key evaluation indicators are some of the multiple indicators, and the probability value corresponding to the key evaluation indicators is not lower than the probability value corresponding to any indicator. Determine the threshold values corresponding to the key evaluation indicators.
6. The method according to claim 5, characterized in that, Determining the threshold values corresponding to the key evaluation indicators includes: For each key evaluation indicator, a second probability distribution of multiple thresholds corresponding to the key evaluation indicator is determined. The second probability distribution is used to represent the probability value of each of the multiple thresholds in the current driving scenario. Based on the second probability distribution of multiple thresholds corresponding to the key evaluation indicators, the indicator thresholds corresponding to the key evaluation indicators are determined. The indicator thresholds are some of the multiple thresholds, and the probability value corresponding to the indicator threshold is not lower than the probability value corresponding to any threshold.
7. An evaluation device for an autonomous driving algorithm, characterized in that, include: The acquisition module is used to acquire scene features, evaluation index features and performance gap features. The scene features are the features corresponding to the current driving scene. The evaluation index features include the features corresponding to multiple evaluation indicators. The multiple evaluation indicators are the indicators of the autonomous driving algorithm in the current driving scene. The performance gap features are used to represent the gap between the driving behavior corresponding to the autonomous driving algorithm and the corresponding human driving behavior. The fusion module is used to fuse the scene features, the features to be evaluated, and the performance gap features to obtain fused features. The weights of the features to be evaluated in the fused features are dynamically associated with the scene features and the performance gap features. The evaluation module is used to determine key evaluation indicators associated with the current driving scenario and the corresponding indicator thresholds based on the fusion features. The key evaluation indicators and the corresponding indicator thresholds are used to evaluate the capabilities of the autonomous driving algorithm in the current driving scenario.
8. An electronic device, characterized in that, Includes processor and memory, of which: Memory, used to store computer programs; A processor for executing a program stored in memory to implement the method described in any one of claims 1-6.
9. A vehicle, characterized in that, Including the electronic device as claimed in claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as claimed in any one of claims 1-6.