Support Tools for Self-Driving Vehicle Testing

The proposed system addresses the challenge of evaluating autonomous vehicle performance by using a computer system that analyzes real-time recognition errors and driving performance, providing a comprehensive visualization to improve system performance.

JP7692501B2Active Publication Date: 2025-06-13FIVE AI LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023575621
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-17
Filing Date
2022-06-08
Publication Date
2025-06-13
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

Current methods for evaluating the performance of autonomous vehicle systems and trajectory planners lack effective tools for real-time recognition error analysis and driving performance assessment, making it difficult to identify and address recognition errors that impact overall driving performance.

Method used

A computer system and method for testing real-time recognition systems in autonomous vehicles, which includes data input from real-world driving runs, a rendering component for generating graphical user interfaces, a ground-truthing pipeline for processing sensor data, and a recognition oracle for identifying recognition errors and generating error timelines.

Benefits of technology

The system provides a comprehensive visualization of recognition errors and driving performance, enabling experts to correlate recognition errors with driving performance and identify causes of recognition errors, thereby improving the overall performance of autonomous vehicle systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692501000001
    Figure 0007692501000001
  • Figure 0007692501000002
    Figure 0007692501000002
  • Figure 0007692501000003
    Figure 0007692501000003
Patent Text Reader

Abstract

1. A computer-implemented method for assessing performance of an autonomous vehicle, comprising: receiving at an input performance data of at least one autonomous driving run, the performance data comprising a time series of at least one recognition error and a time series of at least one driving performance result; and generating at a rendering component rendering data for rendering a graphical user interface, the graphical user interface for visualizing the performance data, the graphical user interface comprising a recognition error timeline and a driving assessment timeline, the timelines being aligned in time and divided into a plurality of time steps of the at least one driving run, and for each time step, the recognition error timeline comprises a visual indication of whether a recognition error occurred at that time step, and the driving assessment timeline comprises a visual indication of the driving performance at that time step.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to tools and methods for evaluating the performance of autonomous vehicle systems and trajectory planners in real or simulated scenarios, as well as computer programs and systems for implementing them. Application examples include performance tests of ADS (Autonomous Driving System) and ADAS (Advanced Driver Assist System).

Background Art

[0002] In the field of autonomous vehicles, there has been a large and rapid development. An autonomous vehicle (AV) is a vehicle equipped with sensors and a control system that enables it to operate without human control of its behavior. Autonomous vehicles are equipped with sensors that enable them to perceive their physical environment, such sensors including, for example, cameras, radars, and lidars. Autonomous vehicles are equipped with a properly programmed computer that can process data received from the sensors and make safe and predictable decisions based on the context recognized by the sensors. Autonomous vehicles may be fully autonomous or semi-autonomous (in that they are designed to operate without human supervision or intervention at least in certain situations). Semi-autonomous systems require various levels of human monitoring and intervention, such systems including advanced driver assistance systems and level 3 autonomous driving systems. Testing the behavior of sensors and control systems installed in a particular autonomous vehicle or a certain type of autonomous vehicle has various aspects.

[0003] Since "Level 5" vehicles are always guaranteed to meet a minimum safety level, they can operate completely autonomously in any situation. Such vehicles do not require any manual control (such as steering wheels, pedals, etc.).

[0004] In contrast, Level 3 and Level 4 vehicles can operate completely autonomously, but only within specific defined situations (e.g., within a geofenced area). Level 3 vehicles must be equipped to autonomously handle any situation that requires immediate response (such as emergency braking), but a change in situation may trigger a "handover request" that requires the driver to take control of the vehicle within a certain limited time frame. Level 4 vehicles have similar limitations, but if the driver fails to respond within the required time frame, the Level 4 vehicle must also be able to autonomously perform a "minimum risk maneuver" (MRM), i.e., take appropriate measures (such as slowing down and parking the vehicle) to bring the vehicle to a safe state. Level 2 vehicles require the driver to be ready to intervene at any time, and it is the driver's responsibility to intervene whenever the autonomous system is unable to respond appropriately. In Level 2 automation, it is the driver's responsibility to determine when intervention is required, while in Levels 3 and 4, this responsibility shifts to the vehicle's autonomous system, and the vehicle must warn the driver when intervention is required.

[0005] As the level of autonomy increases and more responsibility shifts from humans to machines, safety becomes an increasingly difficult challenge. In autonomous driving, the importance of guaranteed safety is recognized. Guaranteed safety does not necessarily imply zero accidents, but rather means ensuring that a minimum safety level is met in defined situations. It is generally considered that for autonomous driving to be achievable, this minimum safety level must significantly exceed that of human drivers.

[0006] According to Shalev-Shwartz et al., "On a Formal Model of Safe and Scalable Self-driving Cars" (2017), arXiv:1708.06374 (RSS paper), which is hereby incorporated by reference in its entirety, human driving is estimated to cause on the order of 10 -6 accidents per hour. Based on the assumption that an autonomous driving system needs to reduce this by at least three orders of magnitude, the RSS paper concludes that a minimum safety level of on the order of 10 -9 accidents per hour needs to be guaranteed, and thus points out that a purely data-driven approach requires an enormous amount of driving data to be collected every time there is a change to the AV system's software or hardware.

[0007] The RSS paper provides a model-based approach to guaranteed safety. The rule-based Responsibility-Sensitive Safety (RSS) model is constructed by formalizing the following few "common sense" driving rules. "1. Do not rear-end people. 2. Do not cut in recklessly. 3. Right of way is given, not taken. 4. Be careful in places with poor visibility. 5. If an accident can be avoided without causing another accident, then it must be done." The RSS model has been proven to be safe in the sense that no accidents will occur if all agents always comply with the rules of the RSS model. The aim is to reduce by several orders of magnitude the amount of driving data that needs to be collected to demonstrate the required safety level.

[0008] A safety model (e.g., RSS) can be used as a basis for evaluating the quality of trajectories planned or realized by an agent in a real or simulated scenario under the control of an autonomous system (stack). The stack is tested by exposing it to various scenarios and evaluating the resulting self-trajectories for compliance with the rules of the safety model (rule-based testing). The rule-based testing approach can also be applied to other aspects of performance, such as comfort or progress towards a defined goal.

Summary of the Invention

Problems to be Solved by the Invention

[0009] Techniques are described that enable an expert to assess both the recognition errors and driving performance of an AV system. Evaluating the recognition output of an AV recognition system by comparison with a ground truth recognition output enables an expert to assess the impact of recognition issues on the overall performance of a given AV system. Herein, a UI is described that presents recognition errors and driving performance in a single visualization to provide a correlation between recognition and driving performance and assist an expert in identifying the causes of recognition errors that may affect overall driving performance.

Means for Solving the Problems

[0010] A first aspect herein is a computer system for testing a real-time recognition system, the real-time recognition system being for deployment in a vehicle equipped with sensors, the computer system comprising At least one input configured to receive data from at least one real-world driving run performed by a sensor-equipped vehicle, the data comprising: (i) a time series of sensor data captured by the sensor-equipped vehicle, and (ii) a time series of at least one associated on-road recognition output extracted from the time series of sensor data by a real-time recognition system under test. A rendering component configured to generate rendering data for rendering a graphical user interface (GUI), the graphical user interface comprising a recognition error timeline having a visual indication of recognition errors occurring at each of a plurality of time steps of at least one real-world driving run. A ground-truthing pipeline configured to process at least one of (i) the time series of sensor data and (ii) the time series of on-road recognition outputs by applying at least one non-real-time and / or non-causal recognition algorithm to extract a time series of at least one ground-truth recognition output (a “pseudo ground-truth”) for comparison with the on-road recognition outputs. A recognition oracle configured to identify recognition errors occurring in one or more time intervals to generate a recognition error timeline by comparing the time series of on-road recognition outputs with the time series of ground-truth recognition outputs. A computer system comprising the above components is targeted.

[0011] In an embodiment, recognition errors may be identified by calculating a numerical error value between the time series of on-road recognition outputs and the time series of ground-truth recognition outputs and comparing the numerical error value with at least one recognition error threshold.

[0012] For example, a numerical error value may be identified as a recognition error only if the numerical error value exceeds an error threshold value.

[0013] The error threshold value may be fixed or variable. For example, different recognition error threshold values may be applied to different actors / agents, or different types thereof (e.g., different thresholds for vehicles vs. pedestrians, etc.).

[0014] The error threshold value may be adjustable, for example, via a GUI or via rule definition instructions provided to a recognition oracle (e.g., those encoded in a domain-specific language (DSL)), or may be set in another way. A rule editor may be provided to code the rule definition instructions in DSL in the form of a recognition error specification. The latter approach provides what is referred to herein as a "recognition error framework".

[0015] The error threshold value may also be changed according to one or more scene variables (driving variables) of a driving operation, for example, variables of an object to which the error threshold value is applied. For example, for a given object (e.g., an agent or a static object), the recognition error threshold value for that object may be increased along with the distance between that object and the self-agent (based on the fact that a smaller recognition error is more critical for nearby objects). The same effect can be achieved using a fixed threshold, provided that the numerical error value is weighted according to the scene variable (e.g., weighted by the reciprocal of the distance). Unless otherwise indicated herein, references to a "variable threshold" include the latter implementation.

[0016] (Weighted) numerical recognition errors may be normalized, i.e., optionally transformed to a certain scale, for example, a range [-1,1] where the failure threshold is set to zero, together with a fixed error threshold value. Normalized recognition errors may sometimes be referred to as recognition "robustness" scores.

[0017] The weighting criteria / variable thresholds may be settable, for example, via a GUI or DSL.

[0018] In addition to the identified recognition errors, the (normalized) error values may be made accessible via the GUI.

[0019] More complex rules may be applied to identify recognition errors based on one or more error thresholds, for example, by mapping multiple recognition error values or combinations thereof.

[0020] A "recognition error" can be a binary indicator of a recognition error (error / no error), or a non-binary category indicator (e.g., a "traffic light" style classification of red, green, blue).

[0021] The recognition error can be, for example, the number of recognition errors aggregated across multiple objects and / or sensors and / or sensor modalities.

[0022] For example, the recognition error rules may be defined hierarchically. For example, using multiple sensors and / or sensor modalities (e.g., lidar, radar, camera, etc.) and / or multiple objects, aggregated recognition errors aggregated across multiple modalities / objects may be extracted. In this case, multiple recognition error timelines may be derived. For example, by applying certain rules to "lower level" timelines (e.g., those related to specific objects, sensors, and / or sensor modalities), the "top-level" aggregated timeline is input. The top-level timeline may be expandable to view the lower level timelines. The recognition errors may be aggregated over a time window to provide a "zoomed out" view of the driving operation.

[0023] The recognition oracle may be configured to filter at least one time interval of driving, the time interval being omitted from the recognition error timeline, and the filtering is based on one or more filtering criteria applied to recognition errors (e.g., to filter time intervals in which no recognition error occurred), and / or one or more tags / labels associated with driving in the real world (e.g., to include only intervals in which there are specific types of scene elements such as vulnerable road users). For example, the tags may include ontology tags related to dynamic and / or static scene elements or conditions (actors, weather, lighting, etc.). Such filtering may sometimes be referred to as "slicing" the timeline.

[0024] The timeline may aggregate multiple driving runs. Slicing is a useful tool in this context as a way to reduce the scope of "uninteresting" information displayed on the timeline.

[0025] The tags may be accessible via the GUI.

[0026] An overview representation of the driving operation may be displayed on the GUI. The static representation may display a static snapshot of the driving operation at the current time step, which is selectable via commands to the GUI. When the current time step is changed, a visual indicator may be changed to mark the current time step on the recognition error timeline. Along with the overview representation, (raw) data of at least one real-world driving operation may also be displayed. For example, a schematic top-down view may be displayed with at least one 3D point cloud (e.g., lidar, radar, or monocular / stereo depth point cloud, or any combination / aggregation thereof) of the real-world driving operation overlaid. Alternatively or additionally, at least one captured image from one real-world driving operation regarding the current time step may be displayed (changing the current time step causes the GUI to be updated with the corresponding image accordingly).

[0027] The overview representation of the driving operation may be rendered using a time series of the in-driving recognition output. For example, the time series of the in-driving recognition output may include a time series of ground truth bounding boxes (position, orientation, size) for each of a plurality of detected objects and the identified object type of each object, which are used to render a visual icon of the object on a known road layout (e.g., derived from a map) of the driving operation.

[0028] The time series of the in-driving recognition output may also be displayed via the GUI for a visual comparison with the ground truth recognition output. For example, the time series of the in-driving recognition output may be overlaid on a schematic representation derived from the latter. For example, the in-driving recognition output may include a plurality of time series of detected real-time bounding boxes, and a subset of the in-driving bounding boxes associated with the current time step may be overlaid on a snapshot of the current time step.

[0029] The recognized ground truth may be in the form of a trace of each agent (the self-agent and / or other agents), where the trace is a time sequence of spatial states and motion states (e.g., a bounding box and a detected velocity vector or other motion vector).

[0030] The extracted trace may be used to visualize the run in a GUI.

[0031] An option may be provided to dynamically "replay" the scenario in the GUI, and as the scenario progresses, a video indicator moves along the recognition error timeline.

[0032] A second driving performance timeline may also be displayed on the GUI to convey the results of a driving performance assessment applied to the same ground truth recognition output (e.g., a trace). For example, a test oracle may be provided for this purpose.

[0033] The driving data may include two or more of a plurality of sensor modalities, such as lidar, radar, and images (e.g., depth data from stereo or monocular imaging).

[0034] In some embodiments, a certain sensor modality (or combination of sensor modalities) may be used to provide the ground truth of other sensor modalities (or combination of sensor modalities). For example, a more accurate lidar may be used to derive a pseudo ground truth used as a baseline for detection results or other recognition outputs derived from radar or image (monocular or stereo) data.

[0035] A relatively small amount of manually labeled ground truth may be used within the system, for example, as a baseline for verifying or measuring the accuracy of pseudo ground truth or recognition outputs during driving.

[0036] The above considered recognition errors derived from pseudo ground truth. In other aspects of the present invention, however, the above GUI may be used to render recognition errors derived in other ways (derived from real-world data without using pseudo ground truth, and recognition errors of simulated driving runs generated by a simulator). In the case of simulated driving, the above description equally applies to the ground truth directly provided by the simulator (without requiring a ground-truthing pipeline) and the scene variables of the simulated driving.

[0037] A second aspect herein is a computer system for assessing the performance of an autonomous vehicle, at least one input configured to receive performance data of at least one autonomous driving run, the performance data including at least one time series of recognition errors and at least one time series of driving performance results, at least one input, a rendering component configured to generate rendering data for rendering a graphical user interface, the graphical user interface being for visualizing the performance data, (i) a recognition error timeline, and (ii) a driving assessment timeline, comprising a rendering component, comprising, The timeline is temporally aligned and divided into a plurality of time steps of at least one driving run. For each time step, the recognition error timeline includes a visual indication of whether a recognition error occurred at that time step, and the driving assessment timeline includes a visual indication of the driving performance at that time step. A computer system is provided.

[0038] The driving assessment timeline and the recognition error timeline may be parallel to each other.

[0039] The above tools visually link driving performance to recognition errors and assist an expert in determining when the ADS / ADAS performance is low / unacceptable. For example, by focusing on areas of the driving performance timeline where a major driving rule violation occurred, the expert can view the recognition error timeline at the same instant to check whether a recognition error might have influenced the violation of that rule.

[0040] In an embodiment, the driving performance may be assessed with respect to one or more predefined driving rules.

[0041] The driving performance timeline may aggregate driving performance across a plurality of individual driving rules and may be expandable to view the respective driving performance timelines of the individual driving rules.

[0042] The driving performance (or each driving performance) may be expandable to view the computational graph representation of the rules (as described below).

[0043] The driving run may be a real-world run, and the driving rules are applied to the real-world trajectory.

[0044] In some cases, a (pseudo) ground truth trajectory / recognition output may be extracted using a ground truth pipeline, which is used to determine recognition errors and assess performance with respect to driving rules (similar to the first aspect above).

[0045] Alternatively, recognition errors may be identified without using pseudo ground truth. For example, such errors may be identified from "flickering" objects (which appear / disappear when the in-motion object detector fails), or "jumping" objects (which appear to jump within the scene in a kinematically infeasible way, e.g., the in-motion detector may "swap" two nearby objects at a certain point during driving).

[0046] The performance data may include a time series of at least one numerical recognition score indicating the recognition area of interest, and the graphical user interface may comprise at least a corresponding timeline of numerical recognition scores, where for each time step, the numerical recognition score timeline includes a visual indication of the numerical recognition score associated with that time step.

[0047] The time series of numerical recognition scores may be a time series of hardness scores indicating a measure of the difficulty for the recognition system at each time step.

[0048] The performance data may include a time series of at least one user-defined score, and the graphical user interface may comprise at least one corresponding custom timeline, where for each time step, the custom timeline includes a visual indication of the user-defined score evaluated at that time step.

[0049] Alternatively, the driving may be simulated driving, and the recognition errors may be simulated.

[0050] For example, one or more recognition error (or recognition performance) models may be used to sample recognition errors or, more generally, to transform the state of the ground truth simulator into more realistic recognition errors provided to the top-level components of the stack of the device under test during the simulation.

[0051] As another example, synthetic sensor data may be generated in the simulation and processed by the recognition system of the stack in the same way as real sensor data. In this case, the simulated recognition errors can be derived in the same way as real-world recognition errors (however, in this case, the recognition errors can be identified by comparison with the ground truth specific to the simulator, so the ground truth of the pipeline is not necessary).

[0052] Filtering / slicing may also be applied to the timeline so that, for example, only the period near the failure for a particular rule / rule combination is presented. Thus, the recognition error timeline can be filtered / sliced based on the rules applied to the driving performance timeline, and vice versa.

[0053] The graphical user interface may comprise a progress bar aligned with the timeline, the progress bar having one or more markers indicating a certain time interval, each interval including one or more time steps of the driving run. A subset of the markers may be labeled with numerical time indicators.

[0054] The graphical user interface may include a scrubber bar that spans the timeline and indicates a selected time step of the driving operation. In response to the user clicking on a point on the timeline to select a new time step of the driving operation, the scrubber bar may move along the timeline such that the scrubber bar spans the timeline at the selected point.

[0055] The graphical user interface may include a zoom input that can be used to increase or decrease the number of time steps of the driving operation included in the timeline. When using the zoom input to increase or decrease the number of time steps within the timeline, the timeline may be configured such that the visual indicators of each time step are respectively reduced or enlarged so that the timeline maintains a constant length.

[0056] The progress bar may be configured such that when using the zoom input to decrease the number of time steps within the timeline below a threshold, the marker is adjusted to indicate a shorter time interval. When using the zoom input to increase the number of time steps within the timeline beyond the threshold, the marker may be adjusted to indicate a longer time interval.

[0057] When using the zoom input to adjust the number of time steps of the driving operation, the timeline may be adjusted to include only the time steps within a range defined from a reference point on the timeline. The reference point may be the starting point of the driving operation. Alternatively, the reference point may be the currently selected time step of the driving operation. The currently selected point may be indicated by the scrubber bar.

[0058] The zoom input may comprise a zoom slider bar, which may be used to adjust the number of time steps within the timeline by moving an indicator along the slider bar. The indicator may be moved by clicking and dragging the slider along the bar, or by clicking on a point on the slider where the indicator is to be moved to. The zoom input may include a pinch gesture on a touch screen, which adjusts the number of time steps within the timeline based on a change in the distance between two fingers in contact with the screen. Alternatively, the zoom input may include a mouse wheel, which adjusts the number of time steps within the timeline in response to the user rotating the wheel forwards and backwards.

[0059] The timeline may be scrollable, and the plurality of time steps displayed on the timeline may be adjusted to shift temporally forwards and backwards in response to the user's scroll action.

[0060] A portion of the driving may be selected by clicking on a first point on a progress bar indicating the start time of that portion and dragging along the progress bar to a second point defining the end time of that portion. The driving data corresponding to the selected portion may be extracted and stored in a database.

[0061] The first aspect described above refers to testing a real-time recognition system by comparing the derived (simulated) ground truth recognition output set of the recognition output during driving. In other aspects, any of the above features of the embodiments can be applied more generally to evaluate any sequence of recognition outputs by comparison with a corresponding sequence of ground truth recognition outputs. In this context, the ground truth may be any baseline that is considered accurate for the purpose of evaluating the recognition output by comparison with the baseline.

[0062] The third aspect in this specification is At least one input configured to receive data regarding at least one driving run, the data including (i) a time series of first recognition outputs and (ii) a time series of second ground truth recognition outputs, the time series of ground truth recognition outputs and the time series of recognition outputs during driving being related to at least one time interval, at least one input and A rendering component configured to generate rendering data for rendering a graphical user interface (GUI), the graphical user interface comprising a recognition error timeline having a visual indication of recognition errors occurring at that time step for each of a plurality of time steps of at least one driving run, a rendering component and A recognition oracle configured to identify recognition errors occurring in one or more time intervals to generate a recognition error timeline by comparing the time series of recognition outputs with the time series of ground truth recognition outputs and A computer system comprising.

[0063] Note that the term "recognition output" is widely used in this context and includes human annotation as well as recognition data obtained from the output of the vehicle's recognition stack.

[0064] The computer system may further comprise a ground-truthing pipeline. The ground-truthing pipeline may be configured to generate a time series of first recognition outputs by applying at least one non-real-time and / or non-causal recognition algorithm to data of at least one driving run and processing the data, where the data includes a time series of sensor data from the driving run and a time series of related on-road recognition outputs extracted from the time series of sensor data by the recognition system. The ground-truth recognition output may be generated by manual annotation of at least one driving run. In this embodiment, the recognition output generated by the recognition system is a "pseudo" ground-truth recognition output, which may be compared with the manually annotated ground-truth recognition output received for the same driving run to identify recognition errors in the pseudo ground-truth recognition output. This comparison may be used as a method for evaluating the appropriateness of the pseudo ground-truth recognition output obtained from the ground-truthing pipeline, which is used as the ground truth for comparison with other sets of recognition outputs to be evaluated. This comparison may be based only on a subset of the manually annotated driving data, which enables the use of pseudo GT to assess the recognition output of larger data sets for which human annotation is not available.

[0065] Alternatively, the recognition system may comprise a real-time recognition system for deployment in a sensor-equipped vehicle, and the recognition output may include a time series of in-driving recognition outputs extracted from a time series of sensor data for a given driving run by the real-time recognition system. The ground-truth recognition output may be generated by applying at least one non-real-time and / or non-causal recognition algorithm to at least one of the time series of sensor data or the time series of in-driving recognition outputs by a ground-truthing pipeline to process the at least one. Alternatively, the ground-truth recognition output may be generated by manual annotation of the driving run.

[0066] The driving run may be a real-world driving run.

[0067] Alternatively, the driving run may be a simulated driving run, the sensor data may be generated by a simulator, and the in-driving recognition output may be obtained by applying the real-time recognition system to the simulated sensor data. The ground-truth recognition output may be obtained directly from the simulator for comparison with the in-driving recognition output.

[0068] A further aspect herein is a computer-implemented method for testing a real-time recognition system, the real-time recognition system being for deployment in a sensor-equipped vehicle, the method comprising receiving, at an input, data of at least one real-world driving run performed by a sensor-equipped vehicle, the data comprising (i) a time series of sensor data captured by the sensor-equipped vehicle and (ii) a time series of at least one associated in-driving recognition output extracted from the time series of sensor data by the real-time recognition system under test, Generating rendering data for rendering a graphical user interface (GUI) with a recognition error timeline, the recognition error timeline having a visual indication of recognition errors that occurred at a time step for each of a plurality of time steps of at least one real-world driving run, by a rendering component; Processing at least one of (i) a time series of sensor data and (ii) a time series of runtime recognition outputs in a ground truth pipeline by applying at least one non-real-time and / or non-causal recognition algorithm to at least one of them to extract a time series of at least one ground truth recognition output for comparison with the runtime recognition output; Identifying recognition errors that occurred in one or more time intervals to generate a recognition error timeline by comparing a time series of runtime recognition outputs with a time series of ground truth recognition outputs; Providing a computer-implemented method comprising.

[0069] A further aspect provides executable program instructions for programming a computer system to implement any of the methods described herein.

[0070] For a better understanding of the present disclosure, and to show how embodiments thereof may be carried out, reference is made, by way of example only, to the following figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0071]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6A

Figure 6B

Figure 7A

Figure 7B

Figure 8A

Figure 8B

Figure 8C

Figure 9A

Figure 9B-9D

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11

Figure 12A

Figure 12B

Figure 12C

Figure 12D

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18A-18B

Figure 19A

Figure 19B

DETAILED DESCRIPTION OF THE INVENTION

[0072] FIG. 11 shows an exemplary architecture in which a "recognition oracle" 1108 receives recognition error data from multiple sources (real sources and / or simulated sources) and inputs that data into a "recognition triage" graphical user interface (GUI) 500.

[0073] The test oracle 252 assesses driving performance, and a particular implementation of the GUI 500 enables driving performance assessment in conjunction with recognition information on each timeline.

[0074] Particular recognition errors may be derived from the ground truth trajectory of a real or simulated drive, and those same ground truth trajectories are used by the test oracle to assess driving performance.

[0075] The test oracle 252 and the recognition oracle 1108 mirror each other insofar as each applies configurable rule - based logic to input onto the timeline on the GUI 500. The former applies a hierarchical rule tree to the (simulated) ground truth trajectory to assess driving performance over a drive (or multiple drives), and the latter applies similar logic to identify significant recognition errors. The rendering component 1120 generates rendering data for rendering the GUI on a display.

[0076] Our co-pending international patent applications PCT / EP2022 / 053406 and PCT / EP2022 / 053413, which are incorporated herein by reference, describe a domain-specific language (DSL) for encoding rules in a test oracle. An extension of the DSL for encoding rules for identifying significant recognition errors in a recognition oracle is described below.

[0077] The described embodiments provide a test pipeline for facilitating rule-based testing of a mobile robot stack in a real or simulated scenario, which incorporates additional functionality for identifying and communicating the presence of recognition errors in a flexible manner.

[0078] Typically, a "full" stack includes everything from the processing and interpretation (recognition) of lower-level sensor data, to the input to major higher-level functions such as prediction and planning, and the control logic for generating appropriate control signals for making decisions at the planning level (e.g., for controlling brakes, steering, acceleration, etc.). In the case of an autonomous vehicle, a level 3 stack includes logic for implementing a transition request, and a level 4 stack additionally includes logic for implementing a minimum-risk maneuver. The stack may also implement secondary control functions such as signaling, headlights, windshield wipers, etc.

[0079] The term "stack" may refer to individual sub-systems (sub-stacks) of the full stack, such as a recognition, prediction, planning, or control stack, which may be tested individually or in any desired combination. The stack may refer to purely software, i.e., one or more computer programs that can be executed on one or more general-purpose computer processors.

[0080] The test framework described below provides a pipeline for generating scenario ground truth from real-world data. This ground truth may be used as the basis for recognition tests by comparing the generated ground truth to the recognition output of the recognition stack under test, as well as by assessing driving behavior against driving rules.

[0081] The behavior of agents (actors) in real or simulated scenarios is evaluated by a test oracle based on defined performance evaluation rules. Such rules may evaluate various aspects of safety. For example, a set of safety rules may be defined to assess the performance of the stack against specific safety criteria, regulations, or safety models (such as RSS), or a bespoke set of rules may be defined to test any aspect of performance. The test pipeline is not limited in its use to safety and can be used to test any aspect of performance, such as comfort or progress towards a defined goal. A rule editor enables performance evaluation rules to be defined or changed and passed to the test oracle.

[0082] Similarly, the recognition of the vehicle can be evaluated by a "recognition oracle" based on defined recognition rules. These may be defined within a recognition error specification that provides a standard format for defining recognition errors.

[0083] Figure 1 shows a set of possible use cases for an error recognition framework. Defining rules in the error recognition framework enables areas of interest in real-world driving scenarios to be highlighted for the user (1602), for example, by flagging these areas in a replay of the scenario presented to the user interface. This enables the user to review obvious errors in the recognition stack and identify possible reasons for the errors, such as occlusions in the original sensor data. Such evaluation of recognition errors also enables a "contract" to be defined between the recognition component and the planning component of the AV stack (1604), where the requirements for recognition performance can be specified, and a stack that meets these recognition performance requirements promises to enable safe planning. An integrated framework may be used to evaluate real recognition errors from real-world driving scenarios and to evaluate simulated errors, either directly simulated using an error recognition model or calculated by applying the recognition stack to simulated sensor data, such as a realistic simulation of a camera image (1606).

[0084] The ground truth determined by the pipeline itself can be evaluated within the same error recognition specification by comparing it to the "true" ground truth determined by manually reviewing and annotating the scenario according to defined rules (1608). Finally, the results of applying the error recognition test framework can be used to guide a test strategy for testing both the recognition subsystem and the prediction subsystem of the stack (1610).

[0085] A scenario, whether real or simulated, requires the self-agent to move within a real or modeled physical context. The self-agent is a real or simulated mobile robot that moves under the control of the stack under test. The physical context includes static and / or dynamic elements that the stack under test is required to effectively handle. For example, the mobile robot may be a fully or semi-autonomous vehicle (ego vehicle) under the control of the stack. The physical context may include a static road layout and a set of given environmental conditions (e.g., weather, time, lighting conditions, humidity, pollution / particle level, etc.) that can be maintained or changed as the scenario progresses. An interactive scenario additionally includes one or more other agents ("external" agents, e.g., other vehicles, pedestrians, cyclists, animals, etc.).

[0086] Consider the following example as applied to testing an autonomous vehicle. However, the principles are equally applicable to other forms of mobile robots.

[0087] A scenario may be represented or defined at various levels of abstraction. A more abstract scenario is more adaptable to a greater degree of variation. For example, an "interruption scenario" or a "lane change scenario" is an example of a highly abstract scenario characterized by an operation or behavior of interest that is adaptable to many variations (e.g., start positions and speeds of various agents, road layout, environmental conditions, etc.). "Scenario run" refers to a specific occurrence where an agent moves within a physical context, optionally in the presence of one or more other agents. For example, multiple runs of an interruption or lane change scenario can be performed (in the real world and / or in a simulator) with different agent parameters (e.g., start position, speed, etc.), different road layouts, different environmental conditions, and / or different stack configurations. The terms "run" and "instance" are used interchangeably in this context.

[0088] In the following example, the performance of the stack is at least partially evaluated by evaluating the behavior of the self-agent within the test oracle against a given set of performance evaluation rules over one or more runs. The rules are applied to the "ground truth" of the scenario run (or each scenario run), which simply means an appropriate representation of the scenario run (including the behavior of the self-agent) that is generally considered trustworthy for the purposes of the test. The ground truth is specific to the simulation, and the simulator calculates a sequence of scenario states, which by definition is a perfectly trustworthy representation of the simulated scenario run. In a real-world scenario run, there is no "perfect" representation of the scenario run in the same sense, but nevertheless, a ground truth that provides appropriate information can be obtained in a number of ways, for example, through manual annotation of in-vehicle sensor data, automated / semi-automated annotation of such data (using, for example, offline / non-real-time processing), and / or the use of external information sources (such as external sensors, maps, etc.).

[0089] Scenario ground truth typically includes the "trajectory" of the self-agent and, if applicable, any other (notable) agents. A trajectory is the history of an agent's position and movement over a scenario. There are many ways in which a trajectory can be represented. Trajectory data typically includes the spatial and movement data of agents within an environment. This term is used in relation to both real scenarios (with real-world trajectories) and simulated scenarios (with simulated trajectories). A trajectory typically records the actual path realized by an agent within a scenario. In terms of the terms, "trajectory" and "path" may include the same or similar types of information (such as a series of spatial and movement states over time). The term "path" is generally often used in the context of planning (and may refer to a future / predicted path), and the term "trajectory" is generally often used in the context of testing / evaluation in relation to past behavior.

[0090] In a simulation context, a "scenario description" is provided as input to the simulator. For example, the scenario description may be encoded using a scenario description language (SDL) or in any other format that can be used by the simulator. The scenario description is typically a more abstract representation of a scenario and can cause multiple simulated runs. Depending on the implementation, the scenario description may have one or more configurable parameters that can be changed to increase the degree of possible variation. The degree of abstraction and parameterization is a design choice. For example, the scenario description may encode a fixed layout using parameterized environmental conditions (e.g., weather, lighting, etc.). However, further abstraction is possible, for example, using configurable road parameters (e.g., road curvature, lane configuration, etc.). The input to the simulator includes the scenario description along with a set of selected parameter values (if applicable). The latter is sometimes referred to as the parameterization of the scenario. The configurable parameters define a parameter space (also called the scenario space), and the parameterization corresponds to a point within the parameter space. In this context, a "scenario instance" may refer to the instantiation of a scenario in the simulator based on the scenario description and (if applicable) the selected parameterization.

[0091] For the sake of brevity, the term "scenario" may also be used to refer to not only a scenario in a more abstract sense but also a scenario run. The meaning of the term "scenario" will be clear from the context in which it is used.

[0092] Trajectory planning is an important function in the context of the present invention, and the terms "trajectory planner", "trajectory planning system", and "trajectory planning stack" may be used interchangeably herein to refer to one or more components capable of planning the trajectory of a mobile robot for the future. The determination of the trajectory plan ultimately determines the actual trajectory realized by the self-agent (however, in some test contexts, this may be affected by other factors such as the implementation of those decisions in the control stack and the actual or modeled dynamic response of the self-agent to the resulting control signals).

[0093] The trajectory planner may be tested alone or in combination with one or more other systems (e.g., recognition, prediction, and / or control). In the full stack, the planning generally refers to the higher-level autonomous decision-making ability (e.g., trajectory planning), while the control generally refers to the lower-level generation of control signals for implementing those autonomous decisions. However, in the context of performance testing, the term "control" is also used in a broader sense. To avoid misunderstanding, when it is stated that the trajectory planner controls the self-agent in a simulation, it does not necessarily imply that the (narrower sense) control system is being tested in combination with the trajectory planner.

[0094] Exemplary AV stack: To provide context related to the described embodiments, further details of an exemplary form of the AV stack are described here.

[0095] Figure 2A shows a very schematic block diagram of the stack 100 during AV driving. The driving stack 100 is shown to include a recognition (sub)system 102, a prediction (sub)system 104, a planning (sub)system (planner) 106, and a control (sub)system (controller) 108. As described above, the term (sub)stack may also be used to describe the aforementioned components 102-108.

[0096] In a real-world context, the recognition system 102 receives sensor outputs from the in-vehicle sensor system 110 of the AV, uses those sensor outputs to detect external agents, and measures their physical states, such as their position, speed, acceleration, etc. The in-vehicle sensor system 110 can take various forms, but generally includes various sensors such as image capture devices (cameras / optical sensors), lidar and / or radar units, satellite positioning sensors (such as GPS), motion / inertial sensors (accelerometers, gyroscopes, etc.). Therefore, the in-vehicle sensor system 110 provides rich sensor data from which it is possible to extract detailed information about the surrounding environment and the states of the AV and any external actors (vehicles, pedestrians, people on bicycles, etc.) within that environment. Typically, the sensor outputs include sensor data from multiple sensor modalities, such as stereo images from one or more stereo optical sensors, lidar, radar, etc. The sensor data from multiple sensor modalities may be combined using filters, fusion components, etc.

[0097] The recognition system 102 typically includes multiple recognition components that cooperate to interpret the sensor outputs and provide a recognition output to the prediction system 104.

[0098] In a simulation context, depending on the nature of the test, particularly depending on where the stack 100 is "sliced" for the test (see below), there are cases where it is necessary to model the in-vehicle sensor system 100 and cases where it is not. In high-level slicing, since simulated sensor data is not required, complex sensor modeling is not necessary.

[0099] The recognition output from the recognition system 102 is used by the prediction system 104 to predict the future behavior of external actors (agents) such as other vehicles in the vicinity of the AV.

[0100] The prediction calculated by the prediction system 104 is provided to the planner 106, and the planner 106 uses the prediction to make a decision on autonomous driving to be executed by the AV in a given driving scenario. The input received by the planner 106 typically indicates the drivable area and also captures the predicted movement of external agents (obstacles from the perspective of the AV) within the drivable area. The drivable area can be determined by combining the recognition output from the recognition system 102 with map information such as an HD (high-resolution) map.

[0101] The core function of the planner 106 is to plan the trajectory (ego-trajectory) of the AV considering the predicted movement of the agents. This may be referred to as trajectory planning. The trajectory is planned to accomplish the desired goal within the scenario. The goal can be, for example, entering a roundabout and exiting at the desired exit, overtaking the vehicle in front, or staying in the current lane at the target speed (lane following). The goal may be determined, for example, by an autonomous route planner (not shown).

[0102] The controller 108 executes the decisions made by the planner 106 by providing appropriate control signals to the in-vehicle AV actor system 112. Specifically, the planner 106 plans the AV's trajectory, and the controller 108 generates control signals for implementing the planned trajectory. Typically, the planner 106 plans ahead so that the planned trajectory can be implemented only partially at the control level, and then a new trajectory is planned by the planner 106. The actor system 112 includes "primary" vehicle systems such as the brake, acceleration, and steering systems, as well as secondary systems (e.g., signals, wipers, headlights, etc.).

[0103] Note that there may be a difference between the planned trajectory at a given instant and the actual trajectory followed by the self-agent. The planning system typically operates over a sequence of planning steps, updating the trajectory planned at each planning step to account for changes in the scenario since the previous planning step (or, more precisely, changes that deviate from the predicted changes). The planning system 106 may reason forward in time such that the trajectory planned at each planning step deviates from the next planning step. Thus, the individual planned trajectories may not be fully realized (when the planning system 106 is tested alone in a simulation, the self-agent may simply follow the planned trajectory accurately up to the next planning step, but as described above, in other real-world contexts and simulation contexts, the planned trajectory may not be followed accurately up to the next planning step because the behavior of the self-agent may be affected by other factors such as the operation of the control system 108 and the real or modeled dynamics of the host vehicle). In many test contexts, ultimately what matters is the actual trajectory of the self-agent, specifically whether the actual trajectory is safe and other factors such as comfort and progress. However, the rule-based test approach herein can also be applied to the planned trajectories (even if those planned trajectories are not fully or accurately realized by the self-agent). For example, even if the actual trajectory of the agent is considered safe according to a given safety rule, the instantaneous planned trajectory may have been unsafe, and the fact that the planner 106 considered an unsafe course of action may become apparent even if it did not lead to unsafe behavior of the agent within the scenario. The instantaneous planned trajectory constitutes one form of internal state that can usefully be evaluated in addition to the actual behavior of the agent in the simulation. Other forms of internal stack state can similarly be evaluated.

[0104] The example of FIG. 2A contemplates a relatively “modular” architecture having separable perception, prediction, planning, and control systems 102-108. The sub-stack itself may also be modular, for example, having separable planning modules within the planning system 106. For example, the planning system 106 may comprise a plurality of trajectory planning modules that can be applied to different physical contexts (e.g., simple lane driving vs. complex intersections or roundabouts). This is relevant to simulation testing for the reasons above, as it allows components (e.g., the planning system 106 or its individual planning modules, etc.) to be tested individually or in different combinations. To avoid misunderstanding, in a modular stack architecture, the term stack may refer not only to the full stack, but also to its individual sub-systems or modules.

[0105] The degree to which various stack functions are integrated or separable can vary significantly between different stack implementations, and in some stacks, certain aspects may be so closely coupled as to be indistinguishable. For example, in other stacks, planning and control may be integrated (e.g., such a stack may be able to directly perform planning in terms of control signals), while other stacks (e.g., those shown in FIG. 2A) may be designed in a way that makes a clear distinction between these two (e.g., perform planning in terms of trajectories and perform independent control optimization to determine the best way to execute the trajectories planned at the control signal level). Similarly, in some stacks, prediction and planning may be more closely coupled. In extreme cases, in so-called “end-to-end” driving, perception, prediction, planning, and control may be essentially inseparable. Unless otherwise specified, the terms perception, prediction, planning, and control as used herein do not imply any particular coupling or modularization of these aspects.

[0106] The term "stack" encompasses software, but it will be understood that it can also encompass hardware. In simulation, the stack software may be tested on a "general-purpose" non-vehicle computer system before ultimately being uploaded to the in-vehicle computer system of the physical vehicle. However, in "hardware-in-the-loop" testing, the testing may extend to the hardware that forms the basis of the vehicle itself. For example, the stack software may be executed on an in-vehicle computer system (or a replica thereof) coupled to a simulator for testing purposes. In this context, the stack being tested extends to the computer hardware that forms the basis of the vehicle. As another example, a particular function of the stack 100 (e.g., a recognition function) may be implemented in dedicated hardware. In a simulation context, hardware-in-the-loop testing can include supplying synthetic sensor data to the recognition component of the dedicated hardware.

[0107] Exemplary test paradigm: Figure 2B shows a very general overview of a test paradigm for an autonomous vehicle. For example, an ADS / ADAS stack 100 of the type shown in Figure 2A undergoes repeated testing and evaluation in simulation by running multiple scenario instances in a simulator 202 and evaluating the performance of the stack 100 (and / or its individual sub-stacks) with a test oracle 252. The output of the test oracle 252 is useful to an expert 122 (team or individual), enabling the expert 122 to identify issues within the stack 100 and modify the stack 100 to mitigate those issues (S124). This result also helps the expert 122 select additional scenarios for testing (S126), and the process continues by repeatedly modifying, testing, and evaluating the stack 100 in simulation. The improved stack 100 is ultimately incorporated into a real-world AV101 equipped with a sensor system 110 and an actuator system 112 (S125). The improved stack 100 typically includes program instructions (software) executed by one or more computer processors of an in-vehicle computer system (not shown) of the vehicle 101. The software of the improved stack is uploaded to the AV101 in step S125. Step S125 may also include changes to the underlying vehicle hardware. Once the improved stack 100 is installed on the AV101, it receives sensor data from the sensor system 110 and outputs control signals to the actuator system 112. Real-world testing (S128) can be used in combination with simulation-based testing. For example, once an acceptable level of performance is reached through the simulation testing and stack improvement process, appropriate real-world scenarios may be selected (S130), the performance of the AV101 in those real scenarios may be captured and similarly evaluated with the test oracle 252.

[0108] Scenarios can be obtained in various ways, including manual coding, for simulation purposes. This system is also capable of extracting scenarios from real-world driving for simulation purposes, enabling real-world situations and their variations to be recreated within the simulator 202.

[0109] Figure 2C shows a very schematic block diagram of the scenario extraction pipeline. Real-world driving data 140 is passed to a "ground-truthing" pipeline 142 for the purpose of generating scenario ground truth. The driving data 140 can include, for example, sensor data and / or recognition outputs captured / generated on one or more vehicles (which can be autonomous, human-driven, or a combination thereof), and / or data captured from other sources such as external sensors (e.g., CCTV). The driving data is processed within the ground-truthing pipeline 142 to generate appropriate ground truth 144 (trajectory and context data) for real-world driving. As discussed, the ground-truthing process can be based on manual annotation of "raw" driving data 140, or the process can be fully automated (e.g., using offline recognition methods), or a combination of manual and automated ground-truthing can be used. For example, 3D bounding boxes can be placed around the vehicles and / or other agents captured in the driving data 140 to determine the spatial and motion states of their trajectories. The scenario extraction component 146 receives the scenario ground truth 144 and processes the scenario ground truth 144 to extract a more abstracted scenario description 148 that can be used for simulation purposes. The scenario description 148 is used by the simulator 202 to enable multiple simulated drives to be performed. The simulated drives are a variation of the original real-world drive, and the degree of possible variation is determined by the degree of abstraction. Ground truth 150 is provided for each simulated drive.

[0110] The actual scenario ground truth 144 and the simulated ground truth 150 may be processed by a recognition triage tool 152 to evaluate the recognition stack and / or may be processed by a test oracle 252 to assess the stack based on the ground truth 144 or the simulated ground truth 150.

[0111] In the current non-vehicle context, it is not necessary for the trajectory to be extracted in real time (or rather, it is not necessary for the trajectory to be extracted to support real-time planning), rather, the trajectory is extracted "offline". Examples of offline recognition algorithms include non-real-time and non-causal recognition algorithms. Offline technology is in contrast to "online" technology that can be implemented executably within the AV stack 100 to facilitate real-time planning / decision-making.

[0112] For example, it is possible to use non-real-time processing that cannot be executed online due to hardware or other practical constraints of the in-vehicle computer system of the AV. To extract the trajectory, for example, one or more non-real-time recognition algorithms can be applied to the real-world driving data 140. A non-real-time recognition algorithm can be an algorithm that is thought to be impossible to operate in real time due to required computational resources or memory resources.

[0113] In this context, it is also possible to use "non-causal" recognition algorithms. Non-causal algorithms may or may not be able to operate in real time at runtime. However, since they require future knowledge, they cannot be implemented in an online context. For example, a recognition algorithm that detects the state of an agent (e.g., position, orientation, speed, etc.) at a specific moment based on subsequent data requires future knowledge (except when constrained to operate with a short look-ahead window), and thus cannot support real-time planning within stack 100 in an online context. For example, filtering using a backward pass may be able to operate in real time, but it is a non-causal algorithm that requires future knowledge.

[0114] The term "recognition" generally refers to techniques for recognizing structures within real-world data 140, such as 2D or 3D bounding box detection, position detection, orientation detection, motion detection, etc. For example, a trajectory may be extracted as a time series of bounding boxes or other spatial states in 3D or 2D space (e.g., in a reference frame of an aerial view), along with associated motion information (e.g., speed, acceleration, jerk, etc.).

[0115] Ground truth pipeline The problem when testing the performance of an autonomous vehicle stack in the real world is that the autonomous vehicle generates a huge amount of data. This data can be used later to analyze or evaluate the performance of the AV in the real world. However, a potential challenge is to search for relevant data within this footage and identify events of interest that occurred during driving. One option is to manually analyze the data and identify events of interest through human annotation. However, this can be costly.

[0116] Figure 3 shows an example of manually tagging real-world driving data during operation. The AV is equipped with sensors such as, for example, a camera. As shown in the exemplary image 1202, video is collected along the drive by the camera. In an exemplary drive on a highway by a human driver, if the driver notices something of interest, the driver can provide a flag to the AV and tag that frame within the data collected by the sensors. The image shows a visualization of the drive on map 1200, and the bubbles indicate points along the drive that the driver has tagged something. In this example, each tagged point corresponds to a frame of the camera image, which is used to filter the data to be analyzed after the drive so that only the tagged frames are inspected later.

[0117] As shown in map 1200, there are large gaps between the tagged frames in the driving route, and none of the data collected in these gaps is tagged, so this data is not used. By filtering the data using manual annotation by the driver of the vehicle, subsequent analysis of the driving data is limited to events that the human driver or test engineer felt were important enough to flag, or events for which there was enough time to flag. However, there may be useful insights regarding the vehicle's performance at other points in the remaining data, and it would be useful to determine an automated method for more fully processing and evaluating driving performance. Furthermore, identifying more issues for the same amount of data than manual tagging provides more opportunities for the AV system to make improvements for the same amount of collected data.

[0118] A possible solution is to create an integrated analysis pipeline that uses the same metrics to evaluate both scenario simulations and real-world driving. The first step is to extract driving trajectories from the actually collected data. For example, the approximate positions of the host vehicle and other agents can be estimated based on on-vehicle detection results. However, the on-vehicle detection results are not perfect, because computing resources are limited and the on-vehicle detection results function in real time, that is, the data informing a given detection result is only what the sensor has observed up to that point. This means that the detection results may contain noise and be inaccurate.

[0119] Figure 4A shows how data is processed and refined within the data ingestion pipeline to determine a pseudo ground truth 144 for a given set of real-world data. Note that the "true" ground truth cannot be extracted from real-world data, and the ground truth pipeline described herein provides an estimate of the ground truth sufficient for evaluation. This pseudo ground truth 144 may sometimes be simply referred to as "ground truth" herein.

[0120] The data ingestion pipeline (or "ingestion" tool) takes in recognition data 140 from a given stack and optionally from other arbitrary data sources 1300 such as manual annotations, purifies the data, and extracts the pseudo ground truth 144 of the real-world driving scenarios captured in the data. As shown in the figure, sensor data and detection results from the vehicle are ingested, optionally along with additional inputs such as offline detection results or manual annotations. These are processed to apply the offline detector 1302 to the raw sensor data and / or to perform purification 1304 of the detection results received from the vehicle's on-board recognition stack. The purified detection results are then output as the pseudo ground truth 144 of the scenario. This can then be used as the basis for various use cases, including evaluating the ground truth against the driving rules by a test oracle (described later), determining recognition errors by comparing the vehicle detection results with the pseudo ground truth, and extracting scenarios for simulation. Other metrics may be calculated for the input data, including a recognition "difficulty" score 1306, which can be applied, for example, to the detection results or the entire camera image and indicates the difficulty for the recognition stack to handle the given data correctly.

[0121] Figure 4B shows an example of a set of bounding boxes before and after purification. In the example of Figure 4B, the upper image shows a set of "unpurified" noisy 3D bounding boxes that define the position and orientation of the vehicle at each time step, and these bounding boxes represent the ground truth with noise added. The example shown corresponds to bounding boxes with added noise, but the same effect is obtained when purifying vehicle detection results from a real-world driving stack. As shown in Figure 4B, the bounding boxes contain noise, and both the position and orientation of the detected bounding boxes vary over time due to recognition errors.

[0122] The refinement pipeline can use various methods to remove this noise. The lower track in Figure 4B shows the pseudo ground truth trajectory 144 of the vehicle with the noise removed. As shown in the figure, the orientation and position of the vehicle are consistent for each frame, forming a smooth driving trajectory. The multiple possible methods used by the pipeline to perform this smoothing are not described in detail. However, the pipeline benefits from a computing power that is superior to that of the online detector, enabling a more accurate detector to be used, and also benefits from using past and future detection results to smooth the trajectory. The real-world detection results collected from the vehicle function in real time and are based only on past data. For example, if an object is partially occluded at time t but becomes fully visible by the vehicle's sensors at time t + n, an offline refinement pipeline can use the detection result at time t + n to inform an earlier detection result based on the partially occluded data, resulting in a more complete detection result overall.

[0123] Various types of offline detectors or detection result refinement methods can be used. Figure 5A shows a table of possible detection result refinement techniques, and Figure 5B shows a table of possible offline detectors that can be applied to sensor data to obtain improved detection results.

[0124] Various techniques are used to refine the detection results. One example is semantic keypoint detection applied to camera images. After refinement, the result becomes a stable detection result with a rectangular cuboid of appropriate size to smoothly track the vehicle, as shown in Figure 4B for example.

[0125] Reference is made to International Patent Publication No. 2021 / 013792, which is incorporated herein by reference. The above-cited document discloses a class of offline annotation methods that can be implemented within a ground-truthing pipeline 400 to extract the ground-truth trajectories of each agent of interest. Trajectories are extracted by applying automated annotation techniques to annotate the real-world driving data 140 with a sequence of refined 3D bounding boxes (in this case, the agent trajectories are composed of the refined 3D boxes).

[0126] This method functions roughly as follows. The real-world driving data 140 is composed of a sequence of frames, and each frame is composed of a set of 3D structural points (e.g., a point cloud). Each agent of interest (ego agent and / or other agents) is tracked as an object over multiple frames (this agent is a "common structure component" in the terms of the above-cited document).

[0127] The "frame" in this context refers to any captured 3D structural representation, i.e., something that includes the captured points (3D structural points) that define the structure in 3D space, which provides a basically static "snapshot" (i.e., a static 3D scene) of the 3D structure captured in that frame. A frame may be said to correspond to a single instant, but this does not necessarily imply that the frame or the underlying sensor data from which the frame is derived needs to be captured instantaneously. For example, lidar measurements may be captured over a short interval (e.g., about 100 ms) by a moving body with a lidar sweep and "untwisted" considering the movement of the moving body to form a single point cloud. Even in that case, the single point cloud may be said to correspond to a single instant.

[0128] Real-world driving data may be composed of multiple frame sequences, for example, two or more distinct sequences among lidar, radar, depth frames (where depth frames in this context refer to 3D point clouds derived by depth imaging such as stereo or monocular depth imaging). The frames can also be composed of fused point clouds calculated by fusing multiple point clouds from different sensors and / or different sensor modalities.

[0129] This method starts with an initial set of 3D bounding box estimates (coarse size / pose estimates) for each agent of interest and uses these estimates to construct a 3D model of that agent from the frame itself. Here, pose refers to a 6D pose (3D position and orientation in 3D space). The following example considers specifically the extraction of 3D models from lidar, but this explanation applies equally to other sensor modalities. Multiple modalities of sensor data can be used, for example, a rough 3D box can be provided by a first or second sensor modality or multiple modalities (e.g., radar or depth imaging). For example, an initial coarse estimate can be calculated by applying a 3D bounding box detector to the point cloud of a second modality (or multiple modalities). The coarse estimate can also be determined from the same sensor modality (in this case lidar) and this estimate is refined using subsequent processing techniques. As another example, a real-time 3D box from the recognition system 102 of the test object can be used as an initial coarse estimate (e.g., when calculated on a vehicle during real-world driving). In the case of the latter approach, this method can be described as a form of detection result refinement.

[0130] To create an aggregated 3D object model for each agent, a subset of the points contained in the rough 3D bounding box of each frame is obtained (or the rough 3D bounding box may be slightly enlarged to provide additional "margin" for object point extraction), and the points belonging to that object are aggregated over multiple frames. Generally speaking, the aggregation functions by first transforming the subset of points from each frame to the agent's reference frame. Since the pose of the agent within each frame is only roughly known, the transformation to the agent reference frame is not accurately known at this point. The transformation is first estimated from the rough 3D bounding box. For example, the transformation can be efficiently implemented by transforming the subset of points according to the axes of the rough 3D bounding box within each frame. Subsets of points from different frames mostly belong to the same object, but may be misaligned within the agent reference frame due to errors in the initial pose estimate. To correct for the misalignment, a registration method is used to register two subsets of points. Such a method generally functions by using some form of matching algorithm (e.g., Iterative Closest Point) to transform (rotate / translate) one of the subsets of object points to align it with the other. The matching uses the knowledge that most of the two subsets of points are from the same object. This process is then repeated over subsequent frames to build a dense 3D model of the object. By building such a dense 3D model, noisy points (those not belonging to the object) are isolated from the actual object points and can be more easily filtered out.Next, by applying a 3D object detector to the densely filtered 3D object model, a 3D bounding box of a more accurate size that fits exactly to the agent in question can be determined (this assumes a rigid agent where the size and shape of the 3D bounding box do not change between frames and the variables of each frame are only its position and orientation). Finally, an aggregated 3D model is matched with the corresponding object points within each frame, and by accurately specifying the position of the more accurate 3D bounding box in each frame, a refined 3D bounding box estimate (forming part of the pseudo ground truth) for each frame is provided. This process can be iteratively repeated, whereby an initial 3D model is extracted, the pose is refined, the 3D object model is updated based on the refined pose, and so on.

[0131] The refined 3D bounding box serves as a pseudo ground truth position state when determining the degree of recognition error of position-based recognition outputs (such as boxes during driving, pose estimates, etc.).

[0132] To incorporate motion information, the 3D bounding box may be optimized congruently with a 3D motion model. The motion model can provide the motion state of the agent in question (such as speed / velocity, acceleration, etc.), and the motion state may be used as a pseudo ground truth for motion detection results during driving (such as speed / velocity, acceleration estimates calculated by the recognition system 102 under test). The motion model can facilitate a realistic (kinematically achievable) 3D box across frames. For example, the congruent optimization can be formulated based on a cost function that penalizes the mismatch between the aggregated 3D model and the points of each frame, while at the same time penalizing kinematically infeasible changes in the agent pose between frames.

[0133] The motion model also enables the 3D box to be accurately localized in frames that include object detection misses (i.e., when a rough estimate is not available, which can occur when the rough estimate is the detection result on the vehicle and the recognition system 102 under test fails at a given frame) by interpolating the poses of 3D agents between adjacent frames based on the motion model. This enables object detection misses to be identified within the recognition triage tool 152.

[0134] The 3D model can be in the form of a point cloud or a surface model (e.g., a distance field) may be fitted to the points. International Patent Publication No. 2021 / 013791, which is incorporated herein by reference, discloses further details of 3D object modeling techniques where the 3D surface of a 3D object model is encoded as a (signed) distance field fitted to the extracted points.

[0135] The use of these refinement techniques is that they can be used to obtain a pseudo ground truth 144 of the agents in a scene including the host vehicle and external agents, and that the refined detection results can be treated as the actual trajectories taken by the agents in the scene. This may be used to assess how accurate the on-vehicle recognition of the vehicle is by comparing the vehicle's detection results with the pseudo ground truth. The pseudo ground truth can also be used to confirm how the system under test (i.e., the host vehicle stack) has violated highway rules while driving.

[0136] The pseudo ground truth detection result 144 can also be used to perform semantic tagging and querying of the collected data. For example, a user can enter a query such as "find all events with interruptions", where an interruption is when an agent enters in front of the host vehicle in the host vehicle's lane. Since the pseudo ground truth has the trajectories of all agents in the scene, including the position and orientation at any given time, it is possible to identify an interruption by searching within the agent trajectories for instances where an agent has entered in front of another vehicle in a given lane. More complex queries may be constructed. For example, a user may enter a query such as "find all interruptions where the agent had a speed of at least x". Since the agent's motion is defined by the pseudo ground truth trajectories extracted from the data, it is straightforward to search within the refined detection results for instances of interruptions where the agent exceeded a given speed. When these queries are selected and executed, the time required to manually analyze the data is reduced. This means that the driver does not need to rely on identifying areas of interest in real time, and instead, areas of interest can be automatically detected within the collected data, from which scenarios of interest can be extracted for further analysis. This allows for more data to be used and, in some cases, enables the identification of scenarios that might otherwise be missed by a human driver.

[0137] Test pipeline: Next, further details of the test pipeline and test oracle 252 are described. The following examples focus on simulation-based testing. However, as noted above, the test oracle 252 can equally be applied to evaluate stack performance in real-world scenarios, and the following related explanations equally apply to real-world scenarios. Specifically, the test pipeline described below may be used with the extracted ground truth 144 obtained from real-world data as described in FIGS. 1-5. Applying the described test pipeline with a recognition evaluation pipeline in a real-world data analysis tool will be described in more detail later. The following explanation refers to stack 100 of FIG. 2A by way of example. However, as noted above, the test pipeline 200 is highly flexible and can be applied to any stack or sub-stack operating at any level of autonomy.

[0138] FIG. 6A shows a schematic block diagram of a test pipeline represented by reference numeral 200. The test pipeline 200 is shown to include a simulator 202 and a test oracle 252. The simulator 202 runs a simulated scenario for the purpose of testing all or part of the stack 100 during AV driving, and the test oracle 252 evaluates the performance of the stack (or sub-stack) in the simulated scenario. As discussed, only a sub-stack of the running stack may be tested, but for simplicity, the following explanation refers throughout to the (full) AV stack 100. However, this explanation equally applies to a sub-stack instead of the full stack 100. The term "slicing" is used herein to refer to the selection of a set or subset of stack components for testing.

[0139] As described above, the idea of simulation-based testing is to run a simulated driving scenario in which the ego-agent has to proceed under the control of the stack 100 under test. Typically, the scenario includes a static drivable area (e.g., a specific static road layout) where the ego-agent has to proceed, typically in the presence of one or more other dynamic agents (e.g., other vehicles, bicycles, pedestrians, etc.). For this purpose, the simulated input 203 is provided from the simulator 202 to the stack 100 under test.

[0140] Stack slicing determines the form of the simulated input 203. As an example, FIG. 6A shows the prediction, planning, and control systems 104, 106, and 108 within the AV stack 100 under test. For testing the full AV stack of FIG. 2A, the recognition system 102 can also be applied during testing. In this case, the simulated input 203 includes synthetic sensor data that is generated using an appropriate sensor model and processed within the recognition system 102 in the same way as real sensor data. This requires the generation of sufficiently realistic synthetic sensor inputs (e.g., realistic image data and / or similarly realistic simulated lidar / radar data, etc.). The resulting output of the recognition system 102 is then fed to the higher-level prediction and planning systems 104, 106.

[0141] In contrast, so-called "planning-level" simulations basically bypass the recognition system 102. Instead, the simulator 202 directly provides a simpler higher-level input 203 to the prediction system 104. In some contexts, it may even be appropriate to bypass the prediction system 104 in order to test the planner 106 based on predictions directly obtained from the simulated scenario (i.e., "perfect" predictions).

[0142] Between these two extremes, there is room for many different levels of input slicing. For example, testing only a subset of the recognition system 102, such as only the "later" (higher-level) recognition components, or testing components such as filters or fusion components that act on the outputs from lower-level recognition components (such as object detectors, bounding box detectors, motion detectors, etc.).

[0143] In any form, the simulated input 203 is used (either directly or indirectly) as the basis for decision-making by the planner 108. The controller 108 then implements the planner's decision by outputting a control signal 109. In the real-world context, these control signals drive the physical actuator system 112 of the AV. In simulation, the vehicle dynamics model 204 is used to simulate the physical response of the autonomous vehicle to the control signal 109 by converting the resulting control signal 109 into realistic movement of the ego-agent within the simulation.

[0144] Alternatively, a simpler form of simulation assumes that the ego-agent exactly follows each planned trajectory between planning steps. This approach bypasses the control system 108 (to the extent separable from the planning) and removes the need for the vehicle dynamics model 204. This may be sufficient for testing certain aspects of the plan.

[0145] Within the range where the external agent exhibits autonomous behavior / decision-making within the simulator 202, some form of agent decision logic 210 is implemented to make those decisions and determine the agent's behavior within the scenario. The agent decision logic 210 may be of the same complexity as the self-stack 100 itself or may have more limited decision-making capabilities. The aim is to provide a sufficiently realistic external agent behavior within the simulator 202 in order to usefully test the decision-making capabilities of the self-stack 100. In some contexts, this may not require any agent decision logic 210 at all (open-loop simulation), while in other contexts, relatively limited agent logic 210 such as basic adaptive cruise control (ACC) can be used to provide useful tests. Where appropriate, one or more agent dynamics models 206 may be used to provide more realistic agent behavior.

[0146] The scenario is run according to the scenario description 201a of the scenario and (where applicable) the selected parameterization 201b. The scenario typically has both static and dynamic elements, which may be "hard-coded" within the scenario description 201a or may be configurable and thus determined by the scenario description 201a in combination with the selected parameterization 201b. In a driving scenario, the static elements typically include a static road layout.

[0147] The dynamic elements typically include one or more external agents within the scenario, such as other vehicles, pedestrians, bicycles, etc.

[0148] The range of dynamic information provided to the simulator 202 for each external agent can be varied. For example, a scenario may be described by separable static and dynamic layers. To provide various scenario instances, a given static layer (e.g., defining a road layout) can be used in combination with various dynamic layers. The dynamic layer may include, for each external agent, the spatial path traversed by that agent, along with one or both of the motion data and behavior data associated with that path. In a simple open-loop simulation, the external actor simply follows the spatial path and motion data defined in the dynamic layer that is non-reactive, i.e., does not react to the ego agent within the simulation. Such an open-loop simulation can be implemented without the agent decision logic 210. However, in a closed-loop simulation, the dynamic layer instead defines at least one behavior (e.g., the ACC behavior) to be followed along the static path. In this case, the agent decision logic 210 implements that behavior in a reactive manner within the simulation, i.e., reactively to the ego agent and / or other external agents. The motion data may still be associated with the static path, but in this case is less prescriptive and may, for example, serve as a goal along the path. For example, in the ACC behavior, a target speed can be set along the path that the agent attempts to match, but the agent decision logic 210 may be permitted to reduce the speed of the external agent below the target at any point along the path to maintain a target distance from the leading vehicle.

[0149] As will be appreciated, a scenario can be described in many ways with any degree of configurability for the purposes of the simulation. For example, the number and type of agents, as well as their motion information, may be configurable as part of the scenario parameterization 201b.

[0150] The output of the simulator 202 for a given simulation includes the self-trajectory 212a of the self-agent and one or more agent trajectories 212b (trajectories 212) of one or more external agents. Each trajectory 212a, 212b is a complete history of the behavior of the agent within the simulation having both a spatial component and a motion component. For example, each trajectory 212a, 212b may take the form of a spatial path having motion data associated with points along the path, such as speed, acceleration, jerk (rate of change of acceleration), snap (rate of change of jerk), etc.

[0151] Additional information is also provided to supplement the trajectories 212 and provide context thereto. Such additional information is referred to as "context" data 214. The context data 214 relates to the physical context of the scenario and can have both static components (e.g., road layout) and dynamic components (e.g., weather conditions within a range that changes over the simulation). Since the context data 214 is directly defined by the selection of the scenario description 201a or the parameterization 201b, it may be somewhat "pass-through" in that it is not affected by the results of the simulation. For example, the context data 214 may include a static road layout directly provided by the scenario description 201a or the parameterization 201b. However, typically, the context data 214 includes at least some elements derived within the simulator 202. This can include, for example, simulated environmental data such as weather data, and the simulator 202 can freely change the weather conditions as the simulation progresses. In that case, the weather data may be time-dependent and that time-dependence is reflected in the context data 214.

[0152] Test oracle 252 receives trajectory 212 and context data 214 and scores their outputs with respect to a set of performance evaluation rules 254. The performance evaluation rules 254 are shown to be provided as inputs to test oracle 252.

[0153] Rule 254 is typically categorical (e.g., pass / fail type rules). Certain performance evaluation rules are also associated with numerical performance metrics used to "score" the trajectory (e.g., degree of achievement or failure, or other quantities that help explain or are otherwise related to the categorical result), such as the degree of achievement or failure, or other quantities that help explain or are otherwise related to the categorical result. The evaluation of Rule 254 is time-based, and a given rule may have different results at different points in time within a scenario. Scoring is also time-based, and for each performance evaluation metric, the test oracle 252 tracks how the value (score) of that metric changes over time as the simulation progresses. The test oracle 252 provides an output 256 that includes, as will be described in more detail later, a time sequence 256a of the categorical (e.g., pass / fail) results for each rule and a score-time plot 256b for each performance metric. The results and scores 256a, 256b are useful to the expert 122 and can be used to identify and mitigate performance issues within the tested stack 100. The test oracle 252 also provides an overall (aggregate) result for the scenario (e.g., overall pass / fail). The output 256 of the test oracle 252 is stored in the test database 258 in association with information about the scenario to which the output 256 pertains. For example, the output 256 may be stored in association with the scenario description 210a (or its identifier) and the selected parameterization 201b. Similar to the time-dependent results and scores, an overall score may also be assigned to the scenario and stored as part of the output 256. For example, an aggregate score for each rule (e.g., overall pass / fail), and / or an aggregate result (e.g., pass / fail) across all rules 254.

[0154] Figure 6B shows another selection for slicing and uses reference numerals 100 and 100S to represent a full stack and a sub-stack, respectively. The sub-stack 100S is the subject of testing within the test pipeline 200 of Figure 6A.

[0155] Some of the "late-stage" recognition components 102B form part of the sub-stack 100S being tested and are applied to the simulated recognition input 203 during testing. The late-stage recognition components 102B can include filtering or other fusion components that fuse recognition inputs from multiple early recognition components.

[0156] In the full stack 100, the late-stage recognition components 102B receive the actual recognition input 213 from the early recognition components 102A. For example, the early recognition components 102A may comprise one or more 2D or 3D bounding box detectors, in which case the simulated recognition input provided to the late-stage recognition components can include simulated 2D or 3D bounding box detection results derived by ray tracing in the simulation. The early recognition components 102A generally include components that act directly on the sensor data. In the slicing of FIG. 6B, the simulated recognition input 203 formally corresponds to the actual recognition input 213 normally provided by the early recognition components 102A. However, instead of being applied as part of the test, the early recognition components 102A are used to train one or more recognition error models 208, and the recognition error models 208 can be used to statistically introduce realistic errors in a rigorous manner into the simulated recognition input 203 supplied to the late-stage recognition components 102B of the sub-stack 100 under test.

[0157] Such an error recognition model may be referred to as a Perception Statistical Performance Model (PSPM), or equivalently as "PRISM". Further details of the principles of PSPM, and suitable techniques for constructing and training PSPM, can be found in International Patent Publications Nos. 2021037763, 2021037760, 2021037765, 2021037761, and 2021037766, the entireties of each of which are incorporated herein by reference. The idea behind PSPM is to efficiently introduce realistic errors into the simulated recognition inputs provided to sub-stack 100S (i.e., to reflect the types of errors that would be expected if the early recognition component 102A were applied in the real world). In the simulation context, the simulator provides "perfect" ground truth recognition inputs 203G, but these are used to derive more realistic recognition inputs 203 that have realistic errors introduced by the error recognition model 208.

[0158] As described in the foregoing cited references, PSPM can depend on one or more variables ("interference factors") that represent physical conditions, enabling different levels of error to be introduced that reflect the various real-world conditions that can occur. Thus, the simulator 202 can simulate different physical conditions (e.g., different weather conditions) by simply changing the values of the weather interference factors and thereby changing how the introduction of recognition errors occurs.

[0159] The late recognition component 102b within sub-stack 100S processes the simulated recognition inputs 203 in exactly the same way as it processes real-world recognition inputs 213 within full-stack 100, and its output drives prediction, planning, and control.

[0160] Alternatively, PRISM can be used to model the entire recognition system 102, including the late recognition component 102B, in which case the PSPM is used to generate a realistic recognition output that is passed directly to the prediction system 104 as input.

[0161] Depending on the implementation, there may or may not be a deterministic relationship between a given scenario parameterization 201b and the results of a simulation with a given configuration of the stack 100 (i.e., the same parameterization may or may not always lead to the same results with the same stack 100). Nondeterminism can occur in various ways. For example, if the simulation is based on PRISM, PRISM may model the distribution of possible recognition outputs for each given time step of the scenario, from which a realistic recognition output is probabilistically sampled. This leads to nondeterministic behavior within the simulator 202, such that different recognition outputs are sampled and thus different results may be obtained for the same stack 100 and scenario parameterization. Alternatively or additionally, the simulator 202 may be inherently nondeterministic, e.g., weather, lighting, or other environmental conditions may be randomized / probabilistic to some extent within the simulator 202. As will be understood, this is a design choice and in other implementations, instead, various environmental conditions can be fully specified in the scenario parameterization 201b. In a nondeterministic simulation, multiple scenario instances can be run for each parameterization. For a given choice of parameterization 201b, aggregate pass / fail results can be assigned, e.g., as a count or percentage of pass / fail results.

[0162] The test orchestration component 260 is responsible for selecting scenarios for simulation purposes. For example, the test orchestration component 260 may automatically select scenario descriptions 201a and appropriate parameterizations 201b based on the test oracle output 256 from previous scenarios.

[0163] Test oracle rules: The performance evaluation rule 254 is constructed as a computational graph (rule tree) applied within the test oracle. Unless otherwise specified, the term "rule tree" in this specification refers to a computational graph configured to implement a given rule. Each rule is constructed as a rule tree, and a set of multiple rules may be referred to as a "forest" of multiple rule trees.

[0164] FIG. 7A shows an example of a rule tree 300 constructed from a combination of extractor nodes (leaf objects) 302 and accessor nodes (non-leaf objects) 304. Each extractor node 302 extracts a time-varying numerical (e.g., floating-point) signal (score) from a set of scenario data 310. The scenario data 310 is a form of scenario ground truth in the sense described above and may be referred to as such. The scenario data 310 is obtained by deploying a trajectory planner (e.g., planner 106 of FIG. 2A) to a real or simulated scenario and is shown to include self and agent trajectories 212 as well as context data 214. In the simulation context of FIG. 6 or FIG. 6A, the scenario ground truth 310 is provided as the output of the simulator 202.

[0165] Each asserter node 304 is shown to have at least one child object (node), and each child object is one of the extractor nodes 302 or another one of the asserter nodes 304. Each asserter node receives outputs from its child nodes and applies an asserter function to those outputs. The output of the asserter function is a time series of categorical results. The following example considers simple binary pass / fail results, but the technology can be easily extended to non-binary results. Each asserter function evaluates the outputs of its child nodes against predetermined atomic rules. Such rules can be flexibly combined according to the desired safety model.

[0166] In addition, each asserter node 304 derives a time-varying numerical signal from the outputs of its child nodes, which is associated with a categorical result by a threshold condition (see below).

[0167] The topmost root node 304a is an asserter node that is not a child node of any other node. The topmost node 304a outputs a sequence of final results, and its descendants (i.e., the nodes that are direct or indirect children of the topmost node 304a) provide the underlying signals and intermediate results.

[0168] FIG. 7B visually shows an example of a time series of the derived signal 312 and the corresponding results 314 calculated by the asserter node 304. The results 314 are correlated with the derived signal 312 in that a pass result is returned only if the derived signal exceeds a fail threshold 316. As will be understood, this is only an example of a threshold condition that associates a time sequence of results with a corresponding signal.

[0169] The signal directly extracted from the scenario ground truth 310 by the extractor node 302 may be referred to as the "raw" signal to distinguish it from the "derived" signal calculated by the assessor node 304. The results and the raw / derived signals may be discretized in time.

[0170] FIG. 8A shows an example of a rule tree implemented within the test platform 200.

[0171] The rule editor 400 is provided to construct the rules implemented in the test oracle 252. The rule editor 400 receives rule creation input from a user (which may or may not be the end user of the system). In this example, the rule creation input is encoded in a domain specific language (DSL) and defines at least one rule graph 408 implemented within the test oracle 252. In the following example, the rules are logical rules and true and false represent pass and fail respectively (as will be understood, this is purely a design choice).

[0172] Consider the following example of a rule formulated using a combination of atomic logical predicates. Examples of basic atomic predicates include elementary logical gates (OR, AND, etc.) and logical functions such as "greater than", (Gt(a,b)) which returns true if a is greater than b and false otherwise.

[0173] The Gt function is for implementing the safe lateral distance rule between the ego agent and other agents (with agent identifier "other_agent_id") within the scenario. Two extractor nodes (latd, latsd) apply the LateralDistance and LateralSafeDistance extractor functions respectively. These functions act directly on the Scenario Ground Truth 310 to extract a time-varying lateral distance signal (measuring the lateral distance between the ego agent and the identified other agent) and a time-varying safe lateral distance signal regarding the ego agent and the identified other agent respectively. The safe lateral distance signal can depend on various factors such as the speeds of the ego agent (captured in the trajectory 212) and the other agent, as well as the environmental conditions (such as weather, lighting, road type, etc.) captured in the context data 214.

[0174] The assessor node (is_latd_safe), which is the parent of the latd and latsd extractor nodes, is mapped to the Gt atomic predicate. Thus, when the rule tree 408 is executed, the is_latd_safe assessor node applies the Gt function to the outputs of the latd and latsd extractor nodes to calculate a true / false result for each time step of the scenario, returning true for each time step when the latd signal exceeds the latsd signal and false otherwise. In this way, the "safe lateral distance" rule is constructed from atomic extractor functions and predicates, and the ego agent fails the safe lateral distance rule if the lateral distance reaches or falls below the safe lateral distance threshold. As can be understood, this is a very simple example of a rule tree. Rules of any complexity can be constructed following the same principle.

[0175] The test oracle 252 applies the rule tree 408 to the Scenario Ground Truth 310 and provides the results via the user interface (UI) 418.

[0176] Figure 8B shows an example of a rule tree that includes a lateral distance branch corresponding to Figure 8A. Additionally, the rule tree includes a longitudinal distance branch and a top-level OR predicate (safety distance node, is_d_safe) for implementing a safety distance metric. Similar to the lateral distance branch, the longitudinal distance branch extracts a longitudinal distance and a longitudinal distance threshold signal (extractor nodes lond and lonsd, respectively) from the scenario data, and if the longitudinal distance exceeds the safe longitudinal distance threshold, the longitudinal safety assessor node (is_lond_safe) returns true. The top-level OR node returns true if one or both of the lateral distance and the longitudinal distance are safe (below the corresponding threshold), and returns false if neither is safe. In this context, it is sufficient for only one of the distances to exceed the safety threshold (for example, if two vehicles are traveling in adjacent lanes, when they are adjacent, the longitudinal separation is zero or near zero, but if those vehicles have sufficient lateral separation, the situation is not dangerous).

[0177] The numerical output of the top-level node can be, for example, a time-varying robustness score.

[0178] Different rule trees can be constructed to implement, for example, different rules of a given safety model, different safety models, or selectively apply rules to different scenarios (in a given safety model, not all rules necessarily apply to all scenarios, and in this approach, different rules or combinations of rules can be applied to different scenarios). Within this framework, rules can also be constructed to evaluate comfort (for example, based on instantaneous acceleration and / or jerk along a trajectory), progress (for example, based on the time taken to reach a defined goal), etc.

[0179] The above example considers simple logical predicates that are evaluated with a result or signal at a single point in time, such as OR, AND, Gt, etc. However, in practice, it may be desirable to formulate specific rules from the perspective of temporal logic.

[0180] Hekmatnejad et al., "Encoding and Monitoring Responsibility Sensitive Safety Rules for Automated Vehicles in Signal Temporal Logic" (2019), MEMOCODE ’19: Proceedings of the 17th ACM-IEEE International Conference on Formal Methods and Models for System Design (which is incorporated herein by reference in its entirety) discloses a signal temporal logic (STL) encoding of RSS safety rules. Temporal logic provides a formal framework for constructing conditional predicates over time. This means that the result computed by an assessor at a given instant can depend on the results and / or signal values at other instants.

[0181] For example, a requirement of a safety model may be that an ego agent responds to a particular event within a set time frame. Such rules can be encoded in a similar way using temporal logic predicates within a rule tree.

[0182] In the above example, the performance of stack 100 is evaluated at each time step of the scenario. From this, the overall test result (e.g., pass / fail) can be derived. For example, a specific rule (e.g., a rule where safety is critically important) may result in an overall fail if the rule fails at any time step within the scenario (i.e., to obtain an overall pass for the scenario, the rule must pass at all time steps). For other types of rules, the overall pass / fail criteria may be “looser” (e.g., for a specific rule, a fail may be triggered only if the rule fails for a certain number of consecutive time steps), and such criteria may be context-dependent.

[0183] Figure 8C schematically shows the hierarchy of rule evaluation implemented within test oracle 252. A set of rules 254 is received for implementation in test oracle 252.

[0184] A specific rule is applied only to the self-agent (an example is a comfort rule that assesses whether the maximum acceleration or jerk threshold is exceeded by the self-trajectory at any given instant).

[0185] Other rules relate to the interaction between the self-agent and other agents (e.g., a “no collision” rule or the safety distance rule discussed above). Each such rule is evaluated in a pairwise manner between the self-agent and each other agent. As another example, a “pedestrian emergency brake” rule may be activated only when a pedestrian walks in front of the host vehicle and only with respect to that pedestrian agent.

[0186] Not all rules necessarily apply to all scenarios, and some rules may only apply to part of a scenario. The rule activation logic 422 within the test oracle 252 determines whether each of the rules 254 applies to the scenario in question, when it applies, and, if so, selectively activates the rule when it applies. Thus, a rule may remain active throughout a scenario, may never be activated in a given scenario, or may only be activated in part of a scenario. Further, a rule may be evaluated against a different number of agents at different times in a scenario. Selectively activating rules in this way can significantly improve the efficiency of the test oracle 252.

[0187] The activation or deactivation of a given rule may depend on the activation / deactivation of one or more other rules. For example, the "optimal comfort" rule may be considered non - applicable when the pedestrian emergency brake rule is activated (since pedestrian safety is the top concern), and may always be deactivated when the latter is active.

[0188] The rule evaluation logic 424 evaluates each active rule for the duration that it remains active. Each interacting rule is evaluated in a pairwise fashion between the self - agent and the other agents to which it applies.

[0189] Also, there may be some degree of interdependence in the application of rules. For example, another way to handle the relationship between the comfort rule and the emergency brake rule would be to increase the jerk / acceleration threshold of the comfort rule whenever the emergency brake rule is activated for at least one other agent.

[0190] A pass / fail result is contemplated, but the rules may be non-binary. For example, two categories of fail, namely, "acceptable" and "unacceptable" may be introduced. Again, considering the relationship between the comfort rule and the emergency brake rule, an acceptable fail of the comfort rule may occur when it was a fail for that rule but the emergency brake rule was active. Thus, the interdependencies between the rules can be addressed in various ways.

[0191] The activation criteria for Rule 254 can be specified in the rule creation code provided to the rule editor 400, and the same is true for the nature of the rule interdependencies and the mechanisms for implementing those interdependencies.

[0192] Graphical User Interface: FIG. 9A shows a schematic block diagram of the visualization component 520. The visualization component is shown to have an input connected to the test database 258 for rendering the output 256 of the test oracle 252 on the graphical user interface (GUI) 500. The GUI is rendered on the display system 522.

[0193] FIG. 9B shows an exemplary view of the GUI 500. This view pertains to a particular scenario that includes a plurality of agents. In this example, the test oracle output 526 pertains to a plurality of external agents, and the results are organized per agent. For each agent, a time series of the results is available for each rule applicable to that agent at a point in the scenario. In the illustrated example, the summary view of "Agent 01" is selected, and the "top" results calculated for each applicable rule are displayed. There are top results calculated at the root nodes of each rule tree. Color coding is used to distinguish the periods when the rule was inactive for that agent, the periods when it was active and passed, and the periods when it was active and failed.

[0194] A first selectable element 534a is provided for each time series of results. This enables access to the results of lower levels of the rule tree, i.e., the results calculated below the rule tree.

[0195] Figure 9C shows a first expanded view of the results of "Rule 02", and the results of the lower-level nodes are also visualized. For example, regarding the "safe distance" rule in Figure 4B, the results of the "is_latd_safe" node and the "is_lond_safe" node may be visualized (labeled "C1" and "C2" in Figure 9C). In the first expanded view of Rule 02, the achievement / failure of Rule 02 is defined by the logical OR relationship between results C1 and C2, and it can be seen that Rule 02 is only failed when both C1 and C2 result in failure (similar to the case of the "safe distance" rule above).

[0196] A second selectable element 534b is provided for each time series of results, which enables access to the associated numerical performance score.

[0197] Figure 9D shows a second expanded view, where the results of Rule 02 and the results of "C1" are expanded, and the associated scores for the periods during which these rules are active for Agent 01 are visible. The scores are displayed as a visually scored - time plot, also color - coded to represent pass / fail.

[0198] Exemplary scenario: Figure 10A shows a first instance of an interrupt scenario in simulator 202 that ends in a collision event between host vehicle 602 and another vehicle 604. The interrupt scenario is characterized as a multi-lane driving scenario, where host vehicle 602 is moving along a first lane 612 (host lane), and another vehicle 604 is initially moving along a second adjacent lane 604. At some point in this scenario, another vehicle 604 moves from the adjacent lane 614 into the host lane 612, in front of (interrupt distance) host vehicle 602. In this scenario, host vehicle 602 cannot avoid a collision with another vehicle 604. The first scenario instance ends in response to the collision event.

[0199] Figure 10B shows an example of a first oracle output 256a obtained from the ground truth 310a of the first scenario instance. The "no collision" rule is evaluated over the duration of the scenario between host vehicle 602 and another vehicle 604. The collision event results in a failure of this rule at the end of the scenario. Additionally, the "safe distance" rule of Figure 4B is evaluated. When another vehicle 604 approaches host vehicle 602 laterally, a point in time (t1) is reached when both the safe lateral distance threshold and the safe longitudinal distance threshold are violated, which results in a failure of the safe distance rule that persists until the collision event at time t2.

[0200] Figure 10C shows a second instance of the interrupt scenario. In the second instance, the interrupt event does not result in a collision, and host vehicle 602 can reach a safe distance behind another vehicle 604 after the interrupt event.

[0201] Figure 10D shows an example of a second oracle output 256b obtained from the ground truth 310b of a second scenario instance. In this case, it passes the "no collision" rule throughout. The safety distance rule is violated at time t3 when the lateral distance between the host vehicle 602 and the other vehicle 604 becomes unsafe. However, at time t4, the host vehicle 602 somehow reaches a safe distance behind the other vehicle 604. Therefore, the safety distance rule is only non-compliant between time t3 and time t4.

[0202] Recognition error framework As described above, both recognition errors and driving rules can be evaluated based on the extracted pseudo ground truth 144 determined by the ground truth plumbing pipeline 144 and presented to the GUI 500.

[0203] Figure 11 shows an architecture for evaluating recognition errors. A triage tool 152 with a recognition oracle 1108 is used to extract and evaluate recognition errors for both real driving scenarios and simulated driving scenarios, and outputs results that are rendered to the GUI 500 alongside the results from the test oracle 252. Note that the triage tool 152, which is referred to herein as the recognition triage tool, may be more generally used to extract driving data including recognition data and driving performance data useful for testing and improving the autonomous vehicle stack and present it to the user.

[0204] For the real sensor data 140 from a driving run, the output of the online recognition stack 102 is passed to the triage tool 152, and a numerical "real world" recognition error 1102 is determined based on the extracted ground truth 144 obtained by executing both the real sensor data 140 and the online recognition output via the ground truth plumbing pipeline 400.

[0205] Similarly, in the case of a simulated driving run where sensor data is simulated from zero and a recognition stack is applied to the simulated sensor data, the triage tool 152 calculates a simulated recognition error 1104 based on a comparison between the detection result from the recognition stack and the simulation ground truth. However, in the case of simulation, the ground truth can be directly obtained from the simulator 202.

[0206] If the simulator 202 directly models the recognition error to simulate the output of the recognition stack, the difference between the simulated detection result and the simulation ground truth, i.e., the simulated recognition error 1110, is known and is directly passed to the recognition oracle 1108.

[0207] The recognition oracle 1108 receives a set of recognition rule definitions 1106, which may be defined via a user interface or described in a domain-specific language, as will be explained in more detail later. The recognition rule definitions 1106 may apply thresholds or rules that define recognition errors and their limits. The recognition oracle applies the defined rules to the actual or simulated recognition errors obtained for the driving scenario and identifies where the recognition errors break the defined rules. These results are passed to the rendering component 1120, which renders visual indicators of the evaluated recognition rules for display on the graphical user interface 500. For clarity, the input to the test oracle is not shown in FIG. 11, but note that the test oracle 252 also depends on the ground truth scenario obtained from either the ground truth plumbing pipeline 400 or the simulator 202.

[0208] Next, further details of the framework for evaluating the recognition errors of the real-world driving stack against the extracted ground truth are described. As described above, both the recognition errors and the driving rule analysis by the test oracle 252 can be incorporated into the real-world driving analysis tool, which will be described in more detail below.

[0209] Not all errors have the same importance. For example, a 10 cm translational error in an agent 10 meters away from the ego vehicle is much more important than the same translational error in an agent 100 meters away. A simple solution to this problem would be to scale the error based on the distance from the ego vehicle. However, the relative importance of different recognition errors, or the sensitivity of the ego vehicle's driving performance to different errors, depends on the use case of the given stack. For example, when designing a cruise control system for driving on a straight road, it should be sensitive to translational errors but not particularly sensitive to orientation errors. However, an AV dealing with the entrance to a roundabout uses the detected orientation of the agent as an indication of whether the agent is about to exit the roundabout and thus whether it can safely enter the roundabout, and should therefore be very sensitive to orientation errors. Therefore, it is desirable to make it possible to set the sensitivity of the system to different recognition errors according to each use case.

[0210] To define recognition errors, a domain-specific language is used. Using this, recognition rule 1402 (see FIG. 14) can be created, for example, by defining the tolerance for translational errors. This rule implements a settable set of safe error levels for different distances from the host vehicle. This is defined within table 1400. For example, if the vehicle is less than 10 meters away, the error in its position (i.e., the distance between the detected result of the vehicle and the refined pseudo ground truth detection result) can be defined to be 10 cm or less. If the agent is 100 meters away, the allowable error may be defined as up to 50 cm. Using a lookup table, the rules can be defined for any given use case. Based on these principles, more complex rules can be constructed. Rules may be defined such that the errors of other agents are completely ignored based on their positions relative to the host vehicle, for example, an agent in the oncoming lane when the host lane is separated from oncoming traffic by a median. Traffic behind the host vehicle beyond a defined cut-off distance may also be ignored based on the rule definition.

[0211] By defining a recognition error specification 1600 that includes all the rules to be applied, a set of rules can be collectively applied to a given driving scenario. Typical recognition rules that can be included in the specification 1600 measure translational errors in the longitudinal and lateral directions (the average error of the detection results with respect to the ground truth in the longitudinal and lateral directions respectively), orientation errors (defining the minimum angle by which the detection results need to be rotated to match the corresponding ground truth), size errors (the error in each dimension of the detected bounding box, or the ratio of the intersection / union of the aligned ground truth and the detected box to obtain the volume difference), and define thresholds for these. Further rules may be based on the dynamics of the vehicle, which include errors in the speed and acceleration of the agent, and classification errors, such as penalty values in case of misclassifying a car as a pedestrian or a truck. The rules may also include false detections or missed detections, as well as detection delays.

[0212] Based on the defined recognition rules, a robustness score can be constructed. Effectively using this, when the detection results are within the specified thresholds of the rules, the system should be able to drive safely, and if not (e.g., if there is too much noise), it can be said that something bad might happen that the host vehicle may not be able to handle, which needs to be formally captured. For example, complex combinations of rules can be included to evaluate the detection results over time and to incorporate complex weather dependencies.

[0213] Using these rules, errors can be associated with the reproduction of scenarios in the UI. As shown in Figure 14, different recognition rules are displayed in different colors corresponding to different results when applying a given rule definition in DSL to the timeline of that rule. This is the main use case of DSL (i.e., visualization for the triage tool). The user describes the rules in DSL, and the rules are displayed on the timeline of the UI.

[0214] DSL can also be used to define the contract between the recognition stack and the planning stack of the system based on the robustness score calculated for the defined rules. Figure 15 shows an exemplary graph of the robustness score for a given error definition such as translational error. If the robustness score exceeds the defined threshold of 1500, this indicates that the recognition error is within the expected performance and the overall system should promise safe operation. As shown in Figure 15, if the robustness score drops below the threshold, at that level of recognition error, the planner 106 cannot be expected to drive safely, so the error is "out of contract". This contract is essentially the requirement specification for the recognition system. This can be used to assign responsibility to either recognition or planning. If an error is identified as being within the contract when the vehicle is behaving improperly, this points to an issue with the planner rather than a recognition problem, and conversely, for bad behavior when recognition is out of contract, it is due to a recognition error.

[0215] By annotating whether a recognition error is considered within the contract or out of the contract, the contract information can be displayed in the UI500. This uses a mechanism that obtains the contract specification from DSL and automatically flags out-of-contract errors at the front end.

[0216] Figure 16 shows a third use case that integrates recognition errors across different modalities (i.e., the real world and simulation). The above description is related to real-world driving where a real vehicle is driven to collect data, and offline, the refinement technique and the triage tool 152 calculate the recognition errors and calculate whether these errors are within or outside the contract. However, the same recognition error specification 1600 that specifies the recognition error rules for evaluating the errors can be applied to the simulated driving runs. The simulation can generate simulated sensor data processed by the recognition stack as described above with reference to Figure 11, or can be by directly simulating the detection results from the ground truth using the recognition error model.

[0217] In the first case, the detection results based on the simulated sensor data 1112 have errors 1104, and a DSL can be used to define whether these errors are within or outside the contract. This can also be done by simulation based on the recognition error model 208 (i.e., adding noise to the object list), and it is possible to calculate and verify the injected errors 1110 to check whether the simulator 202 is modeling what is expected to be modeled. Using this, instead of injecting out-of-contract errors, it is also possible to avoid the stack failing purely due to recognition errors by intentionally injecting in-contract errors. In one use case, in-contract but near the edge of the contract errors may be injected in the simulation, and it can be verified that the planning system operates correctly when the expected recognition performance is given. This decouples the development of recognition and planning because they can be tested individually in the context of this contract, and these systems should cooperate to a satisfactory level if the recognition meets the contract and the planner functions within the limits of the contract.

[0218] Depending on where the recognition model is sliced, for example when performing fusion, there may be little knowledge about what comes out of the simulator, so evaluating it in terms of in - contract errors and out - of - contract errors is useful for analyzing the simulated scenario.

[0219] Another use of the DSL is to appraise the accuracy of the pseudo ground truth 144 itself. It is impossible to purify imperfect detection results to obtain a perfect ground truth, but there is probably an acceptable accuracy that needs to be reached for the purification pipeline to be used reliably. Using DSL rules, it is possible to appraise the current pseudo ground truth and determine how close it is to the current “true” GT and how much closer it needs to be in the future. This could be the same contract used to check the online recognition errors calculated against the pseudo ground truth, but with stricter limits applied regarding accuracy, so that there is sufficient confidence that the pseudo ground truth is “correct” enough for the online detection results to be appraised. The acceptable accuracy of the pseudo ground truth can be defined as an in - contract error when measured against the “true” ground truth. Within certain thresholds, it is acceptable to generate some errors even after purification. If different systems have different use cases, each system applies a different set of DSL rules.

[0220] The “true” ground truth against which the purified detection results are appraised is obtained by selecting a real - world dataset, manually annotating it, evaluating the pseudo GT against this manual GT according to the defined DSL rules, and determining whether the acceptable accuracy has been achieved. Each time the purification pipeline is updated, the accuracy appraisal of the purified detection results can be re - executed to check whether the pipeline has regressed.

[0221] Another use of the DSL is that, when a contract is defined between the recognition 102 and the plan 106, it becomes possible to divide the types of tests that need to be performed in the recognition layer. This is shown in Figure 17. For example, the recognition layer can be supplied with a set of sensor readings that must contain all the errors that must be within the contract, and DSL rules can be applied to check whether this is the case. Similarly for the planning layer, first a ground-truth test 1702 can be applied, and if it passes, then a contract test 1704 is applied, so that the system is supplied with a list of objects with in-contract errors and checks whether the planner behaves safely.

[0222] In one exemplary test scheme, the planner may be considered as "given", use simulation to generate recognition errors, and search for the limits of acceptable recognition accuracy for the planner to operate as intended. These limits can then be used to semi-automatically create a contract for the recognition system. The set of recognition systems can be tested against this contract to find those that meet it, or the contract can be used as a guide when developing the recognition system.

[0223] Real-world driving analysis tool The above test framework, namely the test oracle 252 and the recognition triage tool 152, may be combined in a real-world driving analysis tool, in which both recognition evaluation and driving evaluation are applied to the recognition ground-truth extracted from the ground-truth pipeline 400 as shown in Figure 2C.

[0224] FIG. 12A shows an exemplary user interface for analyzing a driving scenario extracted from real-world data. In the example of FIG. 12A, a schematic overhead representation 1204 of the scene is shown based on point cloud data (e.g., derived from lidar, radar, or stereo or monocular depth imaging), and a corresponding camera frame 1224 is shown in an inset. Road layout information may be obtained from high-definition map data. The camera frame 1224 may be annotated with detection results. The UI may also show sensor data collected during driving, such as lidar, radar, or camera data. This is shown in FIG. 12B. The scene visualization 1204 is also overlaid with annotations based on the derived pseudo ground truth as well as detection results from in-vehicle recognition components. In the illustrated example, there are three vehicles, each annotated with a box. The solid box 1220 shows the pseudo ground truth of the agents in the scene, and the contour 1222 shows the unrefined detection results from the host vehicle's recognition stack 102. A visualization menu 1218 is shown, in which the user can select which of the sensor data, online and offline detection results to display. These may be switched on and off as needed. Showing the real sensor data alongside both the vehicle detection results and the ground truth detection results enables the user to identify or confirm specific errors in the vehicle detection results. The UI 500 enables the playback of the selected video, a timeline view is shown, and the user can select any point in time 1216 within the video and display the bird's-eye view snapshot and the camera frame corresponding to the selected point in time.

[0225] As described above, the recognition stack 102 can be evaluated by comparing the detection results with the refined pseudo ground truth 144. The recognition is evaluated in light of defined recognition rules 1106 that can depend on the use case of a particular AV stack. These rules specify different ranges of values for the discrepancies between the position, orientation, or scale of the vehicle detection results and those of the pseudo ground truth detection results. The rules can be defined in a domain-specific language (described above with reference to FIG. 14). As shown in FIG. 12A, different recognition rule results are shown along the "top" recognition timeline 1206 of the driving scenario, which aggregates the results of the recognition rules, and a period on the timeline is flagged when any of the recognition rules are violated. This can be expanded to present a set of individual recognition rule timelines 1210 for each of the defined rules.

[0226] The recognition error timeline may be "zoomed out" to present a longer period of the driving run. In the zoomed-out view, it may not be possible to display the recognition errors with the same granularity as when zoomed in. In this case, the timeline may display an aggregation of the recognition errors over a time window to provide a set of recognition errors grouped for the zoomed-out view.

[0227] The second driving assessment timeline 1208 shows how the pseudo ground truth data was assessed against the driving rules. The aggregated driving rules are shown on the top-level timeline 1208, which can be expanded into a set of individual timelines 1212 that display the performance against each defined driving rule. Each rule timeline can be further expanded, as shown, to display a plot 1228 of the numerical performance scores over time for a given rule. This corresponds to the selectable element 534b described above with reference to FIG. 9C. In this case, the pseudo ground truth detection results are considered the actual driving behavior of the agents in the scene. To check whether the vehicle behaved safely for a given scenario, the behavior of the host vehicle can be evaluated against the defined driving rules, for example, based on a digital highway code.

[0228] In summary, both the recognition rule evaluation and the driving assessment are based on using the above-described offline recognition method to refine the detection results from real-world driving. In the case of driving assessment, the refined pseudo ground truth 144 is used to assess the behavior of the host vehicle against the driving rules. As shown in FIG. 2C, this can also be used to generate a simulated scenario for testing. In the case of recognition rule evaluation, the recognition triage tool 152 compares the recorded vehicle detection results with the offline refined detection results to quickly identify and triage potential recognition failures.

[0229] Also, drive notes may be displayed in the drive note timeline view 1214, which may include notable events flagged during the drive. For example, drive notes include when the vehicle braked or changed direction, or when the human driver disengaged the AV stack.

[0230] An additional timeline may be displayed presenting user-defined metrics that help the user debug and triage potential issues. User-defined metrics may be defined to identify errors or stack defects and to triage errors when they occur. The user may define custom metrics according to the goals of a given AV stack. Exemplary user-defined metrics may flag when messages do not arrive in order or message delays in recognition messages. This is useful for triage as it may be used to determine whether the plan was made due to a planner mistake or because the message arrived late or out of order.

[0231] Figure 12B shows an example of a UI visualization 1204 where sensor data is displayed and camera frame 1224 is shown in an inset. Typically, sensor data from a single temporal snapshot is presented. However, each frame may present sensor data aggregated over multiple time steps to obtain a static scene map when high-resolution map data is not available. As shown on the left, there are several visualization options 1218 for showing or hiding data such as camera, radar or lidar data collected during a real-world scenario, or online detection results from the recognition of the host vehicle itself. In this example, the online detection results from the vehicle are shown as a contour 1222 overlaid on a solid box 1220 representing the ground truth refined detection results. An orientation error can be seen between the ground truth and the vehicle's detection results.

[0232] The purification process implemented by the ground truth pipeline 400 is used to generate a pseudo ground truth 144 as the basis for multiple tools. The presented UI displays the results from the recognition triage tool 152, which enables using the test oracle 252 to assess the driving capabilities of the ADAS in a single run case, detect defects, extract scenarios to reproduce issues (see Figure 2C), and send the identified issues to the developers to improve the stack.

[0233] Figure 12C shows an exemplary user interface configured to enable the user to zoom in on a subsection of a scenario. Figure 12C shows a snapshot of the scenario, including the schematic representation 1204 and the camera frame 1224 presented in the inset, as described above with respect to Figure 12A. Also shown in Figure 12C are the set of recognition error timelines 1206, 1210, as well as the expandable driving assessment timeline 1208 and the drive note timeline 1214 described above.

[0234] In the example shown in FIG. 12C, the current snapshot of the driving scenario is shown by a scrubber bar 1230 that spans all the timeline views simultaneously. This may be used instead of the indication 1216 of the current point within the scenario on a single playback bar. The user can click on the scrubber bar 1230 to select the bar and move it to any point in the driving scenario. For example, the user may be interested in a particular error, such as a point within a section that is colored red or otherwise indicated as a section containing an error on the position error timeline, and the indication is determined based on the "ground truth" and the position error observed at that point between the detection results during the period corresponding to the indicated section. The user can click on the scrubber bar and drag the bar to the point of interest within the position error timeline. Alternatively, the user can click on a point on any of the timelines that the scrubber spans to place the scrubber at that point. This updates the overview view 1204 and the inset view 1224 to present the top-down overview view and the camera frame corresponding to the selected point, respectively. The user can then inspect the overview view and the available camera data or other sensor data to confirm the position error and identify possible reasons for the recognition error.

[0235] The "ruler" bar 1232 is presented above the recognition timeline 1206 and below the overview view. This includes a series of "notches" that indicate the time intervals of the driving scenario. For example, if a 10-second time interval is displayed in the timeline view, notches indicating 1-second intervals are presented. At some points, numerical indicators such as "0 seconds", "10 seconds", etc. are also labeled.

[0236] A zoom slider 1234 is provided at the bottom of the user interface. The user can drag the indicator along the zoom slider to change the portion of the driving scenario presented on the timeline. Alternatively, the position of the indicator may be adjusted by clicking on a desired point on the slider bar to which the indicator is to be moved. A percentage indicating the currently selected zoom level is presented. For example, if the length of the entire driving scenario is one minute, timelines 1206, 1208, 1214 present recognition errors, driving appraisals, and drive notes respectively over the one-minute drive, the zoom slider indicates 100%, and the button is at the leftmost position. When the user slides the button until the zoom slider indicates 200%, the timeline is adjusted to present only the results corresponding to a 30-second snippet of the scenario.

[0237] The zoom may be configured to adjust the displayed portion of the timeline according to the position of the scrubber bar. For example, if the zoom is set to 200% in a 1-minute scenario, the zoomed-in timeline presents a 30-second snippet centered on the selected point where the scrubber is located, that is, a 15-second timeline is presented before and after the point indicated by the scrubber. Alternatively, the zoom may be applied based on a reference point such as the start point of the scenario. In this case, the zoomed-in snippet presented on the timeline after zooming always starts from the start point of the scenario. The granularity of the notches and numerical labels of the ruler bar 1232 may be adjusted according to the degree to which the timeline is zoomed in or out. For example, if the scenario is zoomed in from 30 seconds to present a 3-second snippet, before zooming, the numerical labels may be displayed at 10-second intervals and the notches at 1-second intervals, and after zooming, the numerical labels may be displayed at 1-second intervals and the notches at 100ms intervals. The visualization of the time steps of the timelines 1206, 1208, 1214 is "stretched" to correspond to the zoomed-in snippet. A higher level of detail may be displayed by the timeline in the zoomed-in view because a shorter time snippet can be represented by a larger area in the display of the timeline within the UI. Therefore, an error over a very short time within a longer scenario may be visible in the timeline view only when zoomed in.

[0238] Other zoom inputs may be used to adjust the timeline to display shorter or longer snippets of the scenario. For example, if the user interface is implemented on a touch screen device, the user may apply a pinch gesture to apply zoom to the timeline. In other examples, the user may scroll the mouse scroll wheel forward and backward to change the zoom level.

[0239] When the timeline is zoomed in to show only a subset of the driving scenarios, the timeline can be scrolled temporally to shift the display portion temporally, so that various parts of the scenario can be inspected in the timeline view by the user. The user can scroll by clicking and dragging a scroll bar (not shown) at the bottom of the timeline view or by using, for example, the touch pad of a related device on which the UI is operating.

[0240] The user can also select snippets of scenarios, for example, for further analysis or as a basis for simulations, which are exported. FIG. 12D shows how sections of the driving scenario can be selected by the user. The user can click on the relevant points on the ruler bar 1232 with the cursor. This can be done at any zoom level. This sets the first boundary of the user selection range. The user drags the cursor along the timeline to expand the selection range up to the selected point in time. When zoomed in, by continuing to drag to the end of the displayed snippet of the scenario, this scrolls the timeline forward and allows the selection range to be further expanded. The user can stop dragging at any point, and the point where the user stops becomes the end boundary of the user selection range. The bar 1230 at the bottom of the user interface displays the temporal length of the selected snippet, and this value is updated as the user drags the cursor to expand or contract the selection range. The selected snippet 1238 is presented as a shaded section on the ruler bar. Some buttons 1236 are presented that provide user actions such as "Extract Trajectory Scenario" to extract the data corresponding to the selection range. This may be stored in a database of the extracted scenarios. This may be used for further analysis or as a basis for simulating similar scenarios. After making a selection, the user can zoom in or out, and the selection range 1238 on the ruler bar 1232 also expands and contracts along with the ruler and the timelines of recognition, driving assessment, and drive notes.

[0241] The pseudo ground truth data can also be used with a data search tool to search for data within a database. This tool can be used when a new version of the AV stack is deployed. In the case of new software versions, the vehicle can be driven for a certain period (e.g., one week) to collect data. Within this data, the user may be interested in testing how the vehicle behaves under certain conditions, so queries such as "Show night driving" or "Show when it was raining" may be provided. The data search tool can retrieve the relevant videos and then use a triage tool to investigate. This data search tool functions as a kind of entry point for further analysis.

[0242] For example, when a new software version is implemented and the AV is driven for a while to collect a certain amount of data, a further assessment tool may be used to aggregate the data and understand the overall performance of the vehicle. This vehicle has a set of newly developed functions such as the use of indicators and entering and exiting roundabouts, and an overall performance evaluation of how well the vehicle behaves with respect to these functions may be requested.

[0243] Finally, a regression issue can be checked by using a re-simulation tool to run the sensor data on the new stack and perform an open-loop simulation.

[0244] FIG. 13 shows an exemplary user interface of the recognition triage tool 152 focused on scenario visualization 1204 and recognition error time lines 1206, 1210. As shown on the left side, there are several visualization options 1218 for displaying or hiding data such as camera, radar or lidar data collected in a real scenario, or online detection results from the recognition of the host vehicle itself. In this case, the visualization is limited to only the refined detection results, i.e., only the agents detected offline and the refined results are shown by solid boxes. Each solid box has associated online detection results (not shown) indicating how the agent was recognized prior to error correction / refinement at that temporal snapshot. As described above, there is a certain amount of error between the ground truth 144 and the original detection results. Various errors can be defined, including scale, position, and orientation errors of agents in the scene, as well as "ghost" detections and missed detections due to false detections.

[0245] As described above, not all errors have the same importance. The DSL for recognition rules enables the definition of rules according to the required use cases. For example, when designing a cruise control system for driving on a straight road, this should be sensitive to translational errors but not particularly sensitive to orientation errors. However, an AV dealing with the entrance to a roundabout uses the detected orientation of an agent as an indication of whether the agent is about to exit the roundabout and thus whether it can safely enter the roundabout, and should therefore be very sensitive to orientation errors. The recognition error framework enables the definition of separate tables and rules that indicate the relative importance of a given translational or orientation error in its use case. The boxes shown around the host vehicle in FIG. 13 indicate areas of interest that may be defined for the recognition rules to target, for illustrative purposes. The rule evaluation results may be displayed within the recognition error timeline 1210 of the user interface. For example, a visual indicator of the rule may be displayed within the schematic representation 1204 by flagging the area where a particular rule is defined, which is not shown in FIG. 13.

[0246] It is also possible to not only display the results of a single snapshot of a driving run, but also apply queries and filtering to provide the user who filters and analyzes the data according to the recognition evaluation results with more context.

[0247] Figures 18A and 18B show an example of a graphical user interface 500 for filtering and displaying the recognition results of actual driving runs. For a given run, as described above, a recognition error timeline 1206 with a rule evaluation aggregated for all recognition errors is displayed. A second set of timelines 1226 may be presented that show characteristics of the driving scene, such as weather conditions, road features, other vehicles, and vulnerable road users. These may be defined within the same framework used to define the recognition error rules. Note that the recognition rules may be defined such that different thresholds are applied depending on different driving conditions. Figure 18A also shows a filtering function 1800 where the user can select a query to apply to the evaluation. In this example, the user query is to look for "slices" of driving runs where there are vulnerable road users (VRUs).

[0248] This query is processed and used to filter the frame of the driving scenario representation into a frame tagged with vulnerable road users. Figure 18B shows an updated view of the recognition timeline after the filter has been applied. As shown in the figure, a subset of the original timeline is presented, and in this subset, vulnerable road users are always present, as shown in the "VRU" timeline.

[0249] Figure 19A shows other functions that may be used to perform an analysis within the graphical user interface 500. A set 1900 of error threshold sliders that can be adjusted by the user is shown. The error range may be informed by the recognition error limits defined in the DSL of the recognition rules. The user may adjust a given error threshold by sliding a marker to the desired new threshold for that error. For example, the user may set a fail threshold for a translational error of 31 m. This threshold can then be fed back into the translational error defined within the recognition rule specification described in the aforementioned recognition rule DSL to adjust the rule definition to take into account the new threshold. The new rule evaluation result is passed to the front end and the rule fails currently occurring with respect to the new threshold is shown in the expanded timeline view 1210 of the given error. As shown in Figure 19A, lowering the threshold for unacceptable error values causes more errors to be flagged in the timeline.

[0250] Figure 19B shows a method that enables a user to select and examine the most relevant frame based on the calculated recognition errors by applying intensive analysis to a selected slice of the driving scenario. As described above, the user can use the filtering function 1800 to filter the scenario to present only the frames where vulnerable road users exist. Within the matching frames, the user can use the selection tool 1902 to further "slice" the scenario into specific snippets, and the selection tool 1902 can be dragged along the timeline 1206 and expanded to cover the period of interest. For the selected snippet, aggregated data may be presented to the user within the display 1904. Various attributes of the recognition errors captured within the selected snippet may be selected and graphed against each other. In the illustrated example, the type of error and the magnitude of the error are graphed, enabling the user to visualize the most significant error of each type for the selected part of the scenario. The user may select any point on the graph to display the corresponding camera image 1906 of the frame where the error occurred, along with other variables of the scene such as occlusion, and the user can examine the frame for factors that may have caused the error.

[0251] The ground truth pipeline 400 may be used with the recognition triage tool 152 and the test oracle 252, as well as additional tools for querying, aggregating, and analyzing the performance of the vehicle, including the data search and aggregation assessment tools described above. The graphical user interface 500 may display the results from these tools in addition to the snapshot views described above.

[0252] Although the above example considers testing an AV stack, the present technology can be applied to test components of other forms of mobile robots. For example, other mobile robots for carrying goods in industrial areas inside and outside have been developed. Such mobile robots do not carry people and belong to a class of mobile robots called UAVs (unmanned autonomous vehicles). Autonomous aerial mobile robots (drones) have also been developed.

[0253] References to components, functions, modules, etc. in this specification refer to the functional components of a computer system that can be implemented at the hardware level in various ways. The computer system may comprise execution hardware configured to execute the method / algorithm steps disclosed herein and / or to implement a model trained using this technology. The term execution hardware encompasses any form / combination of hardware configured to execute the relevant method / algorithm steps. The execution hardware may take the form of one or more processors, which may be programmable or non-programmable, or a combination of programmable hardware and non-programmable hardware may be used. Examples of suitable programmable processors include general-purpose processors based on instruction set architectures such as CPUs, GPU / accelerator processors, etc. Such general-purpose processors typically execute computer-readable instructions held in memory coupled to or incorporated within the processor and perform the relevant steps in accordance with those instructions. Other forms of programmable processors include field-programmable gate arrays (FPGAs) having programmable circuit configurations through circuit description code. Examples of non-programmable processors include application-specific integrated circuits (ASICs). The code, instructions, etc. may be stored, if necessary, on a temporary or non-temporary medium (examples of the latter include solid-state, magnetic, and optical storage devices, etc.). The subsystems 102-108 of the runtime stack in Figure 2A may be implemented in a vehicle using a programmable processor, a dedicated processor, or a combination of both, or in a non-vehicle computer system in the context of testing, etc. Similarly, various components in the figures, including Figures 11 and 6, such as the simulator 202 and the test oracle 252, may also be implemented using programmable hardware and / or dedicated hardware.

Claims

Claim 1 A computer-implemented method for assessing the performance of an autonomous vehicle, comprising: receiving, at an input, performance data from at least one autonomous driving run, the performance data comprising a time series of at least one recognition error and a time series of at least one driving performance result; generating, at a rendering component, rendering data for rendering a graphical user interface, the graphical user interface for visualizing the performance data; (i) a recognition error timeline, and (ii) a driving assessment timeline, wherein the timelines are temporally aligned and divided into a plurality of time steps of the at least one driving run, and for each time step, the recognition error timeline comprises a visual indication of whether a recognition error occurred in the time step, and the driving assessment timeline comprises a visual indication of the driving performance in the time step. Claim 2 The method of claim 1, wherein the recognition error timeline and the driving assessment timeline are parallel to each other. Claim 3 The method of claim 1, wherein the driving performance is assessed with respect to one or more predefined driving rules. Claim 4 The method of claim 3, wherein the driving assessment timeline aggregates driving performance across a plurality of individual driving rules, and the driving assessment timeline is expandable to view the respective driving assessment timelines of the individual driving rules. Claim 5 The method of claim 3, wherein the driving assessment timeline or each driving assessment timeline is expandable to view a computational graph representation of the driving rules. Claim 6 The method of claim 3, wherein the driving run is a real-world run and driving rules are applied to a real-world trajectory. Claim 7 The method of claim 1, wherein a ground truth pipeline is used to extract a ground truth recognition output, the ground truth recognition output being used to determine recognition errors and assess driving performance. Claim 8 The method according to claim 7, wherein the ground truth pipeline is automated.

9. The method according to claim 1, wherein at least some of the recognition errors are identified without using the ground truth recognition output.

10. The recognition error is a flickering detection result, or The method according to claim 9, comprising at least one of a jumping detection result.

11. The performance data comprises a time series of at least one numerical recognition score indicating a recognition area of interest, the graphical user interface comprises at least a corresponding timeline of the numerical recognition score, and for each time step, the numerical recognition score timeline comprises a visual indication of the numerical recognition score associated with the time step. The method according to claim 1.

12. The method according to claim 11, wherein the time series of the numerical recognition scores is a time series of difficulty scores indicating a measure of the difficulty for the recognition system at each time step.

13. The performance data comprises a time series of at least one user-defined score, the graphical user interface comprises at least one corresponding custom timeline, and for each time step, the custom timeline comprises a visual indication of the user-defined score evaluated at the time step. The method according to claim 1.

14. The method according to claim 1, wherein the driving run is a simulated driving run and the recognition error comprises a simulated recognition error.

15. One or more recognition error models are used to provide the simulated recognition error and convert the state of the ground truth simulator into a realistic recognition output provided to the upper-level components of the stack of the test object during the simulation. The method according to claim 14.

16. The simulated recognition error is derived based on synthetic sensor data and simulation ground truth, and the synthetic sensor data is generated in the simulation and processed by the recognition system. The method according to claim 14.

17. A computer system comprising one or more computers configured to implement the method according to any one of claims 1 to 16. **Claim 18** A computer program comprising executable program instructions for programming a computer system to implement the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Data recording and analyzing device for vehicle traveling

    JP1999125584A

  • Abnormality detection device and abnormality detection method

    JP2018079732A

  • Method and apparatus for testing automatic driving vehicle, and storage medium

    JP2020042014A

  • Display control device, data collection system, and display control method

    JP2021019275A

  • Error modeling framework

    US20210094540A1