Test Visualization Tool

A graphical user interface for autonomous vehicles enhances interpretability by visualizing time plots of numerical performance scores, allowing users to understand rule failures in driving scenarios, addressing the challenge of complex rule-based evaluation.

JP7701482B2Active Publication Date: 2025-07-01FIVE AI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023575617
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-01
Filing Date
2022-06-08
Publication Date
2025-07-01
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

Existing methods for testing autonomous vehicles struggle with the interpretability of rule-based evaluation results, particularly in complex scenarios, making it difficult for users to understand the relationship between rule evaluations and actual driving events.

Method used

A graphical user interface that visualizes driving scenarios with time plots of numerical performance scores, allowing users to select points in time to see the corresponding events and rule failures, enhancing the interpretability of driving behavior.

Benefits of technology

Facilitates the easy and reliable assessment of autonomous vehicle performance by providing a clear visualization of rule failures and their impact on driving scenarios, improving the understanding of complex rule interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701482000001
    Figure 0007701482000001
  • Figure 0007701482000002
    Figure 0007701482000002
  • Figure 0007701482000003
    Figure 0007701482000003
Patent Text Reader

Abstract

1. A computer system for rendering a graphical user interface for visualizing a run of a driving scenario in which an own agent progresses through a road layout, the computer system comprising: an input configured to receive a map of the road layout and run data, the run data comprising a sequence of time-stamped own agent states and a time-varying numerical score quantifying the own agent's performance with respect to each rule of a set of run evaluation rules; and a rendering component configured to cause the graphical user interface to display a plot of the time-varying numerical scores for each rule, a marker indicating a selected time index of the plot, the marker being moveable along a time axis to change the selected time index; and a scenario visualization comprising a visualization of the run at the selected time index, wherein moving the marker along the time axis causes the scenario visualization to be updated when the time index is changed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a computer system and method for visualizing and evaluating the behavior of a mobile robot.

Background Art

[0002] In the field of autonomous vehicles, there has been a large and rapid development. An autonomous vehicle (AV) is a vehicle equipped with sensors and a control system that enables it to operate without human control of its behavior. Autonomous vehicles are equipped with sensors that enable them to perceive their physical environment, such sensors including, for example, cameras, radars, and lidars. Autonomous vehicles are equipped with a properly programmed computer that can process the data received from the sensors and make safe and predictable decisions based on the context recognized by the sensors. Autonomous vehicles may be fully autonomous or semi-autonomous (in that they are designed to operate without human supervision or intervention at least in certain situations). Semi-autonomous systems require various levels of human monitoring and intervention, such systems including advanced driver assistance systems and level 3 autonomous driving systems. Testing the behavior of the sensors and control systems installed in a particular autonomous vehicle or a certain type of autonomous vehicle has various aspects. For example, other mobile robots for transporting goods in industrial areas inside and outside are being developed. Such mobile robots are not manned and belong to a class of mobile robots called UAVs (unmanned autonomous vehicles). Autonomous aerial mobile robots (drones) are also being developed.

[0003] In autonomous driving, the importance of guaranteed safety is recognized. Guaranteed safety does not necessarily imply zero accidents, but rather means ensuring that a minimum safety level is met in defined situations. In general, for autonomous driving to be viable, this minimum safety level must significantly exceed that of human drivers.

[0004] Rule - based models may be used to test the performance of various aspects of autonomous vehicles in real - world driving scenarios as well as in simulations. These models provide the criteria that the autonomous vehicle stack should meet in order to be considered safe. To expose the test to potentially dangerous scenarios, a large number of real - world driving runs or simulated driving runs need to be evaluated. Therefore, a large amount of real or simulated driving data needs to be processed in the test. The rules defined for rule - based test models are applied to each of the real or simulated driving scenarios to generate a set of test results, which can be complex and difficult for users to interpret. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0005] The RSS model provides a rule-based model for evaluating the planning and control of an autonomous vehicle stack by testing the behavior of the ego-agent. Other aspects of the performance of the autonomous vehicle may also be tested using the rule-based model. For example, recognition errors of a real or simulated autonomous vehicle stack are determined based on recognition ground truth (which may be simulated ground truth or "pseudo" ground truth generated from real-world sensor data). The user can define a set of recognition error rules and evaluate the determined recognition errors against these rules to assess whether the recognition output of the autonomous vehicle stack is within acceptable accuracy criteria.

[0006] In a rule-based test of an autonomous vehicle stack, the driving performance of the self-agent is evaluated against one or more defined rules in both real-world driving scenarios and simulations. These rules can include driving rules that evaluate the behavior of the self-agent based on some model of safe driving behavior expected in similar driving scenarios, and / or perception rules that evaluate the accuracy of the self-vehicle's perception of its surroundings. Many rules may be defined for each scenario, and it is important that these rule evaluation results regarding both individual scenarios and a set of aggregated results representing the performance of the self-agent across multiple scenarios are interpretable by the user at test time. One way to provide results that are interpretable at the scenario level is to provide a graphical user interface for displaying the results for each given scenario instance (or "run") in which the self-agent drives through a given set of conditions (real or simulated). In one exemplary graphical user interface, the visualization of the scenario is presented along with a set of timelines for each rule indicating whether the rule was passed or failed during the run. This visualization provides the user with a useful summary of the rules passed and failed in a given run, providing an overall summary of the performance of the self-vehicle in that run. The rules may be definable by the user and / or may be arbitrarily complex. A numerical score for each rule may be provided in the user interface, and multiple conditions may contribute to that rule and thus to its numerical performance score. This flexibility is desirable for accommodating the subtle differences in driving across a wide range of real-world / realistic driving runs, but in particular for rules based on particularly complex rules and / or multiple conditions, the user may have difficulty interpreting the direct relationship between the rule evaluation results and the events of the scenario. The interpretability of driving evaluation rules in an AV performance test is one of the technical problems addressed herein.

Means for Solving the Problems

[0007] Described herein is a system for visualizing the driving behavior of an autonomous agent, providing a visualization of a scenario along with a set of time plots of the numerical performance of the autonomous agent based on each set of rules. Since time markers are provided to the user for each of the scenario visualization and the plots associated with each rule, the user can select a given point in time within the scenario and visualize what is happening within the scenario at that time, and the time markers for each rule move to the corresponding time steps of the plot of the numerical performance of the host vehicle with respect to that rule, enabling the user to quickly identify how a rule failure corresponds to the actual events of the scenario. This enables the user to visualize the relationship between the defined rules, the behavior of the autonomous agent, and other conditions of the scenario at any given point in time during driving. This novel graphical user interface mechanism makes the numerical performance score more interpretable for the user, regardless of how the underlying rules are defined.

[0008] A first aspect in this specification is a computer system for rendering a graphical user interface for visualizing the driving of a driving scenario in which an autonomous agent progresses along a road layout, the computer system comprising at least one input configured to receive a map of the road layout of the driving scenario and driving data of the driving of the driving scenario, the driving data comprising a sequence of time-stamped autonomous agent states and a time-varying numerical score quantifying the performance of the autonomous agent, calculated by applying a set of driving evaluation rules to the driving for each rule in the set of driving evaluation rules, at least one input, and a rendering component configured to generate rendering data, the rendering data comprising, for each rule in the set of driving evaluation rules, a plot of the time-varying numerical score on the graphical user interface and a marker indicating a selected time index on the time axis of the plot, the marker being movable along the time axis via user input in the graphical user interface to change the selected time index, a scenario visualization comprising a visualization of the road layout overlaid with a visualization of the agent of the driving at the selected time index, wherein moving the marker along the time axis causes the rendering component to update the scenario visualization when the selected time index is changed, and a rendering component for displaying the scenario visualization.

[0009] The input may be further configured to receive second driving data of a second driving of a driving scenario, the second driving data comprising a second sequence of agent states with timestamps and a time-varying numerical score quantifying the performance of the agent, calculated by applying driving evaluation rules for each rule of a set of driving performance and / or recognition rules to the driving. The rendering component is further configured to generate rendering data, the rendering data being a second plot of the time-varying numerical score of the second driving for each rule of a set of driving performance and / or recognition rules on a graphical user interface, wherein the time-varying numerical score of the driving and the time-varying numerical score of the second driving are plotted with respect to a common set of axes having at least a common time axis, and the marker indicates a selected time index on the common time axis, the second plot, and a second agent visualization of the second driving at the selected time index, the scenario visualization being for displaying the second agent visualization overlaid with the second agent visualization.

[0010] When testing an ego-agent in a simulation or real-world driving scenario, for a single scenario, multiple runs where aspects of the agent's configuration or behavior vary from run to run may be evaluated. In this case, evaluating the rules and metrics for each run, as is, does not provide a detailed depiction of how differences in the agent's configuration and / or behavior affect the progression of the scenario. Described herein is a system with a run comparison user interface where two driving runs can be compared within a common scenario visualization along a common time interval, the user can interactively select the time index of the scenario, and the user interface displays a visualization of the state of the vehicle at that point in time for each of the two runs. This enables comparison of the behavior of the vehicle over two runs in the playback of the scenario, allowing the user to identify specific actions or features of each run that affect performance improvement or degradation.

[0011] A time-varying numerical score may be calculated by applying one or more rules to a time-varying signal extracted from the run data, and the change in the signal can be viewed in the scenario visualization.

[0012] The rendering component may be configured to remove a plot of the time-varying numerical score of the deselected run from a set of common axes for each driving rule and remove the agent visualization of the deselected run from a single visualization of the road layout in response to a deselection input in the graphical user interface indicating one of the first run and the second run, and the user can switch from a run comparison view regarding both the first run and the second run to a single run view regarding only one of the first run and the second run.

[0013] The graphical user interface may further include a comparison table having entries for each rule of a set of driving evaluation rules, the entries including an aggregated performance result for that rule in a first drive and an aggregated performance result for that rule in a second drive.

[0014] The entry for each rule may further include an explanation of that rule.

[0015] The rendering component may be configured to, in response to an expansion input in the graphical user interface, hide the plot of the time-varying numerical score for each rule and display a timeline view with an indication of the pass / fail result of the rule over time.

[0016] The rendering component may be configured to display, in the graphical user interface, the numerical score at a selected time index for each rule of a set of driving evaluation rules.

[0017] The driving evaluation rules may include recognition rules, and the scenario visualization includes a set of recognition outputs generated by a recognition component of the host vehicle.

[0018] The scenario visualization may include sensor data superimposed on a visualization of a road layout.

[0019] The scenario visualization may include a scenario timeline having scenario time markers, and moving a marker along the scenario timeline causes the rendering component to update the respective time markers of each plot of the time-varying numerical score when the selected time index is changed.

[0020] The scenario timeline may include a frame index corresponding to a selected time index and a set of controls for moving forward or backward by incrementing or decrementing the frame index respectively.

[0021] The driving scenario may be a simulated driving scenario in which a simulated ego agent progresses through a simulated road layout, and the driving data is received from a simulator.

[0022] The driving scenario may be a real-world driving scenario in which an ego agent progresses through a real-world road layout, and the driving data is calculated based on data generated on the ego agent during driving.

[0023] The plot of the numerically scored time series includes an xy plot of the numerically scored time series.

[0024] Alternatively or additionally, the numerically scored time series is plotted using color coding.

[0025] A second aspect herein is a method for visualizing the driving of a driving scenario in which an autonomous agent progresses along a road layout, the method comprising receiving a map of the road layout of the driving scenario and driving data of the driving of the driving scenario, the driving data comprising a sequence of timestamped autonomous agent states and a time-varying numerical score quantifying the performance of the autonomous agent, calculated by applying a driving evaluation rule to the driving for each rule in a set of driving evaluation rules, generating rendering data, the rendering data comprising, for each rule in a set of driving evaluation rules, a plot of the time-varying numerical score on a graphical user interface and a marker indicating a selected time index on the time axis of the plot, the marker being movable along the time axis via user input in the graphical user interface to change the selected time index, and a scenario visualization comprising a visualization of the road layout overlaid with an agent visualization of the driving at the selected time index, wherein moving the marker along the time axis causes the rendering component to update the scenario visualization when the selected time index is changed, and displaying the scenario visualization.

[0026] A further aspect herein provides a computer program comprising executable instructions for programming a computer system to implement the functionality of the method or system of any preceding claim.

[0027] For a better understanding of the present disclosure, and to show how embodiments thereof may be carried out, reference is made, by way of example only, to the following figures.

Brief Description of the Drawings

[0028]

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0029] In one exemplary graphical user interface disclosed in international patent applications PCT / EP2022 / 053413 and PCT / EP2022 / 053406, the visualization of a scenario is presented together with a set of timelines for each rule indicating whether the rule was passed or failed during a drive. This visualization provides the user with a useful summary of the rules passed and the rules that were failed in a given drive, providing an overall summary of the vehicle's performance in that drive. However, as will be described later, the rules are definable by the user and can be arbitrarily complex, and multiple conditions contribute to the numerical performance score provided to the user interface, making it difficult for the user to interpret the direct relationship between the rule evaluation results and the events of the scenario.

[0030] The described embodiments provide a test pipeline to facilitate rule-based testing of a mobile robot stack in a real or simulated scenario. A set of interactive graphical user interface (GUI) features improves the interpretability of the rules applied, enabling an expert to more easily and reliably assess the performance of the stack in a given driving scenario from the GUI output.

[0031] Typically, a "full" stack includes everything from the processing and interpretation (recognition) of lower-level sensor data to the input to major higher-level functions such as prediction and planning, and the control logic to generate appropriate control signals for making planning-level decisions (such as for controlling brakes, steering, acceleration, etc.). In the case of an autonomous vehicle, a level 3 stack includes logic for implementing a takeover request, and a level 4 stack additionally includes logic for implementing a minimum-risk maneuver. The stack may also implement secondary control functions such as signaling, headlights, windshield wipers, etc.

[0032] The term "stack" may also refer to individual sub-systems (sub-stacks) of the full stack, such as recognition, prediction, planning, or control stacks, which may be tested individually or in any desired combination. A stack may refer to purely software, i.e., one or more computer programs that can be executed on one or more general-purpose computer processors.

[0033] The test framework described below provides a pipeline for generating scenario ground-truth from real-world data. This ground-truth may be used as a basis for recognition tests by comparing the generated ground-truth with the recognition output of the recognition stack under test, as well as by evaluating driving behavior against driving rules.

[0034] The behavior of an agent (actor) in a real or simulated scenario is evaluated by a test oracle based on defined performance evaluation rules. Such rules may evaluate various aspects of safety. For example, a set of safety rules may be defined to assess the performance of a stack against specific safety criteria, regulations, or safety models (such as RSS), or a custom set of rules may be defined to test any aspect of performance. The test pipeline is not limited in its use to safety and can be used to test any aspect of performance, such as comfort or progress towards a defined goal. A rule editor enables performance evaluation rules to be defined or modified and passed to the test oracle.

[0035] Similarly, the recognition of a vehicle can be appraised / evaluated by a "recognition oracle" based on defined recognition rules. These may be defined within a recognition error specification that provides a standard format for defining recognition errors.

[0036] Defining rules in the perception error framework enables areas of interest in real-world driving scenarios to be highlighted to the user, for example, by flagging these areas in a replay of the scenario presented in the user interface, as will be described in more detail below. This enables the user to review the obvious errors in the perception stack and identify the possible reasons for the errors, such as occlusion in the original sensor data. Such an evaluation of perception errors also enables a "contract" to be defined between the perception component and the planning component of the AV stack, where the requirements for perception performance can be specified, and a stack that meets these perception performance requirements promises to enable safe planning. An integrated framework may be used to evaluate real perception errors from real-world driving scenarios and to evaluate simulated errors, either directly simulated using a perception error model or calculated by applying the perception stack to simulated sensor data, such as a realistic simulation of camera images.

[0037] The ground truth determined by the pipeline itself can be evaluated within the same perception error specification by comparing it to the "true" ground truth determined by manually reviewing and annotating the scenario according to defined rules. Finally, the results of applying the perception error test framework can be used to guide the test strategy for testing both the perception subsystem and the prediction subsystem of the stack.

[0038] A scenario, whether real or simulated, requires the self-agent to move within a real or modeled physical context. The self-agent is a real or simulated mobile robot that moves under the control of the stack under test. The physical context includes static and / or dynamic elements that the stack under test is required to effectively address. For example, the mobile robot may be a fully or semi-autonomous vehicle (ego vehicle) under the control of the stack. The physical context may include a static road layout and a set of given environmental conditions (e.g., weather, time, lighting conditions, humidity, pollution / particle level, etc.) that can be maintained or changed as the scenario progresses. An interactive scenario additionally includes one or more other agents ("external" agents, e.g., other vehicles, pedestrians, people on bicycles, animals, etc.).

[0039] Consider the following example as applied to testing an autonomous vehicle. However, the principles are equally applicable to other forms of mobile robots.

[0040] A scenario may be represented or defined at various levels of abstraction. A more abstract scenario is more adaptable to a greater degree of variation. For example, an "interruption scenario" or "lane change scenario" is an example of a highly abstract scenario characterized by an operation or behavior of interest that is adaptable to many variations (e.g., start positions and speeds of various agents, road layout, environmental conditions, etc.). "Scenario run" refers to a specific occurrence where the agent moves within the physical context, optionally in the presence of one or more other agents. For example, multiple runs of an interruption or lane change scenario can be performed (in the real world and / or in a simulator) with different agent parameters (e.g., start position, speed, etc.), different road layouts, different environmental conditions, and / or different stack configurations. The terms "run" and "instance" are used interchangeably in this context.

[0041] In the following example, the performance of the stack is at least partially assessed by evaluating the behavior of the self-agent within the test oracle against a given set of performance evaluation rules over one or more runs. The rules are applied to the "ground truth" of the scenario run (or each scenario run), which generally simply means an appropriate representation of the scenario run (including the behavior of the self-agent) that is considered reliable for the purposes of the test. The ground truth is specific to the simulation, and the simulator calculates a sequence of scenario states, which by definition is a perfectly reliable representation of the simulated scenario run. In a real-world scenario run, there is no "perfect" representation of the scenario run in the same sense, but nevertheless, an appropriate ground truth that provides information can be obtained in a number of ways, for example, through manual annotation of in-vehicle sensor data, automated / semi-automated annotation of such data (using, for example, offline / non-real-time processing), and / or the use of external information sources (such as external sensors, maps, etc.).

[0042] Scenario ground truth typically includes the "trajectories" of the self-agent and, if applicable, any other (notable) agents. A trajectory is the history of an agent's position and movement over a scenario. There are many ways in which a trajectory can be represented. Trajectory data typically includes the spatial and movement data of agents within an environment. Agent trajectories are provided with a sequence of timestamped agent states for each agent, enabling the visualization of the agent's state at various time steps. This term is used in relation to both real scenarios (having trajectories in the real world) and simulated scenarios (having simulated trajectories). A trajectory typically records the actual path realized by an agent within a scenario. In terms of the terms, "trajectory" and "path" may include the same or similar types of information (such as a sequence of spatial and movement states over time). The term "path" is generally often used in the context of a plan (and may refer to a future / predicted path), and the term "trajectory" is generally often used in the context of testing / evaluation in relation to past behavior.

[0043] In a simulation context, a "scenario description" is provided as input to the simulator. For example, the scenario description may be encoded using a scenario description language (SDL) or in any other format that can be used by the simulator. The scenario description is typically a more abstract representation of a scenario and can cause multiple simulated runs. Depending on the implementation, the scenario description may have one or more configurable parameters that can be changed to increase the degree of possible variation. The degree of abstraction and parameterization is a design choice. For example, the scenario description may encode a fixed layout using parameterized environmental conditions (e.g., weather, lighting, etc.). However, for example, further abstraction is possible using configurable road parameters (e.g., road curvature, lane configuration, etc.). The input to the simulator includes the scenario description along with a set of selected parameter values (if applicable). The latter may be referred to as the parameterization of the scenario. The configurable parameters define a parameter space (also called a scenario space), and the parameterization corresponds to a point within the parameter space. In this context, a "scenario instance" may refer to the instantiation of a scenario in the simulator based on the scenario description and (if applicable) the selected parameterization.

[0044] For the sake of brevity, the term "scenario" may also be used to refer to a scenario run, not just in the more abstract sense of a scenario. The meaning of the term "scenario" will be clear from the context in which it is used.

[0045] Exemplary AV stack: To provide context related to the described embodiments, further details of an exemplary form of an AV stack are described here.

[0046] FIG. 1A shows a very schematic block diagram of the AV driving stack 100. The driving stack 100 is shown to include a recognition (sub) system 102, a prediction (sub) system 104, a planning (sub) system (planner) 106, and a control (sub) system (controller) 108. As described above, the term (sub) stack may also be used to describe the foregoing components 102-108.

[0047] In the real-world context, the recognition system 102 receives sensor outputs from the AV's in-vehicle sensor system 110 and uses those sensor outputs to detect external agents and measure their physical states, such as their position, velocity, acceleration, etc. The in-vehicle sensor system 110 can take various forms but generally includes various sensors such as image capture devices (cameras / optical sensors), lidar and / or radar units, satellite positioning sensors (such as GPS), motion / inertial sensors (accelerometers, gyroscopes, etc.). Thus, the in-vehicle sensor system 110 provides rich sensor data from which it is possible to extract detailed information about the surrounding environment and the state of the AV and any external actors (vehicles, pedestrians, people on bicycles, etc.) within that environment. Typically, the sensor outputs include sensor data from multiple sensor modalities, such as stereo images from one or more stereo optical sensors, lidar, radar, etc. The sensor data from multiple sensor modalities may be combined using filters, fusion components, etc.

[0048] The recognition system 102 typically includes a plurality of recognition components that cooperate to interpret the sensor outputs and provide a recognition output to the prediction system 104.

[0049] In the simulation context, depending on the nature of the test, particularly depending on where the stack 100 is "sliced" for the test (see below), there are cases where it is necessary to model the in-vehicle sensor system 100 and cases where it is not. In the upper-level slicing, since simulated sensor data is not required, complex sensor modeling is not necessary.

[0050] The recognition output from the recognition system 102 is used by the prediction system 104 to predict the future behavior of external actors (agents) such as other vehicles in the vicinity of the AV.

[0051] The prediction calculated by the prediction system 104 is provided to the planner 106, and the planner 106 uses the prediction to make a decision on autonomous driving to be executed by the AV in a given driving scenario. The input received by the planner 106 typically indicates the drivable area and also captures the predicted movement of external agents (obstacles from the perspective of the AV) within the drivable area. The drivable area can be determined by combining the recognition output from the recognition system 102 with map information such as an HD (high-definition) map.

[0052] The core function of the planner 106 is to plan the trajectory (ego-trajectory) of the AV considering the predicted movement of the agents. This may be referred to as trajectory planning. The trajectory is planned to accomplish the desired goal within the scenario. The goal can be, for example, entering a roundabout and exiting at the desired exit, overtaking the vehicle in front, or staying in the current lane at the target speed (lane following). The goal may be determined, for example, by an autonomous route planner (not shown).

[0053] The controller 108 executes the decisions made by the planner 106 by providing appropriate control signals to the in-vehicle AV actor system 112. Specifically, the planner 106 plans the AV's trajectory, and the controller 108 generates control signals for implementing the planned trajectory. Typically, the planner 106 plans for the future such that the planned trajectory can be implemented only partially at the control level, and then a new trajectory is planned by the planner 106. The actor system 112 includes "primary" vehicle systems such as the brake, acceleration, and steering systems, as well as secondary systems (e.g., signals, wipers, headlights, etc.).

[0054] Note that there may be a difference between the planned trajectory at a given instant and the actual trajectory followed by the self-agent. The planning system typically operates over a sequence of planning steps, updating the planned trajectory at each planning step to account for changes in the scenario since the previous planning step (or, more precisely, changes that deviate from the predicted changes). The planning system 106 may reason forward in time such that the planned trajectory at each planning step deviates from the next planning step. Thus, the individual planned trajectories may not be fully realized (when the planning system 106 is tested alone in a simulation, the self-agent may simply follow the planned trajectory accurately up to the next planning step, but as noted above, in other real-world contexts and simulation contexts, the planned trajectory may not be accurately followed up to the next planning step because the behavior of the self-agent may be affected by other factors such as the operation of the control system 108 and the real or modeled dynamics of the host vehicle). In many test contexts, ultimately what matters is the actual trajectory of the self-agent, specifically whether the actual trajectory is safe, as well as other factors such as comfort and progress. However, the rule-based test approach herein can also be applied to the planned trajectories (even if those planned trajectories are not fully or accurately realized by the self-agent). For example, even if the actual trajectory of the agent is considered safe according to a given safety rule, the instantaneous planned trajectory may have been unsafe, and the fact that the planner 106 considered an unsafe course of action may become apparent even if it did not lead to unsafe behavior of the agent within the scenario. The instantaneous planned trajectory constitutes one form of internal state that can usefully be evaluated in addition to the actual behavior of the agent in the simulation. Other forms of internal stack state can similarly be evaluated.

[0055] The example of FIG. 1A contemplates a relatively “modular” architecture having separable perception, prediction, planning, and control systems 102-108. The sub-stack itself may also be modular, for example, having separable planning modules within the planning system 106. For example, the planning system 106 may comprise multiple trajectory planning modules that can be applied to different physical contexts (e.g., simple lane driving versus complex intersections or roundabouts). This is relevant to simulation testing for the reasons above, as it allows components (such as the planning system 106 or its individual planning modules) to be tested individually or in different combinations. To avoid misunderstanding, in a modular stack architecture, the term stack may refer not only to the full stack but also to its individual sub-systems or modules.

[0056] To what extent various stack functions are integrated or separable can vary significantly between different stack implementations, and in some stacks, certain aspects may be so closely coupled as to be indistinguishable. For example, in other stacks, planning and control may be integrated (e.g., such a stack may be able to directly perform planning in terms of control signals), while other stacks (e.g., those shown in FIG. 2B) may be designed in a way that clearly differentiates between these two (e.g., perform planning in terms of trajectories and perform independent control optimization to determine the best way to execute the planned trajectories at the control signal level). Similarly, in some stacks, prediction and planning may be more closely coupled. In extreme cases, in so-called “end-to-end” driving, perception, prediction, planning, and control may be essentially inseparable. Unless otherwise specified, the terms perception, prediction, planning, and control as used herein do not imply any particular coupling or modularization of these aspects.

[0057] The term "stack" encompasses software, but it should be understood that it can also encompass hardware. In simulation, the stack software may be tested on a "general-purpose" non-vehicle computer system before ultimately being uploaded to the in-vehicle computer system of the physical vehicle. However, in "hardware-in-the-loop" testing, the testing may extend to the hardware that forms the basis of the vehicle itself. For example, the stack software may be executed on an in-vehicle computer system (or a replica thereof) coupled to a simulator for testing purposes. In this context, the stack being tested extends to the computer hardware that forms the basis of the vehicle. As another example, a particular function of the stack 100 (e.g., a recognition function) may be implemented in dedicated hardware. In a simulation context, hardware-in-the-loop testing can include supplying synthetic sensor data to the recognition component of the dedicated hardware.

[0058] Test oracle Figure 1B shows a very high-level overview of the test paradigm for autonomous vehicles. For example, an ADS / ADAS stack 100 of the type shown in Figure 1A undergoes iterative testing and evaluation in simulation by running multiple scenario instances on a simulator 202 and evaluating the performance of the stack 100 (and / or its individual sub-stacks) with a test oracle 252. The output of the test oracle 252 is useful to an expert 122 (team or individual), enabling the expert 122 to identify issues within the stack 100 and modify the stack 100 to mitigate those issues (S124). This result also helps the expert 122 select additional scenarios for testing (S126), and the process continues, repeatedly modifying, testing, and evaluating the stack 100 in simulation. The improved stack 100 is ultimately incorporated into a real-world AV101 equipped with a sensor system 110 and an actuator system 112 (S125). The improved stack 100 typically includes program instructions (software) executed by one or more computer processors of an in-vehicle computer system (not shown) of the vehicle 101. The software of the improved stack is uploaded to the AV101 in step S125. Step 125 may also include changes to the underlying vehicle hardware. Once the improved stack 100 is installed on the AV101, it receives sensor data from the sensor system 110 and outputs control signals to the actuator system 112. Real-world testing (S128) can be used in combination with simulation-based testing. For example, once an acceptable level of performance is reached through the simulation testing and stack improvement process, appropriate real-world scenarios may be selected (S130), the performance of the AV101 in those real scenarios may be captured and similarly evaluated with the test oracle 252.

[0059] Scenarios can be obtained in various ways, including manual coding, for the purpose of simulation. This system is also capable of extracting scenarios from real-world driving for the purpose of simulation, enabling real-world situations and their variations to be recreated within the simulator 202.

[0060] Figure 1C shows a very schematic block diagram of a scenario extraction pipeline. Real-world driving data 140 is passed to a "ground-truthing" pipeline 142 for the purpose of generating scenario ground truth. The driving data 140 can include, for example, sensor data and / or recognition outputs captured / generated on one or more vehicles (which can be autonomous, human-driven, or a combination thereof), and / or data captured from other sources such as external sensors (e.g., CCTV). The driving data is processed within the ground-truthing pipeline 142 to generate appropriate ground truth 144 (trajectory and context data) for real-world driving. As discussed, the ground-truthing process can be based on manual annotation of "raw" driving data 140, or the process can be fully automated (e.g., using offline recognition methods), or a combination of manual and automated ground-truthing can be used. For example, 3D bounding boxes can be placed around the vehicles and / or other agents captured in the driving data 140 to determine the spatial and motion states of their trajectories. The scenario extraction component 146 receives the scenario ground truth 144 and processes the scenario ground truth 144 to extract a more abstracted scenario description 148 that can be used for simulation purposes. The scenario description 148 is used by the simulator 202 to enable multiple simulated drives to be performed. The simulated drives are a variation of the original real-world drive, and the degree of possible variation is determined by the degree of abstraction. Ground truth 150 is provided for each simulated drive.

[0061] The scenario extraction shown in FIG. 1C may be used to extract real-world driving data for testing and visualization, i.e., as described later, the actual vehicle state of the driving may be extracted for visualization of the real-world driving in the recognition visualization user interface. It should be noted that the term "driving data" is used herein to refer to both the "raw" driving data collected by the host vehicle in real-world driving, such as sensor data, and the driving data processed for testing and visualization, including the "ground truth" trajectory and context data of the host vehicle. In the context of rule evaluation, the driving data may also include a numerical score measuring the performance of the host vehicle against one or more driving evaluation rules, which may include recognition error rules or driving rules, as described in more detail below.

[0062] The test oracle 252 applies a rule - based model to evaluate the actual or simulated behavior of an autonomous vehicle stack (also referred to herein as an agent) determined by the planner 106. However, the test paradigm in the simulation context shown in Figure 1B can also be implemented in the context of evaluating recognition errors, in which case a "recognition oracle" takes the place of the test oracle 252 and processes scenarios to evaluate recognition errors against rules defined within a similar rule - based model described below for planning, which is referred to herein as the "recognition error framework". A recognition error is obtained by comparing the recognition output generated by the recognition component 102 with the scenario ground truth, where the scenario ground truth is specific to the simulation and can also be generated for real - world scenarios using the ground - truing pipeline 142. The evaluation of recognition errors within the recognition error framework is described in more detail below. The evaluation of recognition errors is also described in UK Patent Applications Nos. 2108182.3, 2108958.6, 2108952.9, and 2111765.0, which are hereby incorporated by reference in their entirety.

[0063] Simulation context Next, further details of the test pipeline and the test oracle 252 are described. The following example focuses on simulation - based testing. However, as noted above, the test oracle 252 can be similarly applied to evaluate stack performance in real scenarios, and the following related explanations apply equally to real scenarios. The following explanation refers to the stack 100 in Figure 1A by way of example. However, as noted above, the test pipeline 200 is highly flexible and can be applied to any stack or sub - stack operating at any level of autonomy.

[0064] Figure 2A shows a schematic block diagram of a test pipeline 200. The test pipeline 200 is shown to include a simulator 202 and a test oracle 252. The simulator 202 runs a simulated scenario for the purpose of testing all or part of the AV driving stack, and the test oracle 252 evaluates the performance of the stack (or sub-stack) in the simulated scenario. The following description refers to the stack of Figure 1A by way of example. However, the test pipeline 200 is highly flexible and can be applied to any stack or sub-stack operating at any level of autonomy.

[0065] As described above, the idea of simulation-based testing is to run a simulated driving scenario in which the ego-agent must progress under the control of the stack (or sub-stack) being tested. Typically, the scenario includes a static drivable area (e.g., a specific static road layout) in which the ego-agent is required to progress in the presence of one or more other dynamic agents (e.g., other vehicles, bicycles, pedestrians, etc.). The simulated inputs are fed to the stack under test and used for decision-making there. The ego-agents are made to execute those decisions to simulate the behavior of an autonomous vehicle in those situations.

[0066] The simulated input 203 is provided to the stack under test. "Slicing" refers to the selection of a set or subset of stack components for testing. This determines the form of the simulated input 203.

[0067] As an example, FIG. 2A shows the prediction, planning, and control systems 104, 106, and 108 within the AV stack 100 during testing. To test the full AV stack of FIG. 1A, the recognition system 102 can also be applied during testing. In this case, the simulated input 203 includes synthetic sensor data that is generated using appropriate sensor models and processed within the recognition system 102 in a manner similar to real sensor data. This requires the generation of sufficiently realistic synthetic sensor inputs (e.g., realistic image data and / or similarly realistic simulated lidar / radar data, etc.). The resulting output of the recognition system 102 is then supplied to the higher-level prediction and planning systems 104, 106.

[0068] In contrast, so-called "planning-level" simulations basically bypass the recognition system 102. Instead, the simulator 202 directly provides a simpler, higher-level input 203 to the prediction system 104. In some contexts, it may even be appropriate to bypass the prediction system 104 in order to test the planner 106 based on predictions directly obtained from the simulated scenarios.

[0069] Between these two extremes, there is room for many different levels of input slicing, for example, testing only a subset of the recognition system, such as "late-stage" recognition components, i.e., components such as filters or fusion components that act on the outputs from lower-level recognition components (e.g., object detectors, bounding box detectors, motion detectors, etc.).

[0070] As just one example, the description of the test pipeline 200 refers to the in-motion stack 100 of FIG. 1A. As discussed, only a sub-stack of the in-motion stack may be tested, but for simplicity, the following description refers to the AV stack 100 throughout. Thus, in FIG. 2, the reference number 100 may represent either the full AV stack or only a sub-stack depending on the context.

[0071] Regardless of its form, the simulated input 203 is used (either directly or indirectly) as the basis for the decision-making by the planner 108.

[0072] The controller 108 then implements the planner's decision by outputting the control signal 109. In the real-world context, these control signals drive the physical actuator system 112 of the AV.

[0073] In the simulation, the resulting control signal 109 is converted into the realistic movement of the ego-agent within the simulation by using the ego-vehicle dynamics model 204, thereby simulating the physical response of the autonomous vehicle to the control signal 109.

[0074] To the extent that external agents exhibit autonomous behavior / decision-making within the simulator 202, some form of agent decision logic 210 is implemented to make those decisions and determine the behavior of the agents within the scenario. The agent decision logic 210 may be of equal complexity to the self-stack 100 itself or may have more limited decision-making capabilities. The aim is to provide sufficiently realistic behavior of external agents within the simulator 202 in order to usefully test the decision-making capabilities of the self-stack 100. In some contexts, this may not require any agent decision logic 210 at all (open-loop simulation), while in other contexts, useful tests can be provided using relatively limited agent logic 210 such as basic adaptive cruise control (ACC). One or more agent dynamics models 206 may be used to provide more realistic agent behavior.

[0075] The simulation of the driving scenario is run according to the scenario description 201, which has both a static layer 201a and a dynamic layer 201b.

[0076] The static layer 201a defines the static elements of the scenario, which typically includes a static road layout.

[0077] The dynamic layer 201b defines dynamic information about external agents within the scenario, such as other vehicles, pedestrians, bicycles, etc. The range of dynamic information provided can vary. For example, the dynamic layer 201b may include, for each external agent, the spatial path traversed by that agent, along with one or both of the motion data and behavior data associated with that path. In a simple open-loop simulation, the external actors simply follow the spatial paths and motion data defined in the dynamic layer that are non-reactive, i.e., do not react to the ego-agent within the simulation. Such an open-loop simulation can be implemented without the agent decision logic 210. However, in a closed-loop simulation, the dynamic layer 201b instead defines at least one behavior (e.g., the ACC behavior) that is followed along a static path. In this case, the agent decision logic 210 implements that behavior in a reactive manner within the simulation, i.e., reactively to the ego-agent and / or other external agents. The motion data may still be associated with the static path, but in this case is less prescriptive and may, for example, serve as a goal along the path. For example, in the ACC behavior, a target speed can be set along the path that the agent is trying to match, but the agent decision logic 210 may be permitted to reduce the speed of an external agent below the target at any point along the path to maintain a target inter-vehicle distance with the leading vehicle.

[0078] The output of the simulator 202 for a given simulation includes the ego-trajectory 212a of the ego-agent and one or more agent trajectories 212b (trajectories 212) of one or more external agents.

[0079] A trajectory is the complete history of an agent's behavior within a simulation that has both spatial and motion components. For example, a trajectory may take the form of a spatial path with motion data associated with points along the path, such as speed, acceleration, jerk (rate of change of acceleration), snap (rate of change of jerk), etc.

[0080] Additional information is provided to supplement the trajectory 212 and provide context thereto. Such additional information is referred to as "environmental" data 214 which can have both static components (e.g., road layout) and dynamic components (e.g., weather conditions within a range that changes over the simulation). The environmental data 214 may be somewhat "pass-through" in that it is directly defined by the scenario description 201 and not affected by the results of the simulation. For example, the environmental data 214 may include a static road layout directly provided by the scenario description 201. However, typically, the environmental data 214 includes at least some elements derived within the simulator 202. This can include, for example, simulated weather data, and the simulator 202 can freely change the weather conditions as the simulation progresses. In that case, the weather data may be time-dependent and that time-dependence is reflected in the environmental data 214.

[0081] The test oracle 252 receives the trajectory 212 and the environmental data 214 and scores their outputs in a manner described below. The scoring is time-based, and for each performance metric, the test oracle 252 tracks how the value (score) of that metric changes over time as the simulation progresses. The test oracle 252 provides an output 256 that includes a score-time plot for each performance metric, as will be described in more detail later. The scores are output and stored in a database 258, where the scores can be accessed and, for example, the results can be displayed to a user interface as described above. The metric 254 is useful to an expert, and the scores can be used to identify and mitigate performance issues within the tested stack 100.

[0082] Recognition error model FIG. 2B shows a particular form of slicing and uses reference numerals 100 and 100S to represent a full stack and a sub-stack, respectively. It is the sub-stack 100S that is the subject of testing within the test pipeline 200 of FIG. 2A.

[0083] Some “late” recognition components 102B form part of the sub-stack 100S being tested and are applied to the simulated recognition input 203 during testing. The late recognition components 102B can include filtering or other fusion components that fuse recognition inputs from multiple early recognition components.

[0084] In the full stack 100, the late recognition component 102B receives the actual recognition input 213 from the early recognition component 102A. For example, the early recognition component 102A may comprise one or more 2D or 3D bounding box detectors, in which case the simulated recognition input provided to the late recognition component can include simulated 2D or 3D bounding box detection results derived by ray tracing in simulation. The early recognition component 102A generally includes components that act directly on the sensor data.

[0085] In this slicing, the simulated recognition input 203 formally corresponds to the actual recognition input 213 that is normally provided by the early recognition component 102A. However, rather than being applied as part of a test, the early recognition component 102A is instead used to train one or more recognition error models 208, and the recognition error models 208 can be used to introduce realistic errors in a statistically rigorous way into the simulated recognition input 203 that is supplied to the late recognition component 102B of the sub-stack 100 under test.

[0086] Such an error recognition model may be referred to as a Perception Statistical Performance Model (PSPM), or equivalently as "PRISM". Further details of the principles of PSPM, and appropriate techniques for constructing and training PSPM, can be found in International Patent Applications PCT / EP2020 / 073565, PCT / EP2020 / 073562, PCT / EP2020 / 073568, PCT / EP2020 / 073563 and PCT / EP2020 / 073569, which are hereby incorporated by reference in their entirety. The idea behind PSPM is to efficiently introduce realistic errors into the simulated recognition inputs provided to sub-stack 100S (i.e., to reflect the types of errors that would be expected if the early recognition component 102A were applied in the real world). In the simulation context, "perfect" ground truth recognition inputs 203G are provided by the simulator, but these are used to derive more realistic recognition inputs 203 that have realistic errors introduced by the error recognition model 208.

[0087] As described in the foregoing cited documents, PSPM can depend on one or more variables ("interference factors") representing physical conditions, enabling different levels of error to be introduced that reflect the various real-world conditions that can occur. Thus, the simulator 202 can simulate different physical conditions (e.g., different weather conditions) by simply changing the values of the weather interference factors and thereby changing how the introduction of recognition errors occurs.

[0088] The late recognition component 102b within sub-stack 100S processes the simulated recognition inputs 203 in exactly the same way as it processes real-world recognition inputs 213 within the full stack 100, and its output drives prediction, planning, and control. Alternatively, PSPM can be used to model the entire recognition system 102, including the late recognition component 102B.

[0089] One exemplary rule considered herein for evaluation by the test oracle 252 is the "safe distance" rule that applies in a lane-following context and is evaluated between the ego agent and other agents. The safe distance rule requires that the ego agent always maintain a safe distance from other thresholds. Both lateral and longitudinal distances are considered, and for the safe distance rule to pass, it is sufficient if only one of those distances meets the safety threshold (considering a lane driving scenario where the ego agent and other agents are in adjacent lanes, when driving side by side, the longitudinal distance along the road may be zero or near zero, which is safe if a sufficient lateral distance is maintained between the agents, and similarly, when the ego agent is driving behind another agent in the same lane, assuming both agents are following approximately the center of the lane, the lateral distance perpendicular to the direction of the road may be zero or near zero, which is safe if a sufficient longitudinal distance is maintained). A numerical score is calculated for the safe distance rule at a given point in time based on which distance (lateral or longitudinal) is currently determining safety.

[0090] The safe distance rule is chosen to illustrate the specific principles underlying the methodology being described because it is simple and intuitive. However, it will be understood that the techniques described can be applied to any rule designed to quantify some aspect (or aspects) of driving performance, such as safety, comfort, and / or progress towards some defined goal, as a numerical "robustness score". The time-varying robustness score over the duration of the scenario drive is denoted as s(t), and the overall robustness score of the drive is denoted as y. For example, a robustness scoring framework for driving rules based on signal temporal logic may be constructed.

[0091] In general, robustness scores such as those described below with reference to FIGS. 3A and 3B take quantities such as absolute or relative position, velocity, or other quantities of the relative motion of the agent. A robustness score is typically defined based on a threshold for one or more of these quantities according to a given rule (for example, the threshold may define the minimum lateral distance to the nearest agent that is considered acceptable from a safety perspective). Next, the robustness score is defined by whether a given quantity exceeds or falls below its threshold and by how much that quantity exceeds or falls below the threshold. Thus, the robustness score provides a measure of whether, how, or how an agent is expected to operate in relation to other agents and its environment (including, for example, any speed limits defined within the drivable area in a given scenario). Note that numerical scores can be similarly defined for other aspects of driving assessment, including, for example, the assessment of recognition errors.

[0092] FIG. 3A schematically shows the geometric principle of a safety distance rule evaluated between a self-agent E and another agent C (challenger).

[0093] FIGS. 3A and 3B are described in association with each other.

[0094] The lateral distance is measured along a road reference line (which may be straight or curved), and the longitudinal spacing is measured in a direction perpendicular to the road reference line. The lateral spacing and the longitudinal spacing (the distance between the self-agent E and the challenger C) are represented by d lat and d lon respectively. The lateral distance threshold and the longitudinal distance threshold (safety distance) are represented by d slat and d slon respectively.

[0095] The safety distances d slat and d slonTypically not fixed and typically varies as a function of the relative speed of the agent (and / or other factors such as weather, road curvature, road surface, lighting, etc.). Representing the interval and safety distance as a function of time t, the "headroom" distances in the lateral and longitudinal directions are defined as follows.

[0096] D lat (t)=d lat (t)-d slat (t) D lon (t)=d lon (t)-d slon (t) Figure 3A(1) shows the ego-agent E at a safe distance from the challenger C due to the lateral interval d of the agent being greater than the current lateral safety distance d lat for the agent pair. slat (positive D lat ).

[0097] Figure 3A(2) shows the ego-agent E at a safe distance from the challenger C due to the longitudinal interval d of the agent being greater than the current longitudinal safety distance d lon for the agent pair. slon (negative D lon ).

[0098] Figure 3A(3) shows the ego-agent E at a non-safe distance from the challenger C. Since both D lat and D lon are negative, it fails the safety distance rule.

[0099] Figure 3B shows the safety distance rule implemented as a computational graph applied to a set of scenario ground truth 310 (or other scenario data).

[0100] The lateral interval, lateral safety distance, longitudinal interval, and longitudinal safety distance are each extracted as time-varying signals from the scenario ground truth 310 by the first, second, third, and fourth extractor nodes 302, 304, 312, 314 of the calculation graph 300. The lateral and longitudinal headroom distances are calculated by the first and second calculation (assessor) nodes 306, 316 and converted into robustness scores as follows. The following example considers a normalized robustness score over some fixed range such as [-1,1] with 0 as the passing threshold.

[0101] The headroom distance quantifies the extent to which the associated safety distance is violated or not. A positive lateral / longitudinal headroom distance suggests that the lateral / longitudinal interval between the host vehicle E and the challenger C is greater than the current lateral / longitudinal safety distance, and a negative headroom distance suggests the opposite. According to the principle described above, the robustness scores for the lateral distance and the longitudinal distance may be defined, for example, as follows.

[0102] D lat When (t)>0, s lat (t)=min[1,D lat (t) / A lat D lat When (t)≤0, s lat (t)=max[-1,D lat (t) / B lat D lon When (t)>0, s lon (t)=min[1,D lon (t) / A lon D lon When (t)≤0, s lon (t)=max[-1,D lon (t) / B lon Here, A and B represent some predefined normalization distances (these may be the same or different for the lateral score and the longitudinal score). For example, D​​​​lon When (t) varies between A lon and -B lon it can be seen that the longitudinal robustness score s lon (t) varies between 1 and -1. D lon When (t) > A, the longitudinal robustness score is fixed at 1, and s lon (t) < B lon in which case the robustness score is fixed at -1. The longitudinal robustness score s lon (t) varies continuously over all possible values of the longitudinal headroom. The same idea applies to the lateral robustness score. As will be understood, this is merely an example, and the robustness score s(t) can be defined in various ways based on the headroom distance.

[0103] Normalization of the score is advantageous as it makes the rules easier to interpret and facilitates comparison of scores between different rules. However, it is not essential that the score be normalized in this way. The score can be defined over any range having an arbitrary failure threshold (not necessarily zero).

[0104] The robustness score s(t) of the safety distance rule as a whole is calculated by the third assessor node 308 as follows.

[0105] s(t) = min[s lat (t), s lon (t)] The rule passes when s(t) > 0 and fails when s(t) ≤ 0. When s = 0 (which implies that one of the longitudinal and lateral intervals is equal to its safety distance), the rule is "just" failed and represents the boundary between the pass and fail results (performance categories). Alternatively, s = 0 can be defined at the point when the host vehicle E just passes, which is an unimportant design choice, and thus the terms "pass threshold" and "fail threshold" are used interchangeably in this specification to refer to the subset of the parameter space where the robustness score y = 0.

[0106] The pass / fail result (or, more generally, the performance category) may be assigned for each time step of the scenario run based on the robustness score s(t) at that time, which is useful for an expert to interpret the result.

[0107] In addition to rating driving behavior against driving rules, the above rule framework may be used to evaluate other aspects of the autonomous vehicle stack contributing to performance, for example by defining rules for recognition errors. Recognition errors are determined based on a set of ground truth detection results, which are specific to the simulation and may be generated in a real-world driving scenario by manual annotation or by applying an offline recognition pipeline that uses offline detection and refinement techniques not available to the agent in real time to generate a high-quality recognition output, which may be referred to herein as a "pseudo ground truth" recognition output.

[0108] FIG. 8 shows an architecture for evaluating recognition errors. A triage tool 152 with a recognition oracle 1108 is used to extract and evaluate recognition errors for both real and simulated driving scenarios and outputs results that are rendered to a GUI 500 alongside the results from a test oracle 252. Note that the triage tool 152, which is referred to herein as a recognition triage tool, may be more generally used to extract driving data, including recognition data and driving performance data, useful for testing and improving the autonomous vehicle stack and presenting it to the user.

[0109] Regarding the actual sensor data 140 from the driving operation, the output of the online recognition stack 102 is passed to the triage tool 152, and a numerical "real-world" recognition error 1102 is determined based on the extracted ground truth 144 obtained by executing both the actual sensor data 140 and the online recognition output via the ground-truthing pipeline 400.

[0110] Similarly, in the case of a simulated driving operation where sensor data is simulated from zero and the recognition stack is applied to the simulated sensor data, a simulated recognition error 1104 is calculated by the triage tool 152 based on a comparison between the detection result from the recognition stack and the simulation ground truth. However, in the case of simulation, the ground truth can be obtained directly from the simulator 202.

[0111] If the simulator directly models the recognition error to simulate the output of the recognition stack, the difference between the simulated detection result and the simulation ground truth, i.e., the simulated recognition error 1110, is known and is passed directly to the recognition oracle 1108.

[0112] The recognition oracle 1108 receives a set of recognition rule definitions 1106, which may be defined via a user interface or described in a domain-specific language, as will be explained in more detail later. The recognition rule definitions 1106 may apply thresholds or rules that define recognition errors and their limits. The recognition oracle applies the rules defined for actual or simulated recognition errors obtained with respect to a driving scenario and identifies where the recognition errors violate the defined rules. These results are passed to the rendering component 1120, which renders visual indicators of the evaluated recognition rules for display on the graphical user interface 500. For clarity, the input to the test oracle is not shown in FIG. 8, but note that the test oracle 252 also depends on the ground truth scenario obtained from either the ground truth pipeline 400 or the simulator 202.

[0113] Next, further details of the framework for evaluating the recognition errors of the real-world driving stack against the extracted ground truth are described. As noted above, both the recognition errors and the driving rule analysis by the test oracle 252 can be incorporated into the real-world driving analysis tool, which will be explained in more detail below.

[0114] Not all errors have the same importance. For example, a 10 cm translational error in an agent 10 meters away from the ego vehicle is much more important than the same translational error in an agent 100 meters away. A simple solution to this problem would be to scale the error based on the distance from the ego vehicle. However, the relative importance of different perception errors, or the sensitivity of the ego vehicle's driving performance to different errors, depends on the use case of a given stack. For example, when designing a cruise control system for driving on a straight road, this should be sensitive to translational errors but not particularly sensitive to orientation errors. However, an AV dealing with the entrance of a roundabout should be very sensitive to orientation errors because the detected orientation of an agent is used to indicate whether the agent is about to exit the roundabout and thus whether the ego vehicle can safely enter the roundabout. Therefore, it is desirable to make it possible to set the sensitivity of the system to different perception errors according to each use case.

[0115] To define recognition errors, a domain-specific language is used. Using this, recognition rules can be created, for example, by defining tolerance limits for translational errors. This rule implements a settable set of safe error levels for different distances from the host vehicle. For example, when the vehicle is less than 10 meters away, the error in its position (i.e., the distance between the vehicle detection result and the refined pseudo ground truth detection result) can be defined to be 10 cm or less. When the agent is 100 meters away, the acceptable error may be defined up to a maximum of 50 cm. A lookup table can be used to define the rules according to any given use case. Based on these principles, more complex rules can be constructed. Rules may be defined such that the errors of other agents are completely ignored based on their positions relative to the host vehicle, for example, agents in the oncoming lane when the host lane is separated from oncoming traffic by a median strip. Traffic behind the host vehicle beyond a defined cut-off distance may also be ignored based on the rule definition.

[0116] By defining an error recognition specification that includes all the rules to be applied, a set of rules can be collectively applied to a given driving scenario. Typical recognition rules that can be included in the specification are translational errors in the longitudinal and lateral directions (measuring the average error of the detection results with respect to the ground truth in the longitudinal and lateral directions respectively), orientation errors (defining the minimum angle by which the detection results need to be rotated to match the corresponding ground truth), size errors (the error in each dimension of the detected bounding box, or the ratio of the intersection / union of the aligned ground truth and the detected box to obtain the volume difference), and thresholds related to these. Further rules may be based on the dynamics of the vehicle, which include errors in the speed and acceleration of the agent, and classification errors, such as penalty values in case of misclassifying a car as a pedestrian or a truck. The rules may also include false detections or missed detections, as well as detection delays.

[0117] Based on the defined recognition rules, a robustness score can be constructed. Effectively using this, when the detection results are within the specified thresholds of the rules, the system should be able to drive safely, and if not (for example, when there is too much noise), it can be said that something bad may happen that the host vehicle may not be able to handle, and this needs to be formally captured. For example, complex combinations of rules can be included to evaluate the detection results over time and to incorporate complex weather dependencies.

[0118] The error recognition framework is described in more detail in UK Patent Applications Nos. 2108182.3, 2108958.6, 2108952.9, and 2111765.0, which are hereby incorporated by reference in their entirety.

[0119] User Interface The above test framework, i.e., test oracle 252 and recognition triage tool 152, may be combined in a real-world driving analysis tool, in which, as shown in FIG. 1C, both recognition evaluation and driving evaluation are applied to the recognition ground truth extracted from the ground truth pipeline 400.

[0120] The results of the above-described rule-based analysis regarding the planning and recognition of the AV stack provide a numerical score that provides an indicator of the performance of the host vehicle in each scenario. This numerical data can be directly interpreted by an expert as described above to identify stack issues and improve the stack. Next, a user interface is described that provides not only visualization of the scenarios during testing but also the results of the rule evaluation to present the context of the scenarios to the user when identifying stack issues based on the test results. The graphical user interface, described in more detail below, provides a plot of the numerical scores based on applying rules defined on signals extracted from the scenarios, and also provides visualization of the scenario data so that the changes in the signals underlying the numerical scores are visible to the user within the scenario visualization. This is useful in multiple applications.

[0121] In one application example, the user interface may be used to visualize real-world scenarios, and this visualization may include an annotation of the recognition output (e.g., a bounding box) generated by the recognition component 102 and a representation of the scenario with the pseudo ground truth recognition output generated by, for example, a ground truth pipeline or manual annotation. This enables an expert user to easily identify where the recognition of the host vehicle significantly deviates from the "ground truth" recognition output. For example, if the user notices that the orientation of the bounding box representing an agent in front of the host vehicle is significantly different, this represents an orientation error. Using this, the recognition component 102 of the self-stack can visualize the error that caused the recognition error, and thus the recognition stack can be improved. Another possible application example is to identify where the ground truth recognition annotation is incorrect, and this information can be used to improve the ground truth method (whether manual or using an automated ground truth pipeline). The visualization may further display the raw sensor data alongside the recognition output, which may help an expert user identify whether the cause of the error is a failure in the host vehicle's recognition stack or a failure in the ground truth recognition. For example, if there is an orientation error between the bounding box output by the recognition stack 102 and the ground truth bounding box, and a set of camera images or lidar measurements is overlaid on the visual representation of the scenario, the expert user can easily identify the correct orientation of the agent within the scenario and thus identify which recognition output is the cause of the error. However, the user cannot easily determine the cause of the recognition error based solely on the numerical difference in the orientations of the two bounding boxes.

[0122] FIG. 9A shows an exemplary user interface for analyzing driving scenarios extracted from real-world data. In the example of FIG. 9A, a schematic overhead representation 1204 of a scene is shown based on point cloud data (e.g., derived from lidar, radar, or stereo or monocular depth imaging), and a corresponding camera frame 1224 is shown in an inset. Road layout information may be obtained from high-definition map data. The camera frame 1224 may be annotated with detection results. The UI may also show sensor data collected during driving, such as lidar, radar, or camera data. This is shown in FIG. 9B. The scene visualization 1204 is also overlaid with derived pseudo ground truth as well as annotations based on detection results from in-vehicle recognition components.

[0123] In the illustrated example, there are three vehicles, each annotated with a box. The solid box 1220 shows the pseudo ground truth of the agents in the scene, and the contour 1222 shows the unrefined detection results from the host vehicle's recognition stack 102. A visualization menu 1218 is shown, in which the user can select which of the sensor data, online and offline detection results to display. These may be toggled on and off as needed. The raw sensor data can be shown side by side with both the vehicle detection results and the ground truth detection results, enabling the user to identify or confirm specific errors in the vehicle detection results. The UI 500 enables the playback of the selected video, a timeline view is shown, and the user can select any point 1216 in the video to show a bird's-eye view snapshot and the camera frame corresponding to the selected point.

[0124] As described above, the recognition stack 102 can be evaluated by comparing the detection results with the refined pseudo ground truth 144. The recognition is evaluated in light of defined recognition rules 1106 that can depend on the use case of a particular AV stack. These rules specify different ranges of values for mismatches between the position, orientation, or scale of the vehicle detection results and those of the pseudo ground truth detection results. The rules can be defined in a domain-specific language, as described above. As shown in FIG. 9A, different recognition rule results are shown along the "top" recognition timeline 1206 of the driving scenario, aggregating the results of the recognition rules, and a period on the timeline is flagged when any of the recognition rules are violated. This can be expanded to present a set of individual recognition rule timelines 1210 for each defined rule.

[0125] The recognition error timeline may be "zoomed out" to present a longer period of the driving run. In the zoomed-out view, it may not be possible to display recognition errors with the same granularity as when zoomed in. In this case, the timeline may display an aggregation of recognition errors over a time window to provide a set of recognition errors grouped for the zoomed-out view.

[0126] The second driving assessment timeline 1208 shows how the pseudo ground truth data is assessed against the driving rules. The aggregated driving rules are shown in the top-level timeline 1208, which can be expanded into a set of individual timelines 1212 that display the performance against each defined driving rule. Each rule timeline can be further expanded as shown to display a graph of the numerical performance scores over time for a given rule. In this case, the pseudo ground truth detection results 144 are considered the actual driving behavior of the agents in the scene. To check whether the vehicle behaved safely for a given scenario, the behavior of the host vehicle can be evaluated against the defined driving rules, for example, based on a digital highway code.

[0127] In FIG. 9A, each driving rule timeline can be expanded to present a plot of the associated robustness score. The timeline for the "COMFORT_02" driving rule is shown in the expanded state, with an xy-style plot of the robustness score 1240 visible. A pass / fail threshold 1242 is shown, and the fail region 1244 on the timeline corresponds to the region 1246 of the plot where the score is below the threshold 1242. The user can "scrub" along the plot (e.g., near the pass / fail boundary) to visually map the time-varying plot 1240 to corresponding changes in the visualization 1204 of the drive. A marker (scrubber bar) 1248 is shown that extends vertically through all the rule timelines to indicate the current time step of the visualization 1204. The user can scrub the scenario by moving the marker 1248 horizontally along the rule timeline (to view the scenario visualized at different time steps). Color-coding can be applied to the xy plot to show the region above the pass / fail threshold in a different color than the region below. Further details of the scrubbing mechanism are described below with reference to FIGS. 9C and 9D.

[0128] In summary, both the recognition rule evaluation and the driving assessment are based on using the above-described offline recognition method to purify the detection results from real-world driving. In the case of driving assessment, the purified pseudo ground truth 144 is used to assess the behavior of the host vehicle against the driving rules. As shown in FIG. 1C, this can also be used to generate a simulated scenario for testing. In the case of recognition rule evaluation, the recognition triage tool 152 compares the recorded vehicle detection results with the offline purified detection results to quickly identify and triage possible recognition failures.

[0129] Also, a drive note may be displayed in the drive note timeline view 1214, which may include notable events flagged during driving. For example, the drive note includes when the vehicle brakes or changes direction, or when a human driver releases the AV stack.

[0130] An additional timeline may be displayed presenting user-defined metrics that help the user debug and triage potential issues. The user-defined metrics may be defined to identify errors or stack defects and to triage the errors when they occur. The user may define custom metrics according to the goals of a given AV stack. An exemplary user-defined metric may flag when messages do not arrive in order or when there is a message delay in recognition messages. This is useful for triage as it may be used to determine whether the plan was made due to a planner mistake or because the messages arrived late or out of order.

[0131] Figure 9B shows an example of a UI visualization 1204 where sensor data is displayed and a camera frame 1224 is shown in an inset view. Typically, sensor data from a single temporal snapshot is presented. However, each frame may present sensor data aggregated over multiple time steps to obtain a static scene map if high-resolution map data is not available. As shown on the left side, there are several visualization options 1218 for displaying or hiding data such as camera, radar, or lidar data collected in a real-world scenario, or online detection results from the recognition of the host vehicle itself. In this example, the online detection results from the vehicle are shown as a colored box 1222 superimposed on a gray box 1220 representing the ground truth refined detection results. There is an orientation error seen between the ground truth and the vehicle's detection results.

[0132] The refinement process implemented by the ground truth pipeline 400 is used to generate a pseudo ground truth 144 as the basis for multiple tools. The presented UI displays the results from the recognition triage tool 152, which enables using the test oracle 252 to assess the driving capabilities of ADAS in a single driving instance, detect defects, extract scenarios to reproduce issues (see Figure 1C), and send the identified issues to the developers to improve the stack.

[0133] Figure 9C shows an exemplary user interface configured to allow the user to zoom in on a sub-section of a scenario. Figure 9C shows a snapshot of the scenario, including the schematic representation 1204 and the camera frame 1224 shown in the inset view as described above with respect to Figure 9A. Also shown in Figure 9C are the set of recognition error timelines 1206, 1210, as well as the expandable driving assessment timeline 1208 and the drive note timeline 1214 described above.

[0134] In the example shown in FIG. 9C, the current snapshot of the driving scenario is shown by a scrubber bar 1230 that spans all the timeline views simultaneously. This may be used instead of the indication 1216 of the current point within the scenario on a single playback bar. The user can click on the scrubber bar 1230 to select the bar and move it to any point in the driving scenario. For example, the user may be interested in a particular error, such as a point within a section that is colored red or otherwise indicated as a section containing an error on the position error timeline, and the indication is determined based on the "ground truth" and the position error observed at that point between the detected results during the period corresponding to the indicated section. The user can click on the scrubber bar and drag the bar to the point of interest within the position error timeline. Alternatively, the user can click on a point on any of the timelines that the scrubber spans and place the scrubber at that point. This updates the overview view 1204 and the inset view 1224 to present the top-down overview view and the camera frame corresponding to the selected point, respectively. The user can then examine the overview view and the available camera data or other sensor data to confirm the position error and identify the possible reasons for the recognition error.

[0135] The "ruler" bar 1232 is presented above the recognition timeline 1206 and below the overview view. This includes a series of "notches" that indicate the time intervals of the driving scenario. For example, if a 10-second time interval is displayed in the timeline view, notches indicating 1-second intervals are presented. At some points, numerical indicators, such as "0 seconds", "10 seconds", etc., are also labeled.

[0136] The numerical scores associated with recognition error rules may be continuous (e.g., floating point) or discrete (e.g., integer). The number of detection misses (as a function of time) is an example of an integer score. The degree of deviation from the recognition ground truth (e.g., the offset of the position or orientation of the detection result from the corresponding ground truth) is an example of a floating point score. Color coding may be used on the recognition timeline to plot the change (or approximate change) in the score over time. For example, in the case of an integer score, a different color may be used for each integer value. Continuous scores may be plotted using a color gradient or may be "quantized" into discrete buckets shown using discrete color coding. Alternatively or additionally, the recognition error timeline may be "expandable" in the same way as the driving rules (as in FIG. 9A) to display and xy-plot the associated robustness scores.

[0137] A zoom slider 1234 is provided at the bottom of the user interface. The user can drag an indicator along the zoom slider to change the portion of the driving scenario presented on the timeline. Alternatively, the position of the indicator may be adjusted by clicking on a desired point on the slider bar to which the indicator will move. A percentage indicating the currently selected zoom level is presented. For example, if the length of the entire driving scenario is one minute, timelines 1206, 1208, 1214 present recognition errors, driving assessments, and drive notes, respectively, over the one-minute drive, the zoom slider indicates 100%, and the button is at the far left position. As the user slides the button until the zoom slider indicates 200%, the timeline is adjusted to present only the results corresponding to a 30-second snippet of the scenario.

[0138] The zoom may be configured to adjust the displayed portion of the timeline according to the position of the scrubber bar. For example, if the zoom is set to 200% in a one-minute scenario, the zoomed-in timeline presents a 30-second snippet centered on the selected point where the scrubber is located, i.e., a 15-second timeline is presented before and after the point indicated by the scrubber. Alternatively, the zoom may be applied based on a reference point such as the start point of the scenario. In this case, the zoomed-in snippet presented on the timeline after zooming always starts from the start point of the scenario. The granularity of the notches and numerical labels of the ruler bar 1232 may be adjusted according to the degree to which the timeline is zoomed in or out. For example, if the scenario is zoomed in from 30 seconds to present a 3-second snippet, before zooming, the numerical labels may be displayed at 10-second intervals and the notches at 1-second intervals, and after zooming, the numerical labels may be displayed at 1-second intervals and the notches at 100ms intervals. The visualization of the time steps of the timelines 1206, 1208, 1214 is "stretched" to correspond to the zoomed-in snippet. A higher level of detail may be displayed by the timeline in the zoomed-in view because a shorter time snippet can be represented by a larger area in the display of the timeline within the UI. Therefore, an error over a very short period of time within a longer scenario may be made visible in the timeline view only when zoomed in.

[0139] Other zoom inputs may be used to adjust the timeline to display shorter or longer snippets of the scenario. For example, if the user interface is implemented on a touch screen device, the user may apply a pinch gesture to apply zoom to the timeline. In other examples, the user may scroll the mouse scroll wheel forward and backward to change the zoom level.

[0140] When the timeline is zoomed in to show only a subset of the driving scenario, the timeline can be scrolled temporally to shift the display portion temporally, so that various parts of the scenario can be inspected by the user in the timeline view. The user can click and drag a scroll bar (not shown) at the bottom of the timeline view, or use, for example, the touch pad of a related device on which the UI is operating, to scroll.

[0141] The user can also select a snippet of the scenario, for example, for further analysis or as a basis for simulation, which is exported. FIG. 9D shows how a section of the driving scenario can be selected by the user. The user can click on the relevant point on the ruler bar 1232 with the cursor. This can be done at any zoom level. This sets the first boundary of the user selection range. The user drags the cursor along the timeline to expand the selection range up to the selected point in time. When zoomed in, by continuing to drag to the end of the displayed snippet of the scenario, this scrolls the timeline forward and allows the selection range to be further expanded. The user can stop dragging at any point, and the point where the user stops becomes the end boundary of the user selection range. The bar 1230 at the bottom of the user interface displays the temporal length of the selected snippet, and this value is updated as the user drags the cursor to expand or contract the selection range. The selected snippet 1238 is presented as a shaded section on the ruler bar. This section may be shown in a color different from the rest of the ruler bar. Some buttons 1236 are presented that provide user actions such as "Extract Trajectory Scenario" to extract the data corresponding to the selection range. This may be stored in a database of the extracted scenarios. This may be used for further analysis or as a basis for simulating similar scenarios. After making a selection, the user can zoom in or out, and the selection range 1238 on the ruler bar 1232 also expands and contracts along with the ruler and the timelines of recognition, driving assessment, and drive notes.

[0142] DSL can also be used to define a contract between the recognition stack and the planning stack of the system, based on a robustness score calculated with respect to defined rules. Figure 10 shows an exemplary graph of a robustness score with respect to a given error definition, such as a translational error. If the robustness score exceeds a defined threshold 1502, this indicates that recognition errors are within the expected performance and that the overall system should promise safe operation. If the robustness score drops below the threshold 1502 as shown in Figure 10, the error is "out of contract" because at that level of recognition error, the planner 106 cannot be expected to drive safely. This contract effectively becomes the requirements specification for the recognition system. This can be used to assign responsibility either to the recognition subsystem 102 or the planning subsystem 106. If an error is identified as being within contract when the vehicle is behaving incorrectly, this points to an issue with the planner rather than a recognition problem, and conversely, for bad behavior when recognition is out of contract, the cause is a recognition error.

[0143] Contract information can be displayed in the UI500 by annotating whether a recognition error is considered within contract or out of contract. This uses a mechanism that retrieves the contract specification from the DSL and automatically flags out-of-contract errors at the front end.

[0144] Further details of the above exemplary user interface for visualizing recognition errors and driving rules are described in UK Patent Applications Nos. 2108182.3, 2108958.6, 2108952.9, and 2111765.0.

[0145] In other application examples, as described in more detail herein, visualization may be used to enable an expert user to investigate errors in driving behavior generated based on the output of the vehicle's planner 106. As described above, the driving rules may be defined based on safety criteria that specify a safe distance between vehicles in various situations, and breaking these rules indicates the potential for safety risks. However, as described with respect to FIGS. 3A - 3B, the robustness score of the driving rules is not necessarily based on a measurable quantity that is easily interpretable. In the example given above, the robustness scores for the lateral and longitudinal distances are equal to the normalized difference between the actual distance and the threshold distance, or equal to 1 if the normalized difference exceeds a predetermined difference. This numerical value is useful for easily determining the severity of a rule violation, but is not easy to interpret from the perspective of real - world driving. By viewing these results within a visualization of a scenario showing how the actual host vehicle and other agents are traveling along the road, the user can confirm the relative speed of the vehicles and the distance between the vehicles throughout the scenario. The expert user can progress to the point in the scenario corresponding to, for example, the point of failure based on the robustness score, return to the scenario to identify the cause of the rule violation, and in some cases, determine whether it can be avoided in the future by adjusting the AV planner 106.

[0146] What was described above is a framework for evaluating agents within a scenario according to a set of predefined rules and metrics regarding agent behavior and / or recognition errors. As described above, for each of a set of abstract scenarios defined in a scenario description language and parameterized with a set of parameter values, the AV stack 100 may be evaluated in simulation by evaluating the performance of the ego agent over a number of simulated runs (or instances). A given instance of the AV stack is typically tested against a number of scenarios having various parameters within a "test suite". The test suite is defined by a set of parameter ranges of the parameters of the scenarios to be run and a set of rules (or "rule set") for evaluating the ego agent with respect to that test suite. When the test suite is run, a set of self-trajectories is generated, each self-trajectory comprising a time series of the ego vehicle state over the run, and a set of results is also output, which comprises the pass / fail results of the ego agent with respect to each rule of each scenario and a time series of numerical scores (robustness scores) of the ego agent with respect to each rule of each scenario that quantify the degree of success or failure over the entire run. These results may be aggregated over the test suite to obtain an overall picture of the performance of the ego vehicle over the set of scenario parameters being tested.

[0147] It may also be useful to directly compare two runs. In one example, a user testing an AV stack may want to compare the performance of their vehicle in two versions of the same abstract scenario where a few scenario parameters differ, to gain a detailed insight into how a given parameter value affects perception or the behavior of the ego agent in that abstract scenario. In another example, the same scenario with the same parameter values may be run with two different versions of the ego agent stack, for example, where the planner is changed for each instance of a given test suite. In this case, if the pass or fail of a given rule differs between a previous stack version and the current stack version, particularly in scenarios where the vehicle previously passed the rule but fails in the updated version (referred to herein as regression), it is beneficial to view these runs in a common visualization tool to identify at what point in the scenario the behavior of the two versions of the ego agent diverged, enabling the user to identify the cause of the regression.

[0148] Figure 4 shows a schematic block diagram of a computer system for rendering a driving visualization interface according to an embodiment of the present disclosure. Figure 4 shows that data for a first drive 402 and a second drive 404 are provided as input to a renderer. However, the visualization interface can also be implemented to display a single drive. As described above, the first drive and the second drive may be scenario instances where one or more scenario parameters of the two runs take different values, or the scenario parameters may be the same if the two runs belong to tests of two respective versions of the self-stack. Each drive includes a time series 416 of the ego vehicle state, which includes the spatial and motion coordinates of the ego vehicle at each time step of that drive, and each drive also includes a set of ego agent robustness scores 418 defined for each set of rules defined for the perception and / or behavior of the ego agent over that drive, as described above.

[0149] In addition to the driving data, a map defining the static road layout of the scenario is provided to the renderer. This includes representations of road lanes and road features such as junctions and roundabouts. Each scenario instance has an associated map. The map may be retrieved from a map database.

[0150] The rendering component 408 receives both the driving data and the map data 406 for both drives and renders a common visualization 412 showing snapshots of both drives overlaid on the same map, and a plot 414 of the robustness score for each rule of the rule set, with the robustness scores for both drives plotted on a common set of axes. By providing controls for the user to manually align the two drives, the visualization may present equivalent points in time for both drives to enable a direct visual comparison. Both the map visualization 412 and the robustness score plot 414 include a time axis with time markers 410 marking common time instances within both drives. The time markers 410 for the robustness score plot may be implemented in the form of a scrubber bar 1230 as described above with reference to FIG. 9C, or as points, circles, or other indicators along individual timelines as described below with reference to FIGS. 5 and 6.

[0151] By providing a user operation to move the time marker of the map visualization 412 so as to advance the visualization, the visualization is updated to show the state of the self-agent in each run at the moment where the marker has been moved along the time axis. Also, this operation can be used to update the time markers 410 of the plots of each rule and identify the robustness scores of the self-agent in each run at the selected moment indicated by a line within the robustness plot. The robustness plot 414 is shown in an expanded view in FIG. 4, with the numerical robustness scores along the y-axis and time along the x-axis. Another view of the robustness plot provides a binary indicator based on a pass / fail scheme, showing a timeline where sections of time where the robustness score has exceeded or fallen below the pass / fail threshold are identified, e.g., by color-coding, and the parts of the scenario runs where the self-agent has failed a given rule are shown in red on the timeline and the parts where it has followed the rule are shown in green. This is illustrated and described in more detail below with reference to FIGS. 5 and 6.

[0152] In the map visualization 412, the self-agent may be represented by a different color for each run. Although not shown in FIG. 4, scenario runs typically include one or more external agents moving within the same road layout, and these are also represented by different colors in order to visually distinguish the agents of the scenario for each run. The map visualization 412 and the robustness plot are provided within a common user interface display, which is described in more detail with reference to FIGS. 5 and 6.

[0153] FIG. 5 shows an exemplary driving visualization user interface in a single-driving view. Although two drives are available for display, only one drive is selected in the selection pane presented with a first check box 506 for the first drive and a second check box 504 for selecting and displaying the second drive. As shown in FIG. 5, since the second check box is deselected, only the first drive is displayed in the visualization 412, and a rule evaluation timeline 508 showing only the performance of the self-agent in the first drive is provided. In this example, the rule evaluation timeline is displayed in a view that is not expanded as described above, with a single timeline shown as a line, where the non-compliance of the self-agent with a given rule is shown as a red section on the timeline of that rule, and the time when the self-agent did not fail that rule is shown in green. Each rule timeline 508 is identified by the name of the rule (e.g., DR_01) and the title of the rule, such as "Collision", etc. A numeric indicator 512 is also shown that provides the numeric robustness score for the selected time step. An expand control 514 is provided, and when the user clicks on it, an expanded view of the rule timeline with a robustness plot of that rule can be displayed, as will be described in more detail with reference to FIG. 6.

[0154] The time steps within the run are indicated by time markers 410, which are shown as small circles at the start points of both the rule timeline 508 and the timeline provided at the bottom of the display. The markers for the overall timeline may be adjusted by the user clicking on an indicator and dragging it along the timeline to move the visualization to the selected point in the run. Since the display time markers 410 for the set of rule timelines and the time markers for the overall timeline refer to the same underlying data, user operations to adjust the time markers of one timeline also adjust the time markers of all the rule timelines 508. Since the robustness score for each rule is indexed by time, updating the time markers for each rule causes the robustness score displayed in the numerical indicator 512 to be updated to reflect the selected point in time. A search bar is provided where the user can enter a text filter to display only the rules related to a given keyword. For example, if the user enters "collision", the rule evaluation timeline for the rules that contain the word "collision" in their name or description can be returned.

[0155] In the example of FIG. 5, a set of robustness / rule evaluation plots / timelines 508 are displayed within the scenario visualization 412, which comprises a map visualization and an overall timeline. Within the map visualization, the agent in the first run is shown traveling along a highway lane. At the selected instant (in this example, the start time of the run), no other agents are within the view.

[0156] A set of controls 516 is provided for adjusting the map display. These can include controls for changing the orientation of the map, for example, according to a predefined default-oriented layout (for example, adjusting the map so that north corresponds to the upward direction within the visualization). An "agent tracking" control is shown on the left side, which, when clicked, enables the tracking of the self-agent and the vehicle of the self-agent is always presented at the center of the visualization during the playback of the scenario. The visualization of the field of view of each sensor of the host vehicle can be presented by enabling the sensor control. A button with additional controls may be provided to display additional options to the user, including, for example, measurement tools, debug mode, and various camera position views. The scale indicator indicates a reference distance for comparison with distances in the driving scenario.

[0157] In addition to the visualization 412, the user interface further includes a comparison table 502 that shows the corresponding rules of the scenario and the aggregated pass / fail results of the self-agent for each of the selected runs. As shown in FIG. 5, the comparison table 502 defines the instances to be compared at the top, which are identified by their respective indexes. The parameter values of the scenario for each instance are also displayed. In this example, the y speed is set to "1.6" in both runs. Other parameters can include weather conditions, lighting, etc. A table of rules is shown, with each rule displayed in one line along with a brief description of the function of that rule, and for each instance identified in each column of the table, an indicator of fail or pass of that rule is shown. As described above, the pass and fail conditions for each rule are specified in the rule definition. For example, a rule that specifies the minimum distance to other vehicles may fail if the self-agent is less than this minimum distance from other agents even for a short time. To enable the user to quickly identify rules for which the two runs differed, the pass and fail results are shown in green and red. If a given rule fails in the first run and passes in the other run, this rule can be inspected in a single-run view by selecting one of the runs and reviewing the rule evaluation timeline 508 corresponding to that given rule. The user may review each run in turn in the single-run view, which is done by selecting the checkbox corresponding to that run, as shown by the checkboxes 504, 506 in FIG. 5, such that the checkbox for the other run is deselected.

[0158] FIG. 6 shows a user interface of a driving comparison view in which two drives are compared with a common visualization 412. In this example, both the checkbox 506 for the first drive and the checkbox 504 for the second drive are selected, and both drives are displayed. Instead of selecting the checkboxes, the user may alternatively hover over the eye icon corresponding to the first drive to present a visualization of only that drive. In the map view, the ego agent 610a and the external agents 608a in the first drive are presented, and the ego agent 610b and the external agents 608b in the second drive are presented at respective different positions on the same road layout. The ego agent and the external agents may be shown in different colors within the UI. For example, the ego agent in the first drive may be shown in blue and the external agents may be shown in gray. The agents in the second drive may be shown in a different color such as orange. The time marker 410 indicates the progress of both drives along a common timeline. The frame number 602 is also shown, and controls 604a, 604b are provided, by which the user can select and click to move one frame at a time forward or backward in time between the frames of the drive, and each frame corresponds to one ego vehicle state in the time series of the ego vehicle states received by the renderer. On the right side of the timeline, the time (from the start of the scenario) is also shown. In this example, since the frames correspond to a fixed interval of 0.01 seconds, the current time of 4.950 seconds corresponds to frame 495.

[0159] In the driving comparison view, as described above for the single driving view, the rule evaluation timeline for the first drive is displayed. The time markers for each rule evaluation timeline are placed at relative points along that timeline, the same as the selected point in time on the main timeline of the overall visualization 412. FIG. 6 shows an expanded view of the "ALKS_03" rule that checks for inter-vehicle interruption response, which includes a robustness plot 414a with both robustness scores for both the first and second drives plotted on the same axis. When the user moves a time marker along the expanded timeline of a given rule, the time markers for all other rules as well as the timeline of the overall visualization are updated to the corresponding time step selected by the user.

[0160] Another rule, "ALKS_05 - Stable Lateral Position", is shown in an expanded view with a robustness plot 414b. This plot has the robustness scores for both the first and second drives plotted on it. The time marker has an associated line parallel to the y-axis at the selected time, which intersects the plots for each drive. The label shows the value of the robustness score for each drive at the selected time. In this example, for the first drive, the robustness score is 0.24 at the selected time, and for the second drive, the robustness score is 2. The scale of the plot is indicated by the labels 12 and -12 on the y-axis. The robustness plot for the first drive is shown as a line that almost overlaps the x-axis since the robustness score is relatively close to zero over the duration of that drive. In contrast, the plot for the second drive starts at a high value, then goes below zero and stays near zero for the rest of that drive. For this exemplary rule, the ego-agent passed the rule over the duration of the first drive, but for the second drive, the robustness score dropped below zero. The UI may be configured to display the corresponding part of the plot in red. In this example, since the first drive is the drive for which the rule evaluation timeline is being displayed, the rule evaluation timeline 508 for the ALKS_05 rule is displayed in green over the entire drive.

[0161] The user can click on the time marker 410 and drag it along the timeline (referred to herein as "scrubbing") to select and visualize another time within the duration of the two runs. When the user moves the time marker, the visualization of the agents within the road layout is updated to reflect the state of each agent at the selected time within the run. The time marker of the rule timeline and the robustness value 512 displayed next to the rule timeline are also updated to reflect the selected time. The scrubbing mechanism can be applied to the run comparison view (such as in FIG. 6) and can also be applied to the single run view (described above). The scrubbing mechanism is described in more detail above with reference to FIG. 9A. Although FIG. 9A shows a single run view, this description equally applies to the run comparison view (the user can scrub along the rule timelines stacked vertically for multiple runs).

[0162] Accordingly, the user can compare the behavior of their agent in different runs as the run progresses, understand the reasons for deviations in behavior, and provide information for future tests. For example, if the parameters are the same between two runs, but a comparison is made between two different versions of the self-stack during the runs, and a given rule such as a stable lateral position rule fails in the updated version of the stack, the user can review the position of their vehicle within the run corresponding to the updated stack, identify the nature of the error in the lateral position of their vehicle, and attempt to identify its cause. As described above, since the agents of the second run are displayed in a different color from the agents of the first run, the two runs can be easily distinguished within the visualization. Alternatively, some other means of visually identifying the agents of each run, such as a visual effect like lower opacity of the agents, or a label on or near the agents of a given run, may be used.

[0163] The driving comparison interface may be used to evaluate changes added to the stack. For example, if the AV planner is updated to change the behavior of the host vehicle when leaving a junction, the previous version of the stack (before this change was implemented) and the current version can be compared based on the corresponding driving of the scenario where the host vehicle leaves, to identify changes in the behavior of the host vehicle with the same scenario parameters. The new self-stack may be evaluated in scenarios with different scenario parameters, and these can be compared with the new version of the stack to identify how different scenario parameters affect the host vehicle's decision to leave after the changes to the planner have been implemented.

[0164] As described above, in a typical use case of the driving comparison interface, the user may identify from the comparison table 502 that a given rule passed in one of the compared drives but failed in the other. Assuming both drives are selected, the user can deselect the checkbox associated with the drive that passed the given rule and identify the approximate point in time when the failure occurred based on the rule timeline of the drive that failed. Next, the user can move the time marker close to the point in time when the self-agent failed the rule and view the playback of the host vehicle's behavior near the time of the failure. Next, the checkbox corresponding to the drive that passed can be reselected to compare with the behavior of the host vehicle in the other drive, and the scenario can be played back to show how the behavior of the self-agent differed between the two drives.

[0165] The above description relates to the use of a comparison tool for behavior rules, but the user interface can also include recognition rules as described above, and the recognition of the vehicle (e.g., either the detection result simulated using a recognition error model or the actual detection result generated in real time by an autonomous vehicle) is evaluated against the ground truth. For example, if a change is made to the recognition system and the same scenario is re-run, and the ego agent fails a collision rule by colliding with a vehicle ahead, but passed this rule in a previous version of the stack, the user may replay the two runs within the same visualization and determine that the failure to detect the vehicle ahead within sufficient time caused the collision, and the user can review the latest changes to the recognition stack to identify how the regression occurred.

[0166] In the use case of regression comparison, one or more test suites, each defining a set of scenarios, are run on two different versions of the self-stack. Typically, each test suite contains a large number of scenario instances (e.g., tens of thousands or more), and most of the results are the same between stack versions. It is infeasible for the user to manually review these results to identify the rules that differed between the two versions. Instead, an aggregation may be performed that runs the two test suites and identifies and reports only the rules of each scenario that resulted in different outcomes between the two stack versions.

[0167] An interface showing the results of such aggregation is shown in FIG. 7. This regression report interface may be provided as an additional display within a test tool that also includes the driving comparison visualization described above. The report includes columns of test results, which identify the test suite comparison 702, and as described above, each test suite defines a set of scenario parameters and a rule set. For each test suite of this comparison, a pair of IDs 704 is displayed, and each ID 704 identifies a separate run of that test suite. For a given test suite regression comparison, each rule 706 where a regression was found is displayed in a second column, and a summary 708 of the improvements and regressions for that rule is displayed in a third column. Details of the individual regressions and improvements are shown in the lines below, with the scenario name 710 identified in the second column, and the results 712 identifying the two runs of that scenario that were different and the results of each run. In the illustrated example, the run of the current version of the stack is shown first, followed by the run of the previous version of the stack in parentheses. In this example, for a given "STAY_IN_LANE_JCTN" rule, five improvements and one regression were found. The first scenario listed is a regression, the run corresponding to the current stack version is a failure, and the previous stack version was a pass. The remaining five results indicate that the current version of the stack passed when the previous version of the stack passed. A "Compare" link 714 is provided for each improvement and regression, and clicking on this link by the user will direct the user to the driving comparison interface described above with reference to FIGS. 5-6, along with the corresponding instance IDs that identify the first driving data and the second driving data rendered in the driving comparison interface.

[0168] As described above, the evaluation results may be stored in a result database, and the result database may be accessed by the graphical user interface described above to display a plot of the numerical performance scores.

[0169] References to components, functions, modules, etc. in this specification refer to the functional components of a computer system that can be implemented at the hardware level in various ways. The computer system may comprise execution hardware configured to execute the method / algorithm steps disclosed herein and / or to implement a model trained using this technology. The term execution hardware encompasses any form / combination of hardware configured to execute the relevant method / algorithm steps. The execution hardware may take the form of one or more programmable or non-programmable processors, or a combination of programmable and non-programmable hardware may be used. Examples of suitable programmable processors include general-purpose processors based on instruction set architectures such as CPUs, GPU / accelerator processors, etc. Such general-purpose processors typically execute computer-readable instructions held in memory coupled to or embedded within the processor and perform the relevant steps in accordance with those instructions. Other forms of programmable processors include field-programmable gate arrays (FPGAs) having a programmable circuit configuration through circuit description code. Examples of non-programmable processors include application-specific integrated circuits (ASICs). The code, instructions, etc. may be stored, if necessary, on a temporary or non-temporary medium (examples of the latter include solid-state, magnetic, and optical storage devices, etc.). The subsystems 102-108 of the runtime stack in FIG. 1A may be implemented on a vehicle using a programmable processor, a dedicated processor, or a combination of both, or in a non-vehicle computer system in the context of testing, etc. Similarly, the components in FIGS. 2A, 2B, 3B, 4, and 8 may also be implemented using programmable hardware and / or dedicated hardware.

Claims

1. A computer system for rendering a graphical user interface for visualizing the driving of a driving scenario in which an agent progresses along a road layout, at least one input configured to receive a map of the road layout of the driving scenario and driving data of the driving of the driving scenario, the driving data comprising a sequence of time-stamped agent states, and a time-varying numerical score quantifying the performance of the agent, calculated by applying each rule of a set of driving evaluation rules to the driving, the at least one input; a rendering component configured to generate rendering data, the rendering data for a graphical user interface comprising for each rule of the driving evaluation rules, a plot of the time-varying numerical score, a marker indicating a selected time index on the time axis of the plot, the marker being movable along the time axis via user input in the graphical user interface to change the selected time index, the marker; a scenario visualization comprising a visualization of the road layout overlaid with a visualization of the agent of the driving at the selected time index, moving the marker along the time axis causing the rendering component to update the scenario visualization when the selected time index is changed, the rendering component for displaying the scenario visualization, the agent visualization being a visualization of the field of view of the sensors of the agent, the computer system.

2. The input is further configured to receive second driving data of a second driving of the driving scenario, the second driving data including a second sequence of agent states with timestamps and a second time-varying numerical score quantifying the performance of the agent, calculated by applying each rule of a set of driving evaluation rules to the driving, and the rendering component is further configured to generate rendering data, the rendering data including, for each rule of the set of driving evaluation rules, a plot of the second time-varying numerical score on a graphical user interface, wherein the time-varying numerical score and the second time-varying numerical score are plotted with respect to a common set of axes having at least a common time axis, and the marker indicates a selected time index on the common time axis, the plot, and a second agent visualization of the second driving at the selected time index, the scenario visualization being for displaying the second agent visualization overlaid thereon, the computer system of claim 1. **Claim 3** The computer system of claim 1, wherein the time-varying numerical score is calculated by applying one or more rules to a time-varying signal extracted from the driving data, and a change in the signal is visible in the scenario visualization. **Claim 4** The rendering component, in response to a deselection input in the graphical user interface indicating one of the first driving and the second driving, for each driving rule, removes the plot of the time-varying numerical score of the deselected driving from the common set of axes, and is configured to remove the agent visualization of the deselected driving from the single visualization of the road layout, and a user can switch from a driving comparison view regarding both the first driving and the second driving to a single driving view regarding only one of the first driving and the second driving, the computer system of claim 2. **Claim 5** The graphical user interface further includes a comparison table having an entry for each rule of the set of driving evaluation rules, the entry including an aggregated performance result for the rule in the first driving and an aggregated performance result for the rule in the second driving. The computer system according to claim 2.

6. The computer system according to claim 5, wherein the entry for each rule further comprises an explanation of the rule.

7. The rendering component is configured to, in response to an expansion input in the graphical user interface, hide the plot of the time-varying numerical score of each rule and display a timeline view including an indication of the pass / fail result of the rule over time. The computer system according to claim 1.

8. The rendering component is configured to display, in the graphical user interface, the numerical score at the selected time index for each rule of the set of driving evaluation rules. The computer system according to claim 1.

9. The driving evaluation rules include recognition rules, the scenario visualization includes a set of recognition outputs generated by a recognition component of the host vehicle, and the recognition rules define recognition errors and their limits. The computer system according to claim 1.

10. The computer system according to claim 9, wherein the scenario visualization includes sensor data superimposed on the visualization of the road layout.

11. The scenario visualization includes a scenario timeline having scenario time markers, and moving the markers along the scenario timeline causes the rendering component to update the respective time markers of each plot of the time-varying numerical score when the selected time index is changed. The computer system according to claim 1.

12. The scenario timeline includes a frame index corresponding to the selected time index, and a set of controls for moving forward or backward by incrementing or decrementing the frame index respectively, the computer system according to claim 11.

13. The driving scenario is a simulated driving scenario in which a simulated self-agent travels on a simulated road layout, and the driving data is received from a simulator, the computer system according to claim 1.

14. The driving scenario is a real-world driving scenario in which the self-agent travels on a real-world road layout, and the driving data is calculated based on data generated on the self-agent during the driving, the computer system according to claim 1.

15. The plot of the time-varying numerical score includes an xy plot of the time-varying numerical score, the computer system according to claim 1.

16. The time-varying numerical score is plotted using color coding, the computer system according to claim 1.

17. A method for visualizing the driving of a driving scenario in which a self-agent travels on a road layout, receiving a map of the road layout of the driving scenario and driving data of the driving of the driving scenario, the driving data including a sequence of time-stamped self-agent states, and a time-varying numerical score quantifying the performance of the self-agent calculated by applying the driving evaluation rules to the driving for each rule in a set of driving evaluation rules, the receiving; generating rendering data, the rendering data including, on a graphical user interface, for each rule among the driving evaluation rules, a plot of the time-varying numerical score, and a marker indicating a selected time index on the time axis of the plot, the marker being movable along the time axis via user input in the graphical user interface to change the selected time index, the marker. A scenario visualization comprising a visualization of the road layout overlaid with an agent visualization of the driving at the selected time index, wherein moving the marker along the time axis causes the rendering component to update the scenario visualization when the selected time index is changed, and the generating, for displaying the scenario visualization. The method wherein the agent visualization is a visualization of the field of view of the sensors of the self-agent. A computer program comprising executable instructions for programming a computer system to implement the method or system functions according to any one of claims 1 to 17. ​

Citation Information

Patent Citations

  • How to make driving and autonomous driving easier

    JP2019512824A

  • A framework for evaluating predicted trajectories in traffic prediction for autonomous vehicles

    JP2019527813A

  • Automatic operation program evaluation system, and automatic operation program evaluation method

    JP2020123259A