Performance Testing of Robot System
Through the perceived statistical performance model (PSPM), the perception system of autonomous driving vehicles is simulated, and the problems of difficulty in modeling perception errors and low testing efficiency in the prior art are solved, and safety and reliability testing are achieved in complex environments.
Patent Information
- Application Number
- CN202080059358.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-06
- Filing Date
- 2020-08-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-08-21
AI Technical Summary
The prior art is difficult to ensure that the perception system of autonomous vehicles reaches the same safety level as that of human drivers in various complex environments without increasing a large amount of real-world test mileage, and the simulation method has difficulties in modeling computational efficiency and perception errors.
Perceived statistical performance model (PSPM) is used to simulate the perceived ground reality in the scene, construct probability uncertainty distribution, simulate reality perception output, and combine online error estimation to train and test the perception system of autonomous driving vehicles.
Effectively simulate the real error of the perception system, reduces the real-world mileage required for testing, improves testing efficiency and accuracy, and ensures the safety and reliability of the perception system in various environments.
Smart Images

Figure CN114270369B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to performance testing of autonomous vehicles and other robotic systems. Performance testing is crucial to ensure that such systems can operate at a guaranteed level of safety. Background Art
[0002] It is estimated that in order for an autonomous vehicle (AV) to reach a safety level comparable to that of a human driver, at most 1 error must occur per 10^7 autonomous driving decisions across the entire Operational Design Domain (ODD) of the AV.
[0003] Given the complexity of the AV and the ODD itself, this poses a huge challenge. A self-driving system is an extremely complex assembly of interdependent and interacting software and hardware components (each component is vulnerable to limitations or errors). Several components use neural networks for object detection, type classification, action prediction, and other critical tasks. The system needs to operate safely within the ODD. In this context, the ODD describes all possible driving scenarios that the AV may encounter, and thus, it has infinite possibilities in itself, with variables including road topologies, users, appearances, lighting, weather, behaviors, seasons, speeds, randomness, and intentional behaviors.
[0004] The industry standard method for safety testing is based on actual driving test mileage. An autonomous vehicle fleet is driven by test drivers, and when test driver intervention is required, the decision is characterized as unsafe. Once an instance of test driver intervention occurs in a specific real-world driving scenario, the circumstances of that driving scenario can be explored to isolate the factors that led to the unsafe behavior of the AV and take appropriate mitigation actions. Summary of the Invention
[0005] Simulation has been used for safety testing, but it is only useful if the simulation scenario is realistic enough (if the AV planner makes an unsafe decision in a completely unrealistic simulation scenario, then it is far less useful in the context of safety testing than an instance of an unsafe behavior in a realistic scenario).
[0006] One approach is to run simulations based on real-world scenarios that require testing driver intervention. Sensor outputs from the AV are collected, and the sensor outputs can be used to reconstruct driving scenarios in the simulator that require testing driver intervention. Variables of the scenarios can be "fuzzified" at the planning level so as to test variations of real-world scenarios that are still realistic. In this way, more information about the causes of unsafe behaviors can be obtained, analyzed, and used to improve prediction and planning models. However, significant problems arise because as the number of errors per decision decreases, the number of test miles that need to be driven to find a sufficient number of instances of unsafe behavior increases. A typical AV planner may make approximately 1 decision every two seconds on average. At an average speed of 20 miles per hour, this is equivalent to making approximately 90 decisions per mile driven. This in turn implies that there is less than one error per 10^5 miles driven in order to match the human safety level. Robust safety testing requires multiple tests to fully test the AV across its ODD. As the perception stack evolves, this situation will deteriorate further because with each change to the perception stack, more test miles are required. For these reasons, this approach is simply not feasible when testing at a safety level close to that of humans.
[0007] There are other problems with existing simulation methods.
[0008] One approach is simulation at the planning level, but this cannot fully account for the effects of perception errors. Many factors can affect perception errors, such as weather, lighting, distance to another vehicle or the speed of another vehicle, occlusion, etc.
[0009] An alternative would be a full "fidelity" simulation, in which the entire hardware and software stack of the simulated AV is replicated. However, this is a huge challenge in itself. The AV perception pipeline typically consists of multiple perception components that work together to interpret the sensor outputs of the AV.
[0010] One problem is that some perception components (such as Convolutional Neural Networks (CNNs)) are particularly sensitive to the quality of simulated data. Although high-quality simulated image data can be generated, CNNs in perception are extremely sensitive even to minor deviations from real data. Therefore, these would require extremely high-quality simulated image data that covers all possible conditions that the AV might encounter in the real world (e.g., different combinations of simulated weather conditions, lighting conditions, etc.) - otherwise their behavior in simulated scenarios will not fully reflect their behavior in the real world.
[0011] The second problem is that certain types of sensor data are particularly difficult to model (simulate). As a result, even perception systems that are not particularly sensitive to the quality of the input data will give poor results. For example, radar belongs to the category of sensor data that is extremely difficult to simulate. This is because the physical characteristics of radar are inherently difficult to model.
[0012] The third primary problem is the issue of computational efficiency. Based on current hardware constraints, it is estimated that it may be possible to achieve realistic simulations in real time (even if other problems can be overcome).
[0013] The present disclosure provides a distinct approach for simulation-based safety testing using a model referred to herein as a "Perceptual Statistical Performance Model" (PSPM). The core problem addressed in the present disclosure is to simulate realistic perception outputs - i.e., perception outputs with realistic errors - in a manner that is not only more robust than realistic simulations but also significantly more efficient.
[0014] The PSPM models perceptual errors in terms of a probabilistic uncertainty distribution based on a robust statistical analysis of the actual perceptual outputs computed by one or more perceptual components being modeled. The unique aspect of the PSPM is that, given a perceptual ground truth (i.e., the "perfect" perceptual output that would be computed by a perfect but unrealistic perceptual component), the PSPM provides a probabilistic uncertainty distribution that represents the realistic perceptual components that can be provided by the perceptual components it is modeling. For example, given a ground truth 3D bounding box, the PSPM modeling of a simulated 3D bounding box detector will provide an uncertainty distribution representing the realistic 3D object detection output. Even when the perception system is deterministic, it can be effectively modeled as stochastic to account for the epistemic uncertainty of the many hidden variables it depends on in practice.
[0015] Of course, the perceptual ground truth will not be available at runtime in the real world AV (which is why complex perceptual components are needed to reliably interpret imperfect sensor outputs). However, the perceptual ground truth can be directly derived from the simulated scenarios running in the simulator. For example, in the case of a 3D simulation of a driving scenario with an ego vehicle (the simulated AV being tested) in the presence of external actors, the ground truth 3D bounding box can be directly computed from the simulated scenario of the external actors based on their size and pose (position and orientation) relative to the ego vehicle. The PSPM can then be used to derive realistic 3D bounding object detection outputs from these ground truths, which can in turn be processed by the remaining AV stack as if they were at runtime.
[0016] The situation addressed in this document is as follows: The perception system or subsystem being modeled itself provides a perception error estimate, such as a covariance estimate for the perception output. These can be referred to as "online" perception error estimates to distinguish them from the modeling by the PSPM itself. Such online error estimates are important because they can be fed, for example, into higher-level perception components (such as filters or fusion components that fuse perception outputs in a way that involves their relative error levels) as well as probabilistic prediction / planning. This disclosure recognizes that the online error estimates themselves may be subject to error, and it is useful to be able to model this error in a representative way.
[0017] A first aspect of this document provides a computer system for testing and / or training a runtime stack of a robotic system, the computer system comprising:
[0018] A simulator configured to run a simulation scenario in which a simulated agent interacts with one or more external objects;
[0019] A planner of the runtime stack configured to make autonomous decisions for each simulation scenario based on a time series of perception outputs computed for the simulation scenario; and a controller of the runtime stack configured to generate a series of control signals to cause the simulated agent to execute the autonomous decisions as the simulation scenario progresses;
[0020] wherein the computer system is configured to compute each perception output by:
[0021] Computing perception ground truth based on the current state of the simulation scenario;
[0022] Applying a perception statistical performance model (PSPM) to the perception ground truth to determine a probabilistic perception uncertainty distribution; and
[0023] Sampling a perception output from the probabilistic perception uncertainty distribution;
[0024] wherein the PSPM is used to model a perception slice of the runtime stack and is configured to determine the probabilistic perception uncertainty distribution based on a set of parameters learned from a set of actual perception outputs generated using the perception slice to be modeled;
[0025] wherein the perception slice includes an online error estimator, and the computer system is configured to use the PSPM to obtain a predicted online error estimate of the perception output in response to the perception ground truth.
[0026] In an embodiment, the predicted online error estimate can be sampled from the probabilistic perception uncertainty distribution.
[0027] In practice, this means determining the 'covariance of the covariance', or more generally, the statistical distribution of errors in online perception error estimation under different perceived ground truth conditions, such that it is possible to sample realistic online perception error estimates in a statistically useful way.
[0028] The PSPM can take the form of a function approximator that receives the perceived ground truth t and outputs the parameters of a probabilistic perception uncertainty distribution from which to sample the perception output and the predicted online error estimate.
[0029] The PSPM can have a neural network architecture.
[0030] The PSPM can be applied to the perceived ground truth and one or more confounding factors associated with the simulated scenario, each of which is a variable of the PSPM, the value of the variable characterizing the physical conditions applicable to the simulated scenario, and the probabilistic perception uncertainty distribution depending on the variable, and the predicted online error estimate depending on the confounding factor.
[0031] One or more confounding factors can include one or more of the following confounding factors, where the confounding factor at least partially determines the probabilistic uncertainty distribution from which the perception output is sampled:
[0032] The occlusion level of at least one of the external objects;
[0033] One or more lighting conditions;
[0034] An indication of the time of day;
[0035] One or more weather conditions;
[0036] An indication of the season;
[0037] The physical characteristics of at least one of the external objects;
[0038] Sensor conditions, e.g., the position of at least one of the external objects in the field of view of the subject's sensor;
[0039] The number or density of external objects;
[0040] The distance between two external objects;
[0041] The truncation level of at least one of the external objects;
[0042] The type of at least one of the objects, and
[0043] An indication of whether at least one of the external objects corresponds to any external object from an earlier time in the simulated scenario.
[0044] The PSPM may include a time-dependent model such that the sampled perception outputs sampled at the predicted online error estimates depend on at least one of the following: the earlier one of the perception outputs sampled at the previous moment, and the earlier one of the perceived ground truths calculated for the previous moment.
[0045] The computer system may include a scenario evaluation component configured to evaluate the behavior of an external agent in each simulated scenario by applying a set of predetermined rules.
[0046] At least some of the predetermined rules may be related to safety, and the scenario evaluation component may be configured to evaluate the safety of the agent's behavior in each simulated scenario.
[0047] The computer system may be configured to record details of each simulated scenario in a test database, where the details include the decisions made by the planner, the perception outputs on which those decisions are based, and the behavior of the simulated agent when executing those decisions.
[0048] Sampling from the probabilistic perception uncertainty distribution may be non-uniform and biased towards lower-probability perception outputs.
[0049] 1 The computer system may include a scenario blurring component configured to generate at least one blurred scenario for running in a simulator by blurring at least one existing scenario.
[0050] To model false negative detections, the probabilistic perception uncertainty distribution may provide the probability of a visible object among successfully detected objects, which is used to determine whether to provide an object detection output for the object. The object is visible when it is within the sensor field of view of the agent in the simulated scenario, and thus the detection of the visible object is not guaranteed.
[0051] Ray tracing may be used to calculate the perceived ground truth for one or more external objects.
[0052] At least one of the external objects may be a moving actor, and the computer system includes a prediction stack of a runtime stack configured to predict the behavior of the external actor based on perception outputs, and a planner configured to make autonomous decisions according to the predicted behavior.
[0053] A second aspect of the present disclosure provides a computer-implemented method for performance testing a runtime stack of a robotic system, the method including:
[0054] Run a simulation scenario in a simulator, where a simulation subject interacts with one or more external objects, where a planner in a runtime stack makes autonomous decisions for the simulation scenario based on a time series of perceptual outputs calculated for the simulation scenario, and a controller in the runtime stack generates a series of control signals to cause the simulation subject to execute the autonomous decisions as the simulation scenario progresses;
[0055] Where each perceptual output is calculated by:
[0056] Calculate perceptual ground truth based on the current state of the simulation scenario;
[0057] Apply a Perceptual Statistical Performance Model (PSPM) to the perceptual ground truth to determine a probabilistic perceptual uncertainty distribution; and
[0058] Sample a perceptual output from the probabilistic perceptual uncertainty distribution;
[0059] Where the PSPM is used to model a perceptual slice of the runtime stack and determines the probabilistic perceptual uncertainty distribution based on a set of parameters learned from a set of actual perceptual outputs generated using the perceptual slice to be modeled;
[0060] Where the perceptual slice includes an online error estimator and the PSPM is used to obtain an online error estimate of the prediction of the perceptual output in response to the perceptual ground truth.
[0061] A third aspect of the present disclosure provides a computer-implemented method for training a Perceptual Statistical Performance Model (PSPM), where the PSPM models the uncertainty in perceptual outputs calculated by a perceptual slice of a runtime stack of a robotic system, the method comprising:
[0062] Apply the perceptual slice to a plurality of training sensor outputs, thereby calculating a training perceptual output for each sensor output, where each training sensor output is associated with a perceptual ground truth, where the perceptual slice includes an online error estimator that provides an online perceptual error estimate for each training perceptual output;
[0063] Use the training perceptual outputs and their online perceptual error estimates to train the PSPM, where the trained PSPM provides a probabilistic perceptual uncertainty distribution in the form of p(e, E|t), where p(e, E|t) represents the probability that the perceptual slice calculates a specific perceptual output e and a specific online perceptual error estimate E given a perceptual ground truth t.
[0064] Assuming that e and E are independent of each other but each depends on the perceptual ground truth t, the probability distribution p(e, E|t) can include separate component distributions p(e|t) and p(E|t).
[0065] A fourth aspect of the present disclosure provides a Perceptual Statistical Performance Model (PSPM) implemented in a computer system, the PSPM being configured to model a perceptual slice of a runtime stack of a robotic system and being configured to:
[0066] Receive a computed perceptual ground truth t;
[0067] Determine a probabilistic perceptual uncertainty distribution of the form p(e, E|t) from the perceptual ground truth t based on a set of learned parameters, where p(e, E|t) represents the probability of computing a particular perceptual output e and a particular online perceptual error estimate E for the perceptual slice given the perceptual ground truth t, and the probabilistic perceptual uncertainty distribution is defined over a range of possible perceptual outputs and online perceptual error estimates, and the parameters are learned from a set of actual perceptual outputs generated using the perceptual slice to be modeled.
[0068] Another aspect of the present disclosure provides a computer program for programming one or more computers to implement any method or function of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] To better understand the present disclosure and to illustrate how embodiments of the present disclosure may be implemented, reference is made to the accompanying drawings, in which:
[0070] Figure 1 A schematic block diagram of a runtime stack of an autonomous vehicle is shown;
[0071] Figure 2 An example of a real-world driving scenario is shown;
[0072] Figure 3 A test pipeline using photorealistic simulation is shown;
[0073] Figure 4 An alternative PSPM-based test pipeline according to the present disclosure is shown;
[0074] Figure 5 How perceptual performance is affected by confounding factors is shown;
[0075] Figure 6 A high-level overview of certain principles of PSPM-based safety testing is provided;
[0076] Figure 7 A perceptual error data set that can be used to train the PSPM is shown;
[0077] Figure 7A Shows the application to Figure 7 The results of a trained PSPM for the perceptual error data set;
[0078] Figure 8 Shows an engineering pipeline incorporating PSPM;
[0079] Figure 9 Shows an example of a perception stack;
[0080] Figures 9A - 9C Shows different ways in which a perception stack can be modeled using one or more PSPMs Figure 9 ;
[0081] Figure 10 Provides a schematic overview of factors that can lead to perception uncertainty;
[0082] Figure 11 Shows an example of simulated image data to which certain forms of perception components are highly sensitive;
[0083] Figure 12 and Figure 13 Shows an aerial view and a driver's view of a roundabout scenario;
[0084] Figure 14 Schematically depicts stereo imaging geometry;
[0085] Figure 15 Shows an example time series of the additive error of the position component;
[0086] Figure 16 Shows a lag plot of the position error;
[0087] Figure 17 Shows a graphical representation of a time-dependent position error model;
[0088] Figure 18 Shows an example binning scheme for the confounding factors azimuth and distance;
[0089] Figure 19 Shows a lag plot of the position error increment;
[0090] Figure 20 Shows a histogram of the position error increments of the X, Y, and Z components;
[0091] Figure 21 Shows the PDF-fitted position error increments of the X, Y, and Z components;
[0092] Figure 22 Shows an example mean of the error increment distribution in training data (for a single object based on tracking);
[0093] Figure 23 Shows a time series plot of the true perception error and the simulated error;
[0094] Figure 24 Shows a hysteresis plot for real perception error and simulation error;
[0095] Figure 25 Graphically depicts the relative importance of certain confounding factors in a particular (left to right) for a target association state as determined by MultiSURF Relief analysis;
[0096] Figure 26 Graphically depicts the relative importance of confounding factors (left to right) for a target transition as determined by MultiSURF Relief analysis;
[0097] Figure 27 Shows an example node in a neural network;
[0098] Figure 28 Shows a high-level overview of a convolutional neural network architecture;
[0099] Figure 29 Shows that PSPM is implemented as a neural network during training and inference;
[0100] Figure 30 Shows a neural network PSPM having one or more confounding factor inputs at the input layer;
[0101] Figure 31 Shows an example of a time-dependent neural network architecture;
[0102] Figure 32 Shows a "set-to-set" PSPM implemented as a neural network;
[0103] Figure 33A Schematically depicts the spatial encoding of the perceptual output that contributes to processing in a convolutional neural network (CNN);
[0104] Figure 33B Schematically depicts the training phase of a CNN PSPM;
[0105] Figure 33C Schematically depicts a trained CNN PSPM during inference;
[0106] Figure 33D Shows how a CNN PSPM can be architected to encode a perceptual output distribution in an output tensor from which a realistic perceptual output can be sampled; and
[0107] Figure 34 Shows how a PSPM can be configured to model a perceptual slice including an online error estimation component. Detailed Description
[0108] 1. Overview
[0109] The terms "PSPM" and "PRISM" are used interchangeably in the following description.
[0110] When making a safety case for an autonomous vehicle, it is impractical to perform all the required tests in the real world. However, building simulations with such high fidelity that the vehicle's perception system behaves equivalently on real and simulated data is an unsolved problem. The approach, referred to herein as "PRISM", solves this problem by constructing a surrogate model of the perception system, which includes both sensors and (one or more) perception components that interpret the sensor data captured by the sensors. PRISM is the distribution of credible perception outputs given some low-fidelity scene representation (perceptual ground truth).
[0111] Building on the above, to ensure that driverless technology can be proven safe, it is necessary to test driverless technology in a large number of scenarios. Conducting such tests with real vehicles is both expensive and time-consuming. In natural scenarios, most of the miles driven will be uneventful - in 2016 in the UK, 136,621 people were injured and 1,792 people died in road accidents, and all motor vehicles drove 323.7 billion miles, with an accident occurring only once every 2.4 million miles driven. Simulation must be part of the testing strategy for driverless technology. Simulation miles are much cheaper than real miles, and it is easier and safer to increase the number of hazards per mile in simulation than in the real world.
[0112] One way to generate realistic perception outputs is via a high-fidelity simulation of the world, including sensor measurements. In this approach, realistic sensor readings are generated, which are fed into the car's software in place of real sensor readings, e.g., a realistic twin of the real world rendered as an image for perceptual input. This rendering is as Figure 11 shown. The car's software outputs control signals for the car's actuators, which are fed into a physics simulation. New sensor readings are generated based on the output of the physics simulation, thus closing the loop. This approach requires accurate models to be generated for tasks ranging from challenging to unsolved:
[0113] · It is possible to simulate road surfaces, vehicle dynamics, and other physical characteristics with current technology, but they are not well understood.
[0114] · It is possible to simulate GPS, IMU, and wheel-encoding, but it is important to get their error statistics correct.
[0115] · Visual appearance, camera lenses, and image sensors are reasonably well understood, but high-fidelity rendering is slow.
[0116] · Lidar modelling is similar to camera modelling, but has different material reflection properties. The scanning characteristics of lidar are an additional challenge.
[0117] · It is difficult to accurately model radar echoes with current technology because it is difficult to model the relevant material properties and the detailed dependencies on shape and multiple reflections.
[0118] · Worst of all, state-of-the-art neural networks for visual object detection are extremely sensitive to detailed image statistics, and constructing synthetic images that elicit the same network behavior as equivalent real images is an unsolved problem.
[0119] Inaccurate models of the above sensors will affect the output of the perception module in the simulation, resulting in potentially different ego behavior. Such differences in behavior limit the usefulness of these simulations in evaluating real-world performance. Additionally, running many miles of photorealistic simulation required to verify the safety behavior of autonomous vehicles is expensive. This is because rendering photorealistic scenes is a slow, computationally intensive task that requires a GPU. High-fidelity simulation is both difficult and expensive, and conclusions from testing using high-fidelity simulation are unlikely to generalize to the real world.
[0120] Figure 1 A data flow diagram of an autonomous vehicle stack 100 through decomposition is shown. The perception system 102 receives sensor readings from the world and outputs a scene representation. The planning and prediction system (represented by reference numerals 104 and 106 respectively) uses this scene representation and plans a trajectory through the scene. The control system 108 outputs control signals to the world that will cause the vehicle to follow the trajectory.
[0121] The perception system 102, the planning and prediction systems 104, 106, and the control system 108 communicate with each other using well-defined interfaces. The perception system 102 uses raw sensor data and processes the raw sensor data into a more abstract scene representation. This representation includes dynamic object pose, extent, motion, and detection confidence. The planning and prediction systems predict the possible trajectories of other agents in the scene and plan a safe, legal, and comfortable path through the scene. The control system uses the desired trajectory from the planning and prediction systems and outputs control signals for the actuators.
[0122] In many cases, especially in the case of the interface between perception and planning, these internal interfaces are easier to simulate than sensor readings. These interfaces can be used for a second type of simulation called low-fidelity simulation. One can simulate only those aspects of the world that are necessary to reconstruct the abstract scene representation used by the planner and provide this abstract scene representation directly to the planner, thus taking the perception system out of the loop. While this avoids some of the burdens of high-fidelity simulation, it presents new challenges: replicating the behavior of the perception system. It is well known that perception systems are not perfect, and errors in the perception system can affect the prediction, planning, and control systems in a meaningful way. Since the results tested in simulation should generalize to the real world, it must be possible to simulate realistic perception outputs.
[0123] A method is proposed to simulate realistic perception outputs using a model called PRISM. PRISM is a distribution of plausible perception outputs given some low-fidelity scene representations. The mathematical framework guiding the creation of PRISM is outlined, a prototype is created, and the modelling choices are documented. Doing so demonstrates that the modelling approach is reasonable.
[0124] Broadly speaking, in high-fidelity simulation, the simulator replaces the world and the entire vehicle stack is treated as a black box. In low-fidelity simulation, the world and the perception system 102 are replaced (see Figure 4 and the description below).
[0125] Figure 1FIG. 0 shows a highly schematic block diagram of a runtime stack 100 for an autonomous vehicle (AV). The runtime stack 100 is shown as including a perception stack 102, a prediction stack 104, a planner 106, and a controller 108.
[0126] The perception stack 102 receives sensor outputs from an on-board sensor system 110 of the AV.
[0127] The on-board sensor system 110 can take different forms, but typically includes various sensors such as image capture devices (cameras / optical sensors), LiDAR and / or RADAR units, satellite positioning sensors (GPS, etc.), motion sensors (accelerometers, gyroscopes, etc.), etc. The various sensors together provide rich sensor data from which detailed information about the surrounding environment and the state of the AV and any external actors (vehicles, pedestrians, cyclists, etc.) within that environment can be extracted.
[0128] Thus, the sensor outputs typically include sensor data from multiple sensor modalities, such as stereo images from one or more stereo optical sensors, LiDAR, RADAR, etc.
[0129] The perception stack 102 includes a plurality of perception components that cooperate to interpret the sensor outputs and thereby provide a perception output to the prediction stack 104.
[0130] The perception output from the perception stack 102 is used by the prediction stack 104 to predict the future behavior of external actors.
[0131] The predictions computed by the prediction stack 104 are provided to the planner 106, which uses the predictions to make autonomous driving decisions to be executed by the AV in a way that takes into account the predicted behavior of external actors.
[0132] The controller 108 executes the decisions made by the planner 106 by providing appropriate control signals to an on-board motor 112 of the AV. In particular, the planner 106 plans the manoeuvres to be taken by the AV, and the controller 108 generates control signals to execute these manoeuvres.
[0133] Figure 2Shows an example of certain perception components that can form part of the perception stack 102, namely a 3D object detector 204 and a Kalman filter 206.
[0134] The depth estimator 202 captures stereo image pairs and applies stereo imaging (such as Semi-Global Matching) to extract depth estimates from the stereo image pairs. Each depth estimate is in the form of a depth map that assigns depth values to the pixels of one image of the stereo image pair derived from the depth map (the other image is used as a reference). The depth estimator 202 includes a stereo pair of optical sensors and a stereo processing component (hardware and / or software), which are not shown separately in the figure. In accordance with the terminology used herein, both the optical sensors and the stereo processing component of the depth estimator 202 are considered part of the vehicle sensor system 110 (not the perception stack 102). The depth map is a form of sensor output provided to the perception stack 102.
[0135] The 3D object detector 204 receives the depth estimates and uses them to estimate the poses of external actors near the AV (ego vehicle). Two such external actors are shown in the form of two other vehicles. The pose representation in this context is a 6D pose, i.e., (x, y, z, pitch, roll, yaw), representing the position and orientation of each external actor in 3D space.
[0136] For illustrative purposes, Figure 2 is highly simplified. For example, the 3D object detector can be formed by multiple collaborative perception components that jointly operate on the sensor outputs of multiple sensor modalities. The application of PSPM in a more complex stack will be described later. Currently, for illustrative purposes of some of the core principles of PSPM, consider a simplified example where it is assumed that the 3D object detector operates on the sensor output of a single modality (stereo depth).
[0137] In real-world scenarios, multiple physical conditions can affect the performance of the perception stack 102. As noted, physical conditions that are considered variables with respect to a particular PSPM are referred to as "confounders". This allows for the consideration of variable physical conditions that are statistically relevant to a particular perception slice.
[0138] As described above, one simulation approach would be to attempt to realistically simulate not only Figure 1 the entire runtime stack 100 but also the vehicle sensor system 110 and the motor 112. This is illustrated in Figure 3 In this scenario, the challenge is the simulation of sensor data: certain types of sensor data (e.g., RADAR) are inherently difficult to simulate well, while other types of sensor data (image data) are relatively easier to simulate, and certain perception components (such as CNNs) are highly sensitive to even minor deviations from actual sensor data. Another challenge is the simulation of the large amount of computational resources required for sensors and for running complex perception components (such as CNNs).
[0139] For example, for the Figure 3 arrangement, it would be necessary to simulate extremely high-quality depth maps and to run the 3D object detector 204 on those simulated depth maps. During simulation, even extremely minor deviations in the simulated depth maps (compared to the real depth maps provided by the stereo depth estimator 202) can significantly affect the performance of the 3D object detector 204.
[0140] Figure 4 A high-level schematic overview of PSPM-based simulation is provided. In this case, a "headless" simulator setup is used, where there is no need to create simulated sensor data (e.g., simulated images, depth maps, lidar, and / or radar measurements, etc.), and there is no need to apply the perception stack 102 (or at least not to fully apply the perception stack 102 - see below). Instead, one or more PSPMs are used to effectively compute realistic perception outputs, which are in turn fed into the high-level components of the runtime stack 100 and processed as they would be at runtime.
[0141] The PSPM is said to model a "perception slice", which can be all or part of the perception stack 102. A perception slice can be a single perception component or multiple cooperating perception components.
[0142] Mathematically, a perception slice can be represented as a function F, where:
[0143] e = F(x),
[0144] e is the perceptual output of the perceptual slice, and x is the set of sensor outputs operating on the perceptual component.
[0145] At runtime on the AV, e is determined by applying F to x, which in turn is given by the sensors.
[0146] The PSPM mapped to the confounder space C can be represented as a function p, where:
[0147] p(e|t,c) represents the probabilistic uncertainty distribution that provides the probability of F computing the perceptual output e given the perceptual ground truth t and a set of one or more confounders c (i.e., given a particular set of possible real-world conditions represented by the point c in the confounder space C).
[0148] For example, for 2D bounding box detection:
[0149] · F can be a CNN
[0150] · x can be an RGB image
[0151] · t can be the ground truth bounding box that can be directly computed from the simulation using ray tracing (without simulating x or applying F), or a set of multiple such boundaries framed for multiple ground truth objects (set-to-set method)
[0152] · c can be distance and / or weather, etc.
[0153] In Figure 3 the example, e represents one or more 6D pose vectors computed by the 3D object detector 204, and x represents the depth map from which e is derived provided by the stereo depth estimator 202.
[0154] Figure 5 Illustrates how to use the PSPM to simulate Figure 3 the realistic perceptual output of the scene. In this case, the perceptual slice is the 3D object detector 204.
[0155] "Realistic" in the current context means a perceptual output that is more realistic than the perceptual ground truth.
[0156] Provided is a PSPM 500 which essentially models the perception slice 204 as a noisy “channel” affected by both the characteristics of the stereo depth estimator 202 and the physical environment. The physical environment is characterized by a set of confounding factors c which, in this example, are: lighting, weather, occlusion, and distance to each external actor.
[0157] To apply the PSPM 500, the perception ground truth can be directly computed from the simulated scenario under consideration. For example, in a simulated scenario where the simulated AV (ego vehicle) has many external actors in its vicinity, the 6D pose ground truth can be determined by directly computing the 6D pose of each external actor in the ego vehicle's reference frame.
[0158] The PSPM 500 then uses the computed ground truth t to compute the distribution p(e|t,c). Continuing the above example, this will provide, for each simulated external actor, the probability that the actual 3D object detector 204 computes the perception output e [estimated 3D pose of the external actor] given the perception ground truth t [“actual” 6D pose] in a real-world scenario characterized by the same confounding factors c.
[0159] Once p(e|t,c) has been computed, it can be used to run multiple simulations on a series of realistic perception outputs (PSPM samples) obtained by sampling p(e|t,c). Depending on the realistic means of p(e|t,c) with a high enough probability—note that it may be very desirable to test relatively low-probability perception outputs (outliers) if they are still realistic. The extent of testing outliers will depend on the safety level that the AV needs to meet.
[0160] In Figure 5 three realistic perception outputs e1, e2, e3 are shown by way of example. These are sampled from p(e|t,c).
[0161] One way is to sample the perception outputs from p(e|t,c) in a way that favors the most likely perception outputs, for example, using Monte Carlo sampling. Broadly speaking, this will test a larger number of the most likely perception outputs and fewer of the less likely ones.
[0162] However, while this may be useful in some contexts, in other contexts it may be more useful to deliberately test a larger number of "outliers" (i.e., less likely but still realistic perceived outputs), because the potential outliers are more likely to cause or contribute to unsafe behavior. That is, p(e|t,c) can be sampled in a way that deliberately biases towards outliers, to deliberately make a particular scenario more "challenging" or "interesting" as it progresses. This can be achieved by transforming the distribution of the PSPM and sampling from the transformed distribution.
[0163] Figure 6 An overview of the process for constructing the PSPM is provided. A large number of real sensor outputs x are collected and annotated with the perception ground truth t. This is exactly the same process as generating the training data for the perception components of the perception stack 102 (represented by box 602) - and a first subset of the annotated sensor outputs is used for this proposal. A trained perception slice 204 is shown, which (in the real world) is run and used to construct the PSPM 500, which will model the perception slice 204 during simulation.
[0164] Continue Figure 3 Continuing with the example, the real sensor output would be a depth map, and the ground truth would be the ground truth 6D pose of any object captured in the depth map. This annotated data is used not only to train the 3D object detector 204 (the perception slice in this example), but also to construct the PSPM 500, which models the 3D object detector 204 during simulation.
[0165] Box 604 represents the PSPM construction (training), and a second subset of the annotated sensor outputs is used for this purpose. Each sensor output is additionally annotated with a set of confounding factors c, which characterize the physical conditions under which it was captured in the real world. A large number of sensor outputs are required for each set of confounding factors c that the PSPM needs to be able to adapt to. For full "Level 4" autonomous driving, this means capturing and annotating sensor outputs across the entire ODD.
[0166] The PSPM 500 can take the form of a parametric distribution:
[0167] Dist(t,c;θ)
[0168] where t and c are the variables on which the distribution depends, and θ is the set of learning parameters.
[0169] The parameter θ is learned as follows:
[0170] 1) Apply the trained perception slice 204 to each sensor output x to compute the corresponding perceived output e;
[0171] 2) For each sensed output e, determine the deviation (error) Δ between e and the corresponding ground truth t.
[0172] 3) Each error Δ is associated with the ground truth t and a set of confounding factors c associated with the corresponding sensor output x.
[0173] 4) Considering the associated ground truth and the variable confounding factors c, adjust the parameter θ to fit the distribution to the error Δ.
[0174] As will be apparent, various known forms of parametric distributions / models can be applied in this context. Thus, they will not be elaborated further.
[0175] More generally, the training set for PSPM training consists of sensed ground truth (from manual, automatic, or semi-automatic annotation) and the corresponding actual sensed outputs generated by the sensed slices 204 to be modeled. The purpose of training is to learn the mapping between the sensed ground truth and the distribution of the sensed outputs that captures the statistics of the actual sensed outputs. Thus, for a given ground truth t, the sensed outputs sampled from the distribution p(e|t) will be statistically similar to the actual sensed outputs used for training.
[0176] As an example, the sensed slice 204 can be modeled as having zero-mean Gaussian noise. However, it should be emphasized that the present disclosure is not limited to this aspect. The PSPM can well take the form of more complex non-Gaussian models. As an example, the PSPM can take the form of a hidden Markov model, which will allow explicit modeling of the temporal dependencies between the sensed outputs at different times.
[0177] In the Gaussian case, for example, the PSPM 500 can be characterized as:
[0178] e = t + ε
[0179] ε ∼ N(0, Σ(c)),
[0180] where N(0, Σ(c)) represents a Gaussian distribution with zero mean and covariance Σ(c), and the covariance Σ(c) varies according to the confounding factors c. During simulation, then sample Gaussian noise and add it to the sensed ground truth. This will depend on the variance of the Gaussian and thus on the confounding factors applicable to the simulation scenario.
[0181] Exemplary PSPM Error Dataset
[0182] Figure 7An example of a raw error plot of a two-dimensional prediction space is shown—for example, each point can correspond to (x,y) coordinates, which can be estimated by a 2D object detector. Each prediction e is represented by a circle, and each ground truth t is represented by a star. Each error Δ is represented by a line segment between the corresponding prediction e and the corresponding ground truth t (the longer the line segment, the greater the error).
[0183] To construct a PSPM, the aim is to take into account the variable confounding factor c in order to accurately capture Figure 7 the error relationships between the data points (in this context, the data points are the errors Δ) in a way that adjusts the parameter distribution probabilistically.
[0184] Figure 7A Shows the result of a trained PSPM applied to Figure 7 the error data set.
[0185] Selected Confounding Factors
[0186] The decision on which confounding factors to incorporate is observation-driven: when it can be seen that a particular physical property / condition has a material impact on the perceptual uncertainty, this may trigger its introduction as a confounding variable into the applicable PSPM. Only statistically relevant confounding factors should be introduced.
[0187] One way to handle confounding factors is to partition the error data set according to the confounding factors and train a separate model for each partition of the data set. As a very simple example, two confounding factors could be "lighting" and "weather", each of which could take a binary "good / poor" value. In this case, the data set could be divided into four subsets with (lighting, weather) = (good, good), (good, bad), (bad, good), and (bad, bad) respectively, and four separate models could be trained for each subset. In this case, the PSPM consists of four models, where the confounding factor variable c = (lighting, weather) is used as an index to determine the selection of the model.
[0188] Engineering Pipeline Architecture
[0189] Figure 8 A highly schematic overview of the engineering pipeline incorporating the PSPM is shown. The entire pipeline covers everything from data collection, annotation, and extraction; training of the perception components; PSPM characterisation and simulation-based testing.
[0190] A large number of sensor outputs (such as, stereo images, depth maps, lidar measurements, and radar measurements) are collected using a fleet of vehicles each equipped with a sensor system 110 of the type described above. These are collected in the types of environments and driving scenarios that will need to be able to be processed in AV practice, for example, in the target urban areas where AV deployment is desired. The collection vehicles themselves can be AVs or manually driven vehicles equipped with similar sensor systems.
[0191] For the purpose of annotating the captured sensor outputs with ground truth, a ground truth pipeline 802 is provided. This includes annotating the sensor outputs with the type of perceptual ground truth described above. The sensor outputs annotated with perceptual ground truth are stored in an annotated ground truth database 804. Further details are described below.
[0192] In addition, the sensor outputs captured by the fleet are also used to extract driving scenarios, which can then be recreated in a simulator. A high-level structured scenario description language is used to capture the driving scenarios and store them in a scenario database 806.
[0193] The sensor outputs captured from the fleet are not the only source of information from which driving scenarios can be extracted. In addition, CCTV (closed circuit television) data 800 is used as a basis for scenario extraction, typically CCTV data captured in an urban environment, such as an urban environment that shows challenging urban driving scenarios, such as complex roundabouts. This provides a rich source of challenging driving scenarios and thus an excellent basis for safety testing. A collection of backend perception components 808 is used to process the CCTV data 800 to assist in the process of extracting driving scenarios from the CCTV data 800, and the driving scenarios are also stored in the scenario database 806 in a scenario description language format.
[0194] Further details of the scenario description language and the process of extracting scenarios from CCTV data and other data can be found in UK patent application No. 1816852.6, which is incorporated herein by reference in its entirety.
[0195] The driving scenarios captured in the scenario description language format are high-level descriptions of driving scenarios. The driving scenarios have both a static layout (such as road layout (lanes, markings, etc.), buildings, road infrastructure, etc.) and dynamic elements. Figure 8In the pipeline, the static layout is captured in the scene description as a pointer to an HD (High Definition) map stored in the map database 826. The HD map itself can be derived from annotated sensor outputs collected by an AV fleet and / or exported from CCTV.
[0196] Dynamic elements include, for example, the positions and movements of actors (e.g., vehicles, pedestrians, cyclists, etc.) in the static layout and are captured in a scene description language.
[0197] Running Simulation
[0198] The test suite orchestration component 810 uses the captured driving scenarios to formulate test instance specifications 812, which can then be run in a 3D simulator 814 as 3D multibody simulations. The purpose of these simulations is to enable the derivation of accurate perception ground truth, and then apply PSPM. Therefore, they contain a sufficient level of 3D geometric detail to enable the derivation of, for example, ground truth 3D bounding boxes (dimensions, 6D poses of external actors in the ego-vehicle reference frame), odometry, and ego-localization outputs, etc. However, they are not photorealistic simulations because that level of detail is not required. They also do not attempt to simulate conditions such as rain, lighting, etc., as they are modeled as confounding factors c.
[0199] To provide greater scene variation, a scene "fuzzer" 820 is provided that can fuzz the scene in the above sense. Fuzzing the scene means changing one or more variables of the scene to create a new scene that is still realistic.
[0200] Typically, this will involve fuzzing dynamic elements into the static scene, e.g., changing the movements of external actors, removing or adding external actors, etc.
[0201] However, the static layout may also be fuzzed, e.g., to change the curvature of the road, change the positions of static objects, change road / lane markings, etc.
[0202] Figure 8 The training block 602 of... is shown as being able to access the annotated ground truth data database 804, which, as described above, is used to train the perception slice 204 of the runtime stack 100.
[0203] As described above and as Figure 8 shown, the perception slice 204 is not necessarily the entire perception stack 102. In this example, the perception stack 102 is "sliced" before the set of final fusion components (filters), and the set of final fusion components (filters) cooperate to fuse the perception outputs from below the perception stack 102. These form part of one or more remaining prediction slices 205, and the one or more remaining prediction slices 205 do not use PSPM modeling but are applied to PSPM samples. The output of the final (unmodeled) prediction slice 205 is directly fed into the prediction stack 104.
[0204] The PSPM is shown as being stored in the PSPM database 820.
[0205] Running Simulation
[0206] The PSPM sampling orchestration component 816 uses 3D multibody simulation in the 3D simulator 814 to derive ground truth, which in turn forms the input for one or more PSPMs for PSPM modeling of the perception slice 104 and provides PSPM samples 818 for each simulation. The PSPM samples 818 are fed into the remainder of the runtime stack 100 (i.e., in this example, the PSPM samples 818 are fed into the final set of filters 205) and are used as the basis for planning and prediction, and finally cause the controller 108 to generate control signals, which are provided to the set of simulated AV motors.
[0207] The simulated motors are not shown in Figure 8 but are shown in Figure 4 and are denoted by the reference numeral 412. As Figure 4 shown, the 3D multibody simulation in the 3D simulator is driven in part by the simulated motors. These determine how the (in this case, simulated) body moves within the static layout (i.e., they determine the changes in the state of the body (which can be referred to herein as the simulated robot state)). In turn, the behavior of the body can also affect the behavior of the simulated external actors in response to the movement of the AV. As the 3D simulation progresses, new perception ground truth continues to be derived and fed into the PSPM 500 iteratively until the simulation is complete.
[0208] Each completed simulation is recorded as a set of test results stored in the test database 822.
[0209] Note that the same scenario can be run multiple times, but it may not produce the same result. This is due to the probabilistic nature of the PSPM: each time the scenario is run, a different PSPM sample may be obtained from the PSPM. Thus, a large amount of information is obtained by running the same simulation scenario on multiple occasions and observing, for example, the degree to which the simulated agent behaves differently in each instance of the scenario (a large difference in the agent's behavior indicates that the effect of the perception error is significant) or the proportion of scenario instances in which the agent behaves unsafely. If the same scenario is run a large number of times and the agent behaves safely and very similarly in each scenario, it indicates that the planner 106 is able to correctly plan under uncertain conditions in that scenario.
[0210] Test Oracle
[0211] The driving scenarios used as the basis for the simulation are typically based on real-world scenarios or blurred real-world scenarios. This ensures that real-world scenarios are being tested. However, note that these are usually driving scenarios that do not involve any actual autonomous vehicles, i.e., at least in most cases, the driving scenarios being tested are derived from real-life instances of human driving. Thus, it is not known which scenarios may lead to failure.
[0212] To this end, a scenario evaluation component 824 (referred to herein as the "test oracle") is provided, which has the function of evaluating whether the behavior of the simulated AV in the scenario is acceptable once the simulation is complete. The output of the test oracle 824 can include a simple binary (yes / no) output to flag whether the AV behaves safely, or it can be a more complex output. For example, it can include a risk score.
[0213] To do this, the test oracle 824 applies a set of predefined rules, which this document may refer to as the "Digital Highway Code (DHC)". In essence, the rules that define safe driving behavior are hard-coded. If the scenario is completed without violating those rules, the AV is considered to have passed. However, if any of those rules are violated, the AV is considered to have failed and is marked as an instance of unsafe behavior that requires further testing and analysis. Those rules are encoded at the ontological level such that they can be applied to the ontological description of the scenario. The concept of ontology is well-known in the robotics field and, in this context, aims to represent the driving scenario and the behavior of the simulated AV in that scenario at the same level of abstraction, such that the DHC rules can be applied by the test oracle 824. The results of the analysis can quantify how well the agent performs relative to the DHC, e.g., the degree of rule violation (e.g., the rule can specify maintaining a certain distance from a cyclist at all times, and the result can indicate the degree of violation of that rule and the circumstances of the violation).
[0214] Instances of unsafe behavior can also be marked as instances that require "disengagement". For example, this could be the case where a failover mechanism within the runtime stack 100 is activated to prevent a crash or some other critical failure (as in that scenario in the real world).
[0215] The present technique is not limited to detecting unsafe behavior. Behavior can be evaluated based on other metrics, such as comfort, progress, etc.
[0216] Example Perception Stack
[0217] Figure 9A schematic block diagram showing a portion of an example perception stack is shown. A 3D object detector is shown and denoted by reference numeral 204, and the 3D object detector is further shown as including a 2D object detector 902, a 2D tracker filter 904, a size estimation component 906, an orientation estimation component 908, a depth segmentation component 910, and a template fitting component 912. This represents an example architecture of the 3D object detector 204 mentioned above and shown in earlier figures.
[0218] The 2D object detector receives one image of each captured stereo image pair (the right image R in this example) and applies 2D object detection to the image. The output is the 2D bounding boxes of each object detected in the image. This provides the 2D (x, y) positions of each object in the image plane and the bounding boxes indicating the size of the projection of the object onto the image plane. The 2D tracking filter 904 receives the 2D bounding box output and applies filtering to them to refine the 2D bounding box estimates. For example, based on an object behaviour model, the filtering can take into account previous 2D detected bounding boxes and the expected behaviour of the detected objects. The filtered 2D bounding boxes and the image data of the original images contained therein are then used for many different purposes. The 2D object detector 902 can take the form of a trained CNN.
[0219] The depth segmentation component 910 receives the filtered 2D bounding boxes and also receives the depth map extracted from the original stereo image pair by the stereo estimator 202. It uses the filtered 2D boxes to separate the depth points belonging to each object within the depth map. This is a form of depth segmentation.
[0220] The size estimation component 906 also receives the filtered 2D bounding boxes and uses them to estimate the 3D sizes of each detected object based on the image data of the right image contained within the 2D bounding boxes.
[0221] The orientation estimation component 908 similarly receives the filtered 2D bounding boxes and uses them to determine the 3D orientations of each detected object using the image data of the right image contained within the applied 2D bounding boxes. The size estimation component 906 and the orientation estimation component 908 can take the form of trained CNNs.
[0222] For each detected object, the 3D template fitting component 912 receives the separated depth points of the object from the depth segmentation component 910, the 3D dimensions of the object from the dimension estimation component 906, and the 3D orientation of the detected object from the orientation component 908. The 3D template fitting component 902 uses those three pieces of information to fit a template, in the form of a 3D bounding box, to the depth points belonging to the object. Both the 3D dimensions and the 3D orientation of the 3D bounding box are known from the dimension and orientation estimation components 906, 908 respectively, and the points to which the bounding box must be fitted are also known. So, this is just a case of finding the best 3D position of the 3D bounding box. Once this is done for each object, the 3D dimensions and 6D pose (3D position and 3D orientation) of each detected object at a given moment are known.
[0223] Shows the input and output from the 3D template fitting component 912 to the final filter 205. Additionally, the final filter 205 is shown as having inputs that receive sensing outputs from lidar and radar respectively. The lidar and radar sensing components are shown and are denoted by reference numerals 914 and 916 respectively. Each of these provides a sensing output that can be fused with the sensing output (such as 6D pose) from the 3D object detector 204. This fusion occurs in the final filter 205, and the output of the final filter is shown as being connected to the input of the prediction stack 104. For example, this could be a filtered (refined) 6D pose that takes into account all of these stereo, lidar, and radar measurements. It can also take into account the expected object behavior in 3D space as captured in the expected behavior model of the 3D object.
[0224] Slices of the Perception Stack
[0225] Figure 9A Shows an example of how to "slice" the Figure 9 perception stack, namely modeled as a PSPM. The perception stack 102 is considered to be sliced after the final perception component modeled by the PSPM, and the perception output of this perception component can be referred to as the "final output" for the PSPM. The distribution of the PSPM will be defined over those final outputs, i.e., the e in p(e|t,c) corresponds to those final outputs of the component after which the perception stack 102 is sliced. Based on the influence of all perception components and sensors on the uncertainty in the final output e, the PSPM models all perception components and the sensors that provide input to this component (either directly or indirectly) (and is said to be "wrapped" in the PSPM).
[0226] In this case, a single PSPM is provided for each sensor modality, i.e., one for stereo imaging, a second for LiDAR, and a third for RADAR. The three PSPMs are denoted by reference numerals 500a, 500b, and 500c, respectively. To construct the first PSPM 500a, the perception stack 102 is sliced after the 3D template fitting component 912. Thus, the distribution of the first PSPM 500a is defined over the perception output of the template fitting component 912. All perception components and sensors fed into the 3D template fitting component 912 are wrapped within the first PSPM 500a. The second and third PSPMs 914, 916 are sliced after the LiDAR and RADAR perception components 914, 916, respectively.
[0227] The final filter 205 is not modeled as a PSPM but is applied to the PSPM samples obtained from the three PSPMs 500a, 500b, and 500c during testing.
[0228] Figure 9B A second example slice is shown where all three sensor modalities are modeled using a single PSPM 500d. In this case, the distribution p(e|t,c) is defined over all three sensor modalities, i.e., e = (e stereo ,e lidar e lidar ). Thus, each PSPM sample will include the perception output for all three sensor modalities. In this example, the final filter is still not modeled as a PSPM and will be applied during testing to the sampled PSPM obtained using the single PSPM 500d.
[0229] Figure 9C A third example slice is shown where all three sensor modalities along with the final filter 205 are modeled as a single PSPM 500e. In this case, the distribution p(e|t,c) is defined over the filtered perception output of the final filter 205. During testing, the PSPM 500e will be applied to the ground truth derived from the simulation, and the resulting PSPM samples will be fed directly into the prediction stack 104.
[0230] Slice Considerations
[0231] The factor when deciding where to "slice" the perception stack is the complexity of the required ground truth (the required ground truth will correspond to the perception components after the stack is sliced): A potential motivation for the PSPM approach is to have ground truth that is relatively easy to measure. The lowest part of the perception stack 102 operates directly on sensor data, but the information level required for planning and prediction is much higher. In the PSPM approach, the idea is to "bypass" the lower-level details while still providing a statistically representative perception output for prediction and planning during testing. Broadly speaking, the higher the perception stack 102 is sliced, the simpler the ground truth generally is.
[0232] Another consideration is the complexity of the perception components themselves, because any perception components not wrapped in a PSPM must be executed during testing.
[0233] It is generally expected that the slicing will always occur after the CNN in the perception stack, thus avoiding the need to simulate the input to the CNN and avoiding using the computational resources for running the CNN during testing.
[0234] In a sense, it is beneficial to wrap as much of the perception stack 102 into a single PSPM as possible. In the extreme case, this means the entire perception stack 102 is modeled as a single PSPM. The benefit of doing this is the ability to model any correlations between different sensors and / or perception components without having to know about these correlations. However, as more and more of the perception stack 102 is wrapped in a single PSPM, this significantly increases the complexity of the system being modeled.
[0235] For Figure 9A , each individual PSPM 500a, 500b, 500c can be constructed independently of the data of a single sensor modality. This has the benefit of modulation - existing PSPMs can be rearranged to test different configurations of the perception slice 204 without having to retrain. Finally, the optimal PSPM architecture will depend on the context.
[0236] Especially in Figure 9C cases, it may also be necessary to use a time-dependent model to fully capture the dependence on previous measurements / perception outputs introduced by the final filter 205. For example, Figure 9C the PSPM500e can take the form of a hidden Markov model to capture an additional level of time dependence. More generally, such time-dependent PSPMs can be used for any of the above. This is the context where time-dependent models are useful, but in many cases, explicit modeling of time dependence may be useful.
[0237] For Figure 9A and Figure 9B, cutting off before the final filter 205 has the following benefits: it may not be necessary to introduce explicit time dependence, i.e., the form of PSPM can be used, and the form of PSPM has no explicit dependence on the PSPM samples previously obtained from the PSPM.
[0238] Example of PSPM
[0239] The above description mainly focuses on dynamic objects, but PSPM can also be used in the same way for static scene detectors, classifiers, and other static scene perception components (e.g., traffic light detectors, lane offset correction, etc.).
[0240] In fact, PSPM can be built for any part of the perception stack 102, including:
[0241] - Odometry, such as:
[0242] ο IMU,
[0243] ο Visual-odometry,
[0244] ο LIDAR-odometry,
[0245] ο RADAR-odometry,
[0246] ο Wheel encoders;
[0247] -(Self-) Localisation, such as:
[0248] ο Vision-based localisation,
[0249] ο GPS localisation (or more generally satellite localisation).
[0250] "Odometry" refers to the measurement of local relative motion, while "Localisation" refers to the measurement of global position on the map.
[0251] PSPM can be built in exactly the same way to model the perception output of such perception components using appropriate perception ground truth.
[0252] These allow realistic odometry and localisation errors to be introduced into the simulated scene in the same way as detection errors, classification errors, etc.
[0253] Ground Truth Pipeline
[0254] As described above, the generation of annotations in the ground truth pipeline 802 can be manual, automatic, or semi-automatic annotations.
[0255] Automatic or semi-automatic ground truth annotations can use high-quality sensor data that is not typically available at runtime for the AV (or at least not available all the time). In fact, this can provide a way to test whether such components are needed.
[0256] Automatic or semi-automatic annotations can use offline processing to obtain a more accurate perception output, which can be used as the ground truth for PSPM construction. For example, to obtain the perception ground truth for localization or odometry components, offline processing such as bundle adjustment can be used to reconstruct the vehicle's path with high precision, which can then be used as the ground truth to measure and model the accuracy of the AV's online processing. Due to computational resource limitations or because the algorithms used are inherently non-real-time, this offline processing may not be feasible on the AV itself at runtime.
[0257] Examples of Confounding Factors
[0258] Figure 10 A high-level overview of the various factors that can lead to uncertainty in the perception output (i.e., the various sources of potential perception errors) is shown. This includes further examples of confounding factors c that can be incorporated as variables in the PSPM:
[0259] - Occlusion
[0260] - Lighting / time of day
[0261] - Weather
[0262] - Season
[0263] - (Linear and / or angular) distance to the object
[0264] - (Linear and / or angular) speed of the object
[0265] - Position in the sensor field of view (e.g., angle from the image center)
[0266] - Other object characteristics, such as reflectivity, or other aspects of its response to different signals and / or frequencies (infrared, ultrasound, etc.)
[0267] Other examples of possible confounders include the map of the scenario (indicating the environmental structure) and inter-agent variables such as “business” (a measure of the number or density of agents in the scenario), the distance between agents, and the agent types.
[0268] Each can be numerically or categorically characterized in one or more variable components (dimensions) of the confounder space C.
[0269] However, note that a confounding factor can be any variable that represents something about the physical world that may be related to perceptual error. This does not necessarily have to be a directly measurable physical quantity such as speed, occlusion, etc. For example, another example of a confounding factor related to another actor could be "intent" (e.g., whether a cyclist intends to turn left or continue straight ahead at an upcoming turn at a particular moment, which can be determined from real-world data of the actions the cyclist actually takes by looking ahead at a given time). In a sense, variables such as intent are latent variables or unobserved variables, in that at a particular moment (in this case, before the cyclist takes a definitive action), the intent cannot be directly measured using the perception system 102 but can only be inferred through other measurable quantities; the point about confounding factors is that it is not necessary to know or measure those other measurable physical quantities in order to model the effect of intent on confounding factor error. For example, the perceptual error associated with a cyclist with an intent to "turn left" may be statistically significantly increased compared to a cyclist with an intent to "continue straight ahead", which may be due to the multiple, unknown, and potentially complex behavioral changes in the behavior of the cyclist who intends to turn left, meaning that in practice, the perception system has a worse perception of them. By introducing the "intent" variable as a confounding factor in the error model, there is no need to try to determine which observable physical manifestations of intent are related to perceptual error - as long as the "intent" ground truth can be systematically assigned to the training data in a way consistent with the simulation (in this case, the intent of the cyclist is known in the simulation in order to simulate their behavior as the scene unfolds), then such data can be used to build appropriate behavioral models for different intents in order to simulate that behavior, as well as a perception error model that depends on intent, if any, without having to determine which physical manifestations of intent are actually related to perceptual error. In other words, in order to model the effect of intent on perceptual error, it is not necessary to understand why intent is related to perceptual error, because intent itself can be modeled as a perceptual confounding factor (rather than trying to model the observable manifestations of intent as confounding factors).
[0270] Low Level Errors
[0271] Examples of low-level sensor errors include:
[0272] - Registration errors
[0273] - Calibration errors
[0274] - Sensor limitations
[0275] Such errors are not explicitly modeled in the simulation, but their effects are wrapped in the PSPM used to model the perception slices that interpret the applicable sensor data. That is, these effects will be encoded in the parameters θ that characterize the PSPM. For example, for a Gaussian-type PSPM, such errors result in larger covariance representing greater uncertainty.
[0276] High-level Perception Errors
[0277] Other errors may occur in the perception pipeline, such as:
[0278] - Tracking errors
[0279] - Classification errors
[0280] - Dynamic object detection failures
[0281] - Fixed scene detection failures
[0282] When it comes to detection, false positives and false negatives can cause the prediction stack 104 and / or the planner 106 to operate in unexpected ways.
[0283] Build specific PSPMs in a statistically robust way to model such errors. These models can also take into account the effects of the variable confounding factor c.
[0284] Taking object detection as an example, detection probabilities can be measured and used to build a detection distribution that depends on, for example, distance, angle, and occlusion level (the confounding factor c in this example). Then, when running the simulation, through ray tracing from the camera, it can be determined according to the model whether the object "might" be detectable. If so, the measured detection probability is checked, and the object is detected or not detected. This deliberately introduces the possibility of not detecting sensor-sensitive objects in the simulation in a way that reflects the behavior of the perception stack 102 in real life, because the detection failures have been modeled in a statistically robust way.
[0285] This method can be extended in a Markov model to ensure proper modeling of conditional detection. For example, an object can be detected with an appropriate probability only if the object has been detected beforehand, otherwise the probability may be different. In this case, false negatives involve some temporal dependence on the simulated detections.
[0286] False positives can be randomly generated with a density similar to the density in space and time measured by PSPM. That is, in a statistically representative manner.
[0287] 2. Problem Statement
[0288] As a further explanation, this section elaborates on the mathematical framework of PRISM and introduces the specific dynamic object detection problems to be addressed in the subsequent sections. Section 3 discusses the datasets used for training PRISM, the techniques for identifying relevant features, and the description of the evaluation methods. Section 4 describes the specific modeling decisions and how data science informs these decisions.
[0289] Note that in the following description, the symbols x g , y g , z g can be used to represent the coordinates of the position-aware ground truth t. Similarly, x s , y s , z s can be used to represent the coordinates of the position-aware stack output e. Thus, the distribution p(x s , y s , z s |x g , y g , z g ) is a form that the above-mentioned perception uncertainty distribution p(e|t) can take. Similarly, x can be used below to generically refer to the set of confounding factors, which is equivalent to the set of confounding factors c or c' described above.
[0290] The perception system has inputs that are difficult to simulate, such as camera images, lidar scans, and radar echoes. Since these inputs cannot be rendered with perfect realism, the perception performance in the simulation will not match the perception performance in the real world.
[0291] The goal is to construct a probabilistic surrogate model for the perception stack, called PRISM. PRISM uses a low-fidelity representation of the world state (perceptual ground truth) and produces a perceptual output in the same format as the vehicle stack (or more precisely, the perceptual slice 204 being modeled). When the stack runs on real data, the samples drawn from the surrogate model in the simulation should be similar to the output of the perception stack.
[0292] PRISM sampling should be fast enough to be used as part of a simulation system for the validation and development of downstream components such as planners.
[0293] 2.1 Intuition
[0294] For the following considerations, the sections below state the most general case:
[0295] ● There is some stochastic function that maps from the true state of the world to the output of the perception stack.
[0296] ● This function can be modeled using training data. The function is modeled as a probability distribution.
[0297] · Since the world state changes smoothly over time, the sampled perceptual outputs should also change smoothly over time. Since the world state is only partially observed, the appropriate way to achieve this is to make the probability distribution depend on the observed world state and the history of perceptual outputs.
[0298] · The simulator (Genie) is responsible for generating a representation of the world at runtime. The output of Genie is the 6D pose and extent of dynamic objects, as well as some other information such as road geometry and weather conditions.
[0299] · For real-world training data, this world representation is obtained from annotations.
[0300] Mathematical statements
[0301] 2.2.1 Preliminaries
[0302] For any set S, let the set of histories of S be An element (t, h) ∈ histories(S) consists of t (the current time) and h (a function that returns elements of S at any time in the past). The symbol denotes the simulation equivalent of x.
[0303] The perception system is the stochastic function f: histories(World) → histories(Perception). In general, the form of f will be:
[0304]
[0305] sense: histories(World) → histories(SensorReading),
[0306] perceive: histories(SensorReading) → histories(Perception). (1)
[0307] The goal is to simulate some f. The world state can be decomposed into a set of properties ObservedWorld and a set of everything else UnobservedWorld (the exact pixel values of the camera image, the temperature at each point on each surface), such that there is a bijection between World and ObservedWorld × UnobservedWorld, and the set of properties ObservedWorld can be reliably measured (this may include the meshes and textures of each object in the scene, the positions of the light sources, the material densities, etc.). In traditional photorealistic simulation methods, simulating f is equivalent to finding some stochastic function histories(ObservedWorld) → histories(SensorReading), which can be combined with perceive to form
[0308]
[0309] Let observe: World → ObservedWorld be the function that maps world states to their observed counterparts. Note that this function is not one-to-one: there will be many world states that map to a single observed world state. An accurate and useful simulation of f for all histories (t, h) ∈ histories(World) will have
[0310]
[0311] where map: ((S → T) × histories(S)) → histories(T) maps a function over histories.
[0312] Then one must conclude that the best photorealistic simulation has Make
[0313]
[0314] Due to the Combining Equations 1, 2, and 4 gives Equation 3, so sense predicts the history of sensor readings, the joint distribution of the history (SensorReading), and the correlations between different sensor readings allow dependencies on unobserved features of the world to be modeled more effectively. Therefore, in the calculation of Similar correlations should be observed in .
[0315] Because SensorReading has high dimensionality, and sense is a random function (because it depends very much on unobserved properties of the world), finding It is very important to make equation 4 even approximately true. Therefore, it is straightforward to find
[0316] 2.2.2 Creating a surrogate model
[0317] The creation of a proxy model can be characterized as a random function estimation task. Let S+ be a finite sequence of elements of S. Let For an element s i A sequence of length N. Get a dataset of sensor reading sequences
[0318]
[0319] Among them, each I ij ∈SensorReading is the time t in run i ij The sensor readings, and M i is the number of timestamps in a particular run. A new dataset is constructed using the function annotate:SensorReading→ObservedWorld that recovers the observed scene parameters from the sensor readings.
[0320] x ij =annotate(I ij ), y ij =perceive(I ij ).
[0321] The task of PRISM is then to estimate the The distribution of samples in A realization of can be obtained by drawing samples from this distribution.
[0322] Dependencies on previously sampled stack outputs are included because the distribution of y meaningfully depends on the unobserved world, and the unobserved world changes smoothly over time. As discussed in Section 2.2.1, this dependence on the unobserved world means that y will change smoothly over time in a way that is difficult to model solely based on the dependencies. This time-related property of stack outputs was explored for the perception system 102 in Section 4.2.3, where strong correlations over time were found.
[0323] Samples from the learned PRISM distribution give reasonable perception outputs conditioned on the low-fidelity scene representation and the history of previous samples. These factors are independent variables in the generative model, and the dependent variable is the perceived scene. Independent variables that meaningfully affect the distribution of the dependent variable are called confounders in this paper. Part of the process of constructing the PRISM model is to identify the relevant confounders to include in the model and how these confounders should be combined. Section 3.2 explores a method for identifying relevant confounders.
[0324] 2.2.3 Dynamic Object Problem
[0325] A specific example of a perception system is presented—a system that uses RGBD images to detect dynamic objects in a scene. A "dynamic object" is a car, truck, cyclist, or other road user described by an oriented bounding box (6D pose and extent). The observed world is a collection of such dynamic objects. In this setting,
[0326]
[0327] SensorReading = Image = [0, 1] w×h×4 , where, is the set of finite subsets of S, 1 Type represents the object type (Car, Van, Tram, Pedestrian), Spin(3) is the set of unit quaternions, and Info is an arbitrary set whose elements describe additional characteristics of the dynamic object, e.g., the extent to which the object is occluded by other (possibly static) objects in a scene closer to the camera. Dynamic objects are useful when characterizing the behavior of the perception system.
[0328] This example further simplifies the dynamic object problem by only choosing to model the positions of dynamic objects in a given ObservedWorld. This includes fitting a model for the possibility that observable objects are not perceived (false negatives).
[0329] As shown in Section 4.2.8, false negatives are more common errors made by the perception system 102 than false positives (spurious dynamic object detections).
[0330] For simplicity, the following description only considers hazardous objects (poisons) in 3D space and omits the discussion of orientation, range, object type, or other possible perception outputs. However, the principle can equally be applied to such other perception outputs.
[0331] 3 Method
[0332] 3.1 Data
[0333] A specific driving scenario is presented, and the data for this specific driving scenario has been recorded multiple times under similar conditions. The scenario mentioned herein by way of example is a roundabout located southeast of London on a test route. Figure 12 The background of the roundabout and the path of the vehicle passing through can be seen therein, where it is observed from a camera as shown Figure 13 as shown.
[0334] By restricting the PRISM training data to runs on the same roundabout under similar climate conditions, the influence of weather and sunlight as confounding factors on perception performance is minimized. The potential performance of PRISM tested on similar collected data is equally maximized. For example, by estimating how PRISM trained on roundabout data performs in a highway scenario, the performance of PRISM can be tested on out-of-domain data.
[0335] 3.1.1 Dataset Generation
[0336] PRISM training requires a dataset containing sufficient information to learn the distribution of perception errors. For simplicity, this section only considers the errors introduced by the perception system 102 when predicting the center position of a dynamic object in the camera frame of an observed dynamic object. To learn such errors, ground truth center and perceived center estimates are required.
[0337] The ground truth center positions are estimated from the human-annotated 3d bounding boxes present in each frame of the recorded video sequence of the roundabout and applied to all dynamic objects in the scene. These bounding boxes are fitted to the scene using a ground truth tooling suite. The ground truth tooling suite combines camera images, stereo depth pointclouds, and lidar pointclouds into a 3D representation of the scene to maximize annotation accuracy. It is assumed that the annotation accuracy is good enough to be used as ground truth.
[0338] Figure 9 The process of obtaining stack prediction objects from the recorded camera images is shown. It should be noted that the pipeline is stateless and each pair of camera frames is processed independently. This forces any temporal correlations found in the perception error data to be attributed to the behavior of the detector on closely related inputs rather than the internal state of the detection algorithm.
[0339] In general, the set of object predictions indexed by image timestamps combined with a similarly indexed set of ground truth data from the ground truth tooling suite is sufficient for PRISM training data. However, all models considered in this section are trained on data that has been passed through additional processing steps to generate associations between the ground truth and predicted objects. This limits the space of alternative models but simplifies it by breaking the fitting task into the following parts: models that fit the position error; models that fit for generating false negatives; models that fit for generating false positives. The association algorithm used runs independently on each frame. For each timestamp, the set of stack prediction objects and the set of ground truth objects are compared using the intersection over union (IOU), where the prediction object with the highest confidence score (a measure indicating how good the prediction is generated by the perception stack 102) is considered first. For each predicted object, the ground truth object with the highest IOU is associated with it, forming pairs for learning the error distribution. Pairs with an IOU score less than 0.5 (a tunable threshold) do not form any associations. After all predicted objects have been considered for association, the set of unassociated ground truth objects and the set of unassociated predicted objects are retained. The unassociated ground truth objects are stored as false negative examples, while the unassociated predicted objects are stored as false positive examples.
[0340] "Set-to-set" models that do not require such associations will be considered later.
[0341] 3.1.2 Content of Training Data
[0342] The previous section described how to generate PRISM training data and divide it into three sources: association, false negatives, and false positives. Table 1 specifies the data present in each source and provides the following definitions:
[0343] centre_x, centre_y, centre_z: The x, y, and z coordinates of the center of the ground truth 3D box.
[0344] orientation_x, orientation_y, orientation_z: The x, y, and z components of the axis-angle representation of the rotation from the camera frame (right of the front stereo) to the ground truth 3D box frame (box frame).
[0345] height, width, length: The extents of the ground truth 3D box along the z, y, and x axes in the coordinate system of the 3D box.
[0346] manual_visibility: A label applied by a human annotator to indicate which of four visibility categories the ground truth object belongs to. The categories are: fully-occluded (100%), largely-occluded (80 - 99%), somewhat-occluded (1 - 79%), and fully-visible (0%).
[0347] occluded: The portion of the area where the ground truth 2D bounding box overlaps with the 2D bounding boxes of other ground truth objects closer to the camera.
[0348] occluded_category: A combination of manual_visibility and occluded, which can be considered the maximum of the two. It is useful to combine manual_visibility and occluded in this way to maximize the number of correct occlusion labels. To see this, note that the occluded score of an object occluded by the static parts of the scene (bushes, trees, traffic lights) will be 0, but will have manual_visibility correctly set by a human annotator. Objects occluded only by other ground truth objects do not have the manual_visibility field set by a human annotator, but will have the correct occluded field. These two cases can be handled by taking the maximum of the two values. Even with this logic, it is possible for the 2D bounding box of a ground truth object to completely obscure the 2D bounding box of an object behind it, even if some of the background objects are visible. This will generate some fully-occluded cases that can be detected by the perception system.
[0349] truncated: The parts of the eight vertices of the ground truth 3D box that are outside the sensor frustum.
[0350] type: When attached to a ground truth object (false negative, the ground truth part of an association pair), this is the object type annotated by a human, such as Car or Tram. When attached to a predicted object (false positive, the predicted part of an association pair), this is the perception stack's best guess of the object type, with the object type limited to Pedestrian or Vehicle.
[0351] In addition to the above, this section will also mention the following derived quantities:
[0352] distance: The distance from the object center to the camera, calculated as the Euclidean norm of the object center position in the camera frame.
[0353] azimuth: The angle formed between the projection of the ray connecting the camera and the object center on the camera's y = 0 plane and the positive z-axis of the camera. The polarity is defined by the direction of rotation around the camera y-axis. Since objects behind the camera cannot be observed, the range is limited to [-π / 2, π / 2].
[0354] Table 1
[0355]
[0356] The composition of the dataset will be discussed in detail where relevant in a later section. The following presents a high-level summary of the data.
[0357] · 15 traversals of the roundabout scenario, spanning a total footage of approximately 5 minutes.
[0358] · 8,600 unique frames containing 96k ground truth object instances visible to the camera.
[0359] · Among these 96k instances: 77% are cars; 14% are vans; 6% are pedestrians; 3% belong to smaller groups.
[0360] · Among these 96k instances: 29% are fully visible; 43% are somewhat occluded; 28% are mostly occluded.
[0361] In Table 1, specific data elements exist in each of the three generated PRISM data sources. X indicates that the column exists in the given data source. GT = ground truth, FN = false negative, FP = false positive. Each of these is actually three independent variables (e.g., center_x, center_y, center_z), but for readability, they are "compressed" here. *In the case with an asterisk, the type of content can be "Vehicle" or "Pedestrian", which are the only categories predicted by the five perception stacks. In the case without an asterisk, there are more classes (such as "Lorry" and "Van"), which are all the classes reported in the ground truth data.
[0362] 3.1.3 Training and Test Data
[0363] For all the modeling experiments described in this paper, the roundabout dataset was split into roughly equal halves to form a training set and a test set. No hyperparameter optimisation was performed, and thus, no validation set was required.
[0364] 3.2 Identifying Relevant Confounding Factors
[0365] The PRISM model may consider many confounding factors. Instead of optimising the model for every possible combination of confounding factors, it is preferably performed on a limited set of known relevant confounding factors.
[0366] To identify relevant confounding factors, a Relief-based algorithm is used. The general outline of the Relief-based algorithm is given in Algorithm 1. The Relief algorithm produces an array of feature weights in the range [-1, 1], where weights greater than 0 indicate that the feature is relevant, as the feature variation tends to change the target variable. In practice, some features will accidentally have weights greater than 0, and only features with weights greater than some user-defined cutoff 0 < τ < 1 are selected.
[0367]
[0368] The algorithm has the following desirable properties:
[0369] · It is sensitive to the non-linear relationships between features and the target variable. Other feature selection methods are not sensitive to these kinds of relationships, such as naive principal component analysis or comparison of Pearson correlations. Not all irrelevant things are independent.
[0370] · It is sensitive to the interactions between features.
[0371] · It is conservative. It will accidentally include irrelevant or redundant confounding factors rather than accidentally excluding relevant confounding factors.
[0372] It is important to note the following considerations of this method:
[0373] · It identifies the correlations in the data but does not provide in-depth understanding of how or why the target variable is related to the confounding factors under investigation.
[0374] · The results are dependent on the parameterization of the confounding variables.
[0375] There are many extensions of the Relief algorithm. An extension called MultiSURF is used here. It is found that MultiSURF performs well in a wide range of problem types and is more sensitive to the interactions of three or more features than other methods. This implementation is used from scikit-rebate, an open-source Python library that provides implementations of many Relief-based algorithms, and the Relief-based algorithms are extended to cover scalar features and target variables.
[0376] In the experiment, use where n is the size of the dataset, and α = 0.2 is the desired false discovery rate. According to Chebyshev's inequality, we can say that the probability of accepting an irrelevant confounding factor as relevant is less than α.
[0377] Relief-based methods are useful tools for identifying possible confounding factors and their relative importance. However, not all features that affect the error characteristics of the perception system will be captured in the annotated training data. A manual process of examining model failures to hypothesize new features to label as confounding factors is necessary.
[0378] 4 Model
[0379] 4.1 Heuristic Model
[0380] Camera coordinates represent the position of points in an image in pixel space. In binocular vision, the camera coordinates of points in two images are available. This allows the reconstruction of the position of points in a 3D Cartesian world. The camera coordinates of a point p in 3D space are given by:
[0381]
[0382] where (u1, v1), (u2, v2) are the image pixel coordinates of p in the left and right cameras respectively, (x p , y p , z p ) are the 3D world coordinates of p relative to the left camera, b is the camera baseline, and f is the camera focal length. As Figure 14 shown. The disparity d is defined as:
[0383]
[0384] The 3D world coordinates of p can be written as:
[0385]
[0386]
[0387] The heuristic model is obtained by imposing a distribution in camera coordinates and propagating it to 3D coordinates using the above relationships. This distribution can equally be used for object centers or object extents. When the image is discretized into pixels, the model allows for the consideration of the physical sensor uncertainties of the camera. The model is given by:
[0388] p(x s , y s , zs | x g , y g , z g ) = ∫∫∫ p(x s , y s , z s | u1, v, d) p(u1, v, d | x g , y g , z g ) du1 dv d, (12)
[0389] where (x g , y g , z g ) are the coordinates of the ground truth points, and (x s , y s , z s ) are the coordinates of the stack prediction. Given the world coordinates, the probability distribution on the camera coordinates is:
[0390] p(u1, v, d | x g , y g , z g ) = p(u1 | x g , y g , z g ) p(v | x g , y g , z g ) p(d | x g , y g , z g ), (13)
[0391] where the distribution independence in each camera coordinate is assumed:
[0392]
[0393] where σ is a constant, is the normal distribution, and Lognormal is the log-normal distribution. Lognormal is chosen because it only supports positive real numbers. This defines the probability density of a normal distribution centered at the camera coordinates of points in 3D space. The normal distribution is chosen based on mathematical simplicity. If only discretisation error is considered, a uniform distribution might be more appropriate. However, other errors are likely to lead to uncertainties in stereo vision, so the extended tails of the normal distribution are useful for modelling such phenomena in practice. For a front-facing stereo camera, α is determined to be 0.7 by maximum likelihood estimation. p(x s , y s , z s |u1, v, d) is given by a Dirac distribution centered at the point values of x s , y s and z s obtained from equations 9 - 11.
[0394] By forming a piecewise constant diagonal multivariate normally distributed approximation of equation 12, by solving the integral using Monte Carlo simulation and using the mean and variance of the sampled values for different x g , y g and z g values to estimate p(x s , y s , z s |x g , y g , z g ), a runtime model is obtained.
[0395] The model can be improved by considering a more accurate approximation of the conditional distribution in equation 12, or by modelling the uncertainty in the camera parameters f and b (set to their measured values in the model). How to extend the model to include time dependence is an open question.
[0396] 4.2 PRISM
[0397] An attempt is made to build a plausible surrogate model (PRISM) of the perception stack / sub-stack 204 guided by data analysis. The model includes non-zero probabilities of position errors that are time-dependent and of objects not being detected, which are significant features of the data.
[0398] 4.2.1 Positional errors
[0399] The center position of a dynamic object detected by the perception stack will be modeled using an additive error model given by:
[0400] y k = x k + e k
[0401] where y k is the observed position of the object, x k is the ground truth position of the object, and e k is the error term, all at time t k . The phrase "positional error" will be used to refer to the additive noise component e k of this model.
[0402] Figure 15 shows the positional errors of a particular dynamic object detected by the perception stack relative to the human-labeled ground truth. A lag plot of the same data can be found in Figure 16 , indicating strong temporal correlations of these errors. It can be concluded from these plots that the generative model of positional errors must condition each sample on previous samples. An autoregressive model is proposed for the time-dependent positional errors, where each error sample linearly depends on previous error samples and some noise. The proposed model can be written as:
[0403] e k = e k-1 + Δe k (18)
[0404] where e k is the positional error sample at time step k, and Δe k is the random term, and Δe k can be a function of one or more confounding factors, often referred to as "error deltas". Figure 17 shows a diagram visualizing this model, including the dependencies on the hypothesized confounding factors C1 and C2.
[0405] The model is based on several assumptions. First, subsequent error increments are independent. This is explored in Section 4.2.3. Second, the empirical distribution of the error increments can be reasonably captured by a parametric distribution. This is explored in Section 4.2.4. Third, the described model is stationary such that the mean error does not change over time. This is explored in Section 4.2.5.
[0406] 4.2.2 Piecewise Constant Model
[0407] It has been shown that modeling the location error requires subsequent errors to be conditioned on the previous error, but how should the first error sample be chosen? Now consider the task of fitting a time-independent distribution of location errors. If no time correlation is found in the data, the method adopted here can equally be applied to all samples of each dynamic object, not just the first sample.
[0408] In general, such a model would be a complex joint probability distribution of all confounding factors. As discussed in Section 2, due to incomplete scene representation in the perception stack (ObservedWorld≠World) and possible uncertainties, the distribution of possible perceptual outputs given the ground truth scene is expected. The expected variance is heteroskedastic; it varies based on the values of the confounding factors. As a simple example, it should not be surprising that the error in the location estimate of a dynamic object has a variance that increases with the distance of the object from the detector.
[0409] The conditional distribution modeled by PRISM is expected to have a complex functional form. This functional form can be approximated by discretizing each confounding factor. In this representation, categorical confounding factors (e.g., vehicle type) are mapped to bins. Continuous confounding factors (e.g., distance to the detector) are split into ranges, and each range is mapped to a bin. The combination of these discretizations is a multi-dimensional table for which an input set of confounding factors is mapped to bins. It is assumed that within each bin, the variance is homoskedastic, and a distribution with constant parameters can be fitted. Global heteroskedasticity is captured by different parameters in each bin. A model of a distribution with fixed parameters in each bin is referred to in this paper as the Piecewise Constant Model (PCM). Examples of general implementations of similar models can be found in the literature. Mathematically, this can be written as P(y|x) ∼ G(α[f(x)], β[f(x)],...), where y is the set of outputs, x is the set of confounding factors, f(·) is the function that maps the confounding factors to bins, and G is the probability distribution with parameters α[f(x)], β[f(x)],... fixed within each bin.
[0410] In the PCM for PRISM, it is assumed that the error is additive, i.e., the predicted position, pose, and range of the stack of dynamic objects is equal to the ground truth position, pose, and range plus some noise. The noise is characterized by a distribution in each bin. In this PCM, it is assumed that this noise is normally distributed. Mathematically this can be written as:
[0411]
[0412] where y is the stack observation, is the ground truth observation, and ∈ is the noise. The distribution in each bin is characterized by a mean μ and a covariance Σ. μ and Σ can be regarded as functions of the confounding factor bin.
[0413] Figure 18 An example binning scheme is shown. The bins are formed by the azimuth angle and the distance to the center of the ground truth dynamic object.
[0414] Training the model requires the ground truth and stack predictions (actual perception outputs) collected as described in Section 3.1.1. (For example, using the maximum a posteriori method to incorporate the prior) Fit the mean and covariance of a normal distribution to the observations in that bin. For the mean of the normal distribution, use the prior of the normal distribution. For the scale of the normal distribution, use an Inverse Gamma prior.
[0415] To set the hyperparameters of the prior, physical knowledge can be combined with the intuition regarding how quickly the model should ignore the prior when data becomes available. This intuition can be represented by the concept of pseudo-observations, i.e., how much the prior distribution is weighted relative to the actual observations (encapsulated in the likelihood function) in the posterior distribution. Increasing the number of pseudo-observations results in a prior with lower variance. The hyperparameters for the normal distribution prior can be set to μ h = μ p and where μ p and σ p represent the prior point estimates of the mean and standard deviation of the bin under consideration, and n pseudo represents the number of pseudo-observations. The rate and scale hyperparameters of the Inverse Gamma prior can be set to and For this model, choose n pseudo = 1 and use the heuristic model described in Section 4.1 to provide prior point estimates for the parameters of each bin.
[0416] The advantage of the PCM method is that it accounts for global heteroscedasticity, provides a unified framework for capturing different types of confounding factors, and it utilizes simple probability distributions. Additionally, the model is interpretable: the distribution in the bin can be examined, the training data can be directly inspected, and there are no hidden transformations. Moreover, the parameters can be analytically fit, meaning that the uncertainty from lack of convergence in the optimization routine can be avoided.
[0417] Confounding factor selection
[0418] To select appropriate confounding factors for the PCM, use the method described in Section 3.2 and the data described in Section 3.1.2. The findings applied to the position, range, and orientation errors are shown in Table 2.
[0419] Table 2: Shows the confounding factors identified as important for the target variable under consideration
[0420]
[0421] As can be seen from Table 2, for d_centre_x and d_centre_z, the relevant confounding factors are some combination of the object's position relative to the camera and the degree to which the object is occluded. The perception system 102 assumes that the detected object exists on the ground plane, y = 0, which may be the reason why d_centre_y does not show a dependence on distance.
[0422] For the model of the position error of the dynamic object detected by the perception system 102, this analysis identifies position and occlusion as good confounding factors to start with. The data does not show a strong preference for position confounding factors in polar coordinates (distance, azimuth) based on the Cartesian grid (centre_x, center_y, center_z). Distance and azimuth are used in the PRISM prototype described herein, but a more in-depth evaluation of the relative performance of each prototype can be performed. 4.2.3 Temporal correlation analysis of position error increments
[0423] The temporal correlation analysis performed on the position error can be repeated for the time series of error increments, giving Figure 19 the lag plots shown in. These plots show much less temporal correlation in the error increments than was found in the position error. The Pearson correlation coefficients for the error increments are shown in Table 3. For each dimension, they are quite small in magnitude, with -0.35 being the furthest from zero. From this analysis, it can be concluded that a good model of the error increments can be formed from independent samples from the relevant distribution.
[0424] Table 3: Pearson correlation coefficients for error increment samples and samples with a one-time-step delay
[0425]
[0426] Distribution of position error increments
[0427] Typically, the x, y, z error increment dimensions are correlated. Here, they are considered independently, but note that future efforts could consider joint modeling of them. Figure 20 Histograms of the error increment samples are presented in, from which it can be clearly seen that the error increments are more likely to be close to zero, but have long tails with extreme values. Figure 21Shows the maximum likelihood best-fit of this data to some test distributions. Visual inspection of these plots indicates that the Student's t-distribution can be a good modeling choice for generating error increments. Due to the presence of a large number of extreme error increments in the data, the normal distribution does not fit well.
[0428] Bounding the random walk
[0429] The autoregressive error increment model proposed in Section 4.2.1 is generally a non-bounded stochastic process. However, it is well known that the detection location of a dynamic object does not simply deviate but remains near the ground truth. This is an important property that must be captured in a time-dependent model. As a specific example of this point, consider modeling the position error as a Gaussian random walk, setting This results in a position error distribution at time t where the variance increases with time without bound. Such a property must not exist in the PRISM model. k
[0430] AR(1) is a first-order autoregressive process defined by:
[0431] y t = a1y t-1 + ∈ t (19)
[0432] where ∈ t is a sample from a zero-mean noise distribution, and y t is a sample of the variable of interest at time t. This process is known to be wide-sense stationary for | a1 | < 1, otherwise the generated time series is non-stationary. Comparing Equation 18 and Equation 19, it can be seen that, given the known results of AR(1), if Δe k is zero-mean, the error increment model proposed in Equation 18 will be non-stationary. Therefore, such a model is not sufficient to generate a reasonable stack output.
[0433] An extension to the model proposed in Equation 18 is proposed, motivated by the nature of the collected error increment data. The extension is to model Δe k conditional on the previous error, such that a search is made for P(Δe k | e k Best fit for (e
[0434] Following the piecewise-constant modeling approach described in Section 4.2.2, P(Δe k |e k -1) is approximated as follows:
[0435] For the space of e k-1 values, M bins are formed, with boundaries {b0, b1, ..., b M}.
[0436] A separate distribution P m (Δe k ) is characterized for each bin, where 0 < m < M represents the bin index.
[0437] Given the previous time-step position error e k-1 , the next error increment is drawn from P m (Δe k ), where B m-1 < e k-1 < B m .
[0438] Figure 22 Figure [X] shows the sample means computed on the PRISM training data for M = 5. The trends revealed are consistent with expectations and have the following intuitive explanation. Consider a series of samples of error increments of the same polarity, accumulating an absolute position error away from the ground truth. To make the overall process appear stationary, subsequent samples of error increments of the same polarity should be less likely to change in the direction towards the true object position. This observation helps to explain the negative Pearson coefficients presented in Table 3, which indicate that subsequent error increments slightly tend to reverse polarity.
[0439] The binning scheme for P m (Δe k ) suffers from the typical PCM drawback of low sample counts in extreme bins. A simple prior can be used to reduce this risk, e.g., setting the mean of the distribution in each bin to follow μ m = -ae m , where e m is the center value of the m-th bin and a > 0. It is worth noting that if P m (Δe k ) is chosen to be Gaussian such that then the time-dependent model becomes:
[0440]
[0441] The time-dependent model is a typical AR(1) process and is stationary for a < 2. In practice, a good prior will require a ~ 0, so such a model is stationary by construction.
[0442] 4.2.6 Simple Verification
[0443] It is useful to see whether the samples from the proposed time-dependent position error model reproduce the features that motivated its construction. Figure 23 A plot of the position error of a single dynamic object trajectory sampled from the learned distribution is shown. Figure 24 A lag plot of the same data is shown. In both cases, the true perception error data is provided for visual comparison. The similarity between the PRISM samples and the observed stack data is encouraging. Clearly, more quantitative evaluation (which will be the subject of Section 5) is needed to make any meaningful sanity claims.
[0444] False Negatives and False Positives
[0445] The perception system has failure modes beyond noisy position estimates of dynamic objects. The detector may fail to identify an object in the scene, i.e., false negatives, or it may identify an object that does not exist, i.e., false positives. A surrogate model like PRISM must emulate the observed false negative and false positive rates of the detector. This section discusses the importance of false negative modeling, investigates which confounding factors affect the false negative rate, and presents two simple Markov models. It has been shown that using more confounding factors can produce Markov models with better performance and highlights some issues with doing so using a piecewise constant approach.
[0446] A survey was conducted to determine the frequencies of true positives (TP), false negatives (FN), and false positives (FP). The results are summarized in Table 4. False negative events are significantly more numerous than false positives. The counts in Table 4 apply to all object distances. The false negative count at such a distance from the detector seems unfair, i.e., humans would have a hard time identifying. Introducing a distance filter on the events reduces the factor by which false negatives are more prevalent than false positives, but the difference is still significant. When considering objects with a depth less than 50m, the number of TP / FN / FP events is 34046 / 13343 / 843. Reducing the distance threshold to a depth of 20m, the number of TP / FN / FP events is 12626 / 1236 / 201.
[0447] Table 4: Shows the numbers describing false positive and false negative events in the dataset
[0448]
[0449] 4.2.9 False Negative Modelling
[0450] Following the method described in Section 3.2, the importance of different confounding factors for false negatives was explored using the relief algorithm. Milts 2 For 20% of the training data samples selected randomly. The results are as Figure 25 shown. A 20% random sample of the data allows the algorithm to run with manageable memory usage. The target variable is the association category produced by the detector: associated or false negative. The associated category is called the association state. The same list of confounding factors as in Section 3.2 is used, where distance and azimuth are used instead of center_x, center_y, and center_z. It has been found that the binning scheme based on distance and azimuth is as good as the binning of the center values but with lower dimensions. In addition, occluded_category is used, which is the most reliable occlusion variable. In addition, the association state of the object in the previous time step is included as a potential confounding factor. This is labeled as "from" in Figure 25 . Note that "from" has three possible values: associated, false negative, and empty. When an object becomes visible to the detector for the first time, there will be no previous association state; in the previous time step, the detector did not detect the object. The association state for such time steps is considered empty, i.e., true negative. Similarly, for an object that disappears from the field of view, whether by exiting the camera frustum or being completely occluded, the empty association state is used as the previous association state for the first frame in which the object reappears.
[0451] From Figure 25 it can be seen that the most important confounding factor is the "from" category. This means that the strongest predictor of the association state is the association state in the previous time step. This relationship is intuitive; if the detector fails to detect an object in one time step, it is expected to detect the object in multiple frames. The object may be inherently difficult to recognize by the detector, or some characteristics of the scene (e.g., lens flare of the camera) may affect its perception ability and persist over multiple frames. The next most important confounding factor is occluded_category. This is again intuitive - if an object is occluded, it is more difficult to detect and thus more likely to be a false negative. Distance is also important. Again, this is expected; the farther an object is, the less information there is about it (e.g., in a camera image, a farther car is represented by fewer pixels compared to a closer car).
[0452] Guided by this assessment, a false negative model is constructed where the only confounding variable is the association state at the previous time step. This is a Markov model as it assumes that the current state depends only on the previous state. This is modelled by determining the probabilities of transitioning from the state at time step t-1 to the state at time step t. Denoting the association state as X, this amounts to finding the conditional probability P(X t |X t-1 ). To determine these transition probabilities, the frequencies of these transitions in the training data are calculated. This is equivalent to Bayesian likelihood maximisation. Table 5 shows the transition probabilities and the number of instances for each type of transition in the data. Each bin in Table 5 has over 800 entries, indicating that the implied transition frequencies are reliable. The observed transition probability from false negative to false negative is 0.98, and from associated to associated is 0.96. As expected from the Relief analysis results, these values show strong temporal correlation. Is there a reason attributable to the transition to the false negative state and the persistence of the false negative state? There is a 0.65 probability of transitioning from the empty (true negative) state to the false negative state, and a 0.35 probability of transitioning from the empty (true negative) state to the associated state. This means that when an object first becomes visible, it is more likely to be a false negative. Many objects enter the scene from a distance, which is likely to be an important factor in generating these initial false negatives. Some objects enter the scene from the side, particularly in the roundabout scenario considered in this example. Such objects are truncated in the first few frames, which may be a factor in the early false negatives. To explore these points in more detail, a model that depends on other factors is constructed.
[0453] Table 5: Probabilities (left two columns) of transitioning from the association state in the previous time step (row) to the association state in the current time step (column), and counts of the number of transitions in the training dataset (right two columns)
[0454]
[0455] As a first step towards a more complex model, a Relief analysis is performed to identify confounding factors important for the transitions, without considering the previous association state. MultiSURF is used on a randomly selected 20% of the training data samples. The results are as Figure 26 shown.
[0456] Figure 26The most important confounding factors that may affect the transition probability are: occluded_category, distance, and azimuth. In fact, using the criteria listed in Section 3.2, all confounding factors are good confounding factors. Based on this evidence, the next most complex Markov model is created; occluded_category is added as a confounding factor. Representing the association state X and the occluded category C, the conditional probability P(X t |X t-1 ,C t ). As with the first Markov model, these transition probabilities are determined from the training data by counting the occurrence frequencies. Table 6 shows the transition probabilities and the number of instances for each transition in the data.
[0457] Table 6: Probabilities (left two columns) of transitioning from the association state and occluded category in the previous time step (rows) to the association state in the current time step (columns), and counts of the number of transitions in the training dataset (right two columns)
[0458]
[0459] Table 6 shows that some frequencies are determined from very low counts. For example, there are only 27 transitions from a fully occluded false negative to the association state. However, this event is expected to be rare, and any of these transitions may indicate incorrect training data. These counts may come from an incorrect association of the annotated data with the detector's observations; if the object is truly fully occluded, then the detector would not be expected to observe it. Perhaps the least trustworthy transitions come from association and full occlusion; there are only 111 observations in total for this category. The probability of transitioning from association and full occlusion to association is 0.61, which is extremely likely; while the count of transitions from false negative and full occlusion to association is low, it actually has a zero probability (because the count from false negative and full occlusion to false negative is so high). Rows with low sums should be treated with caution.
[0460] Despite these limitations, there are still expected trends. When objects transition from the empty state (i.e., when they are first observed), if they are fully visible, there is a 0.60 chance of transitioning to association, meaning the object is more likely to be transitioned to association than a false negative. However, if the object is mostly occluded, the probability of transitioning to association is only 0.17.
[0461] Given the limitations of identification, it is possible to determine whether adding confounding factors improves the model. To compare these models, the method of using the model with the smaller negative log predictive density (NLPD) to better explain the data is adopted. The corresponding NLPD is calculated on the held-out test set. The NLPD of the simple Markov model is 10,197 compared to 9,189 of the Markov model with confounding factors. Adding the occlusion_category confounding factor improves the model by this metric.
[0462] This comparison shows that including confounding factors can improve the model. To construct a model that includes all relevant confounding factors, following the paradigm used in piecewise constant models, new confounding factors add additional bins (e.g., Table 6 has more rows than Table 5).
[0463] 5. Neural Network PRISM
[0464] This section describes how to implement PRISMS using a neural network or a similar "black box" model.
[0465] As is well known in the art, a neural network is formed by a series of "layers", which in turn are formed by neurons (nodes). In a classical neural network, each node in the input layer receives a component of the network input (e.g., an image), which is typically multi-dimensional, and each node in each subsequent layer is connected to each node in the previous layer and calculates a function of the weighted sum of the outputs of the nodes to which it is connected.
[0466] For example, Figure 27 illustrates node i in a neural network, which receives a set of inputs {u j} and calculates a function of the weighted sum of those inputs as its output:
[0467]
[0468] Here, g is a possibly non-linear "activation function", and {w i,j} is the set of weights applied to node i. The weights of the entire network are adjusted during training.
[0469] Reference Figure 28, it is useful to conceptualize the input and output of the layers of a CNN as "volumes" in a discrete three-dimensional space (i.e., a three-dimensional array), each discrete three-dimensional space being formed by a stack of two-dimensional arrays referred to herein as "feature maps". More generally, a CNN takes "tensors" as input, which can typically have any number of dimensions. The following description may also refer to a layer of a feature map as a tensor.
[0470] For example, Figure 28 shows a sequence of five such tensors 302, 304, 306, 308, and 310, which can be generated, for example, by a series of convolution operations, pooling operations, and non-linear transformations known in the art. For reference, the two feature maps within the first tensor 302 are labeled 302a and 302b, respectively, and the two feature maps within the fifth tensor 310 are labeled 310a and 310b, respectively. In this document, (x,y) coordinates refer to the applicable positions within a feature map or an image. The z-dimension corresponds to the "depth" of the feature map or image and may be referred to as the feature dimension. A color image has three depths corresponding to the three color channels, i.e., the value at (x,y,z) is the value of color channel z at position (x,y). The tensors generated at the processing layers within a CNN have a depth corresponding to the multiple filters applied at that layer, where each filter corresponds to a specific feature that the CNN is learning to recognize.
[0471] A CNN differs from classical neural network architectures in that a CNN has processing layers that are not fully connected. Instead, processing layers are provided that are only partially connected to other processing layers. In particular, each node in a convolutional layer is only connected to a local 3D region of the processing layer, receives input from that local 3D region, and the nodes over that local 3D region perform a convolution with respect to a filter. The nodes to which a particular node is connected are said to be within the "receptive field" of that filter. A filter is defined by a set of filter weights, and the convolution at each node is a weighted sum (weighted according to the filter weights) of the outputs of the nodes within the receptive field of the filter. The local partial connection from one layer to the (x,y) positions of the values within its corresponding tensor such that the (x,y) position information is at least to some extent retained in the CNN as the data passes through the network.
[0472] Each feature map is determined by convolving a given filter over an input tensor. Thus, the depth (range in the z - direction) of each convolutional layer equals the number of filters applied in that layer. The input tensor itself can be an image or a stack of feature maps, where the feature maps themselves are determined by convolution. When convolution is applied directly to an image, each filter acts as a low - level structure detector because "activation" (i.e., a relatively large output value) occurs when the pixels within the receptive field of the filter form some structure (i.e., match the structure of a particular filter). However, when convolution is applied to a tensor that is itself the result of an earlier convolution in the network, each convolution is performed on a set of feature maps for different features, and thus the network is further activated when a particular combination of lower - level features is present within the receptive field. Thus, for each successive convolution, the network detects the presence of increasingly high - level structural features corresponding to particular combinations of features from the previous convolution. Thus, in the early layers, the network effectively performs lower - level structure detection but gradually moves towards more high - level structure semantic understanding in the later layers. The filter weights are learned during training, and this is how the network learns what structures to look for. As is known in the art, convolution can be used in combination with other operations. For example, pooling (a form of dimensionality reduction) and non - linear transformations (such as ReLu, softmax, etc.) are typical operations used in combination with convolution in a CNN.
[0473] Figure 29 A highly schematic overview of a PSPM implemented as a neural network (net) or similar trainable function approximator is shown.
[0474] In this example, neural network A00 has an input layer A02 and an output layer A04. Although neural network A00 is schematically depicted as a simple feed - forward neural network, this is merely illustrative, and neural network A100 can take any form, including for example a Recurrent Neural Network (RNN) and / or a Convolutional Neural Network (CNN) architecture. The terms "input layer" and "output layer" do not imply any particular neural network architecture and include, for example, the input and output tensors in the case of a CNN.
[0475] At input layer A02, neural network A00 receives the sensed ground truth t as input. For example, the sensed ground truth t can be encoded as an input vector or tensor. In general, the sensed ground truth t can be related to any number of objects and any number of underlying sensor modalities.
[0476] Neural network A00 can be mathematically represented as a function as follows:
[0477] y = f(t; w)
[0478] where w is a set of adjustable weights (parameters) that processes the input t according to the set of adjustable weights (parameters). During training, the aim is to optimize the weights w with respect to some loss function defined on the output y.
[0479] In Figure 29 the example of, the output y is a set of distribution parameters that defines a predicted probability distribution p(e|t), where the predicted probability distribution p(e|t) is the probability of obtaining some predicted perceptual output e given the perceptual ground truth t at the input layer A02.
[0480] Taking the simple example of a Gaussian (normal) distribution, the output layer A04 can be configured to provide a predicted mean and variance for a given ground truth:
[0481] y = {μ(t; w), σ(t; w)}.
[0482] Note that either the mean or the variance can vary with the input ground truth t as defined by the learned weights w, giving the neural network A00 the flexibility to learn this dependence during training to the extent that it is reflected in the training data to which it is exposed.
[0483] During training, the aim is to learn the weights w that match p(e|t) to the actual perceptual output A06 generated by the perceptual slice 204 to be modeled. This means that, for example, by optimizing a suitable loss function A08 via gradient descent or ascent, the distribution p(e|t) predicted at the output layer for a given ground truth t can be meaningfully compared to the actual perceptual output corresponding to the ground truth t. As described above, the ground truth input t for training is provided by the ground truth (annotation) pipeline 802, which has been defined by manual, automatic, or semi-automatic annotation of the sensor data to which the perceptual slice 204 has been applied. The set of sensor data to which the perceptual slice 204 is applied can be referred to in the following description as input samples or equivalently as frames and is denoted by the reference numeral A01. The actual perceptual output for each frame A01 is calculated by applying the perceptual slice 204 to the sensor data of that frame. However, according to the teachings above, the neural network A00 is not exposed to the underlying sensor data during training but receives the annotated ground truth t of the frame A01 as the input that conveys the underlying scene.
[0484] Given a sufficient set of example {e, t} pairs, existing neural network architectures can be trained to predict the conditional distribution of the form p(e|t). For simple Gaussian distributions (univariate or multivariate), log-normal or (negative) log-PDF loss functions A08 can be used. One way to extend this to non-Gaussian distributions is to use a Gaussian mixture model, where a neural network A00 and the mixture coefficients for combining these (learned according to the input t in the same way as the means and variances of each Gaussian component) are trained together to predict a multi-component Gaussian distribution. Theoretically, any distribution can be represented as mixed Gaussians, so Gaussian mixture models are a useful way to approximate general distributions. References in this paper to "fitting Normal distributions", etc. include Gaussian mixture models. The related descriptions also apply more generally to other distribution parameterizations. As will be understood, there are various known techniques by which neural networks can be constructed and trained to predict conditional probability distributions given a sufficient representative set of input-output pairs. Therefore, further details are not described in this paper unless specifically relevant to the described embodiments.
[0485] At inference time, the trained network A00 is used as described above. The sensed ground truth t provided by the simulator 814 is provided to the neural network A00 of the input layer A02, which is processed by the neural network A00 to generate a predicted sensed output distribution of the form p(e|t) at the output layer A04, which can then be sampled by a sampling orchestration component (sampler) 816 in the manner described above.
[0486] It is important to note the terminology used in this paper. In this context, "ground truth" refers to the input of the neural network A00 from which its output is generated. In training, the ground truth input comes from the annotations, and at inference time, the ground truth is provided by the simulator 814.
[0487] Although the actual sensed output A06 can be considered in the form of ground truth in the training context - since it is an example of the type of output that the neural network is trained to replicate - this term is generally avoided in this paper to avoid confusion with the input to the PSPM. The sensed output generated by applying the sensed slice 204 to the sensor data is referred to as the "actual" or "target" sensed output. The aim of training is to adjust the weights w via a suitable loss function that optimizes the deviation between the measured network output and the target sensed output, so that the distribution parameters of the output layer A04 match the actual sensed output A06.
[0488] Figure 29It is not necessarily a complete representation of the input or output of neural network A00—it can take additional inputs on which the predictive distribution will depend and / or it can provide the outputs of other functions that take it as their input.
[0489] 5.1 Confounding factors
[0490] Figure 30 An extension of the neural network is shown to incorporate one or more confounding factors c according to the above principles. Confounding factors are easily incorporated into the architecture as they can simply be provided as additional inputs to the input layer A02 (during both training and inference), and thus, during training, neural network A00 can learn the dependence of the output distribution on the confounding factors c. That is, network A00 can learn the distribution p(e|t,c) at the output layer A04, where any parameters of the distribution (e.g., mean, standard deviation, and mixing coefficients) can depend not only on the ground truth t but also on the confounding factors c, to the extent that those dependencies are captured in the training data.
[0491] 5.2 Temporal dependence
[0492] Figure 31 Another extension is shown to incorporate explicit time dependency. In this case, the function (neural network) takes as input at the input layer A02:
[0493] · The current ground truth t t
[0494] · The ground truth t from the previous time step t-1
[0495] · The previous detection output e t-1
[0496] where the subscript t (non-bold, italic) denotes the time instance. The output is the distribution of the current sensed output p(e t |t t ,t t-1 ,e t-1 ), where the current sampled sensed output e t is obtained by sampling from this distribution.
[0497] Here, e t-1 is also obtained by sampling from the distribution predicted at the previous time step, and thus, the distribution predicted at the current step will depend on the output of the sampler 816 at the previous step.
[0498] Figure 31 And Figure 32 the implementations of
[0499] One way to achieve the above is to model the characteristics of each detected object for the perceived ground truth t and the sampled perception output e respectively. For example, these characteristics may include Position, Extent, Orientation, and Type.
[0500] The output layer A04 of the neural network is used to predict the transformed real valued variable, and then parameterize the probability distribution of the variable of interest. Conceptually, this form of neural network models the perception slice 204 as a stochastic function.
[0501] Epistemic uncertainty motivates the stochastic modeling of the perception slice 204: even if the perception slice 204 is deterministic, it exhibits significant randomness, which stems from the lack of knowledge of many unknown variables that will affect its output in practice.
[0502] Typical scenarios may include multiple perception objects. Note that e and t in this article are general symbols that can represent the sampled perception output / perceived ground truth of a single object or multiple objects.
[0503] Another challenge mentioned above is to model false positives (FPs, i.e., false positive detections of objects) and false negatives (FNs, i.e., failures to detect objects). The impact of FPs and / or FNs is that the number of ground truth objects (i.e., the number of objects providing the perceived ground truth) does not necessarily match the number of predicted objects (i.e., the number of objects for which real perception output samples are provided).
[0504] A distinction can be made between the "single object" approach and the "set-to-set approach". In the broadest sense, the single object PSPM relies on a clear one-to-one association between the ground truth object and the predicted object. In terms of the single object ground truth it is associated with, the simplest way to implement the single object PSPM is to consider each object independently during training. At inference time, the PSPM receives the perceived ground truth for a single object and provides a single object perception output. False negatives can be directly accommodated by introducing some mechanisms that can model the failure detection of a single object.
[0505] 5.3 Single-object PSPMs
[0506] An example implementation of the single object PSPM using a neural network will now be described. The normal distribution is fitted to the position and extent variables (these can be multivariate normal if required).
[0507] For the modeling direction, the method of Section 3.2.2 of “Probabilistic Regression of Rotations using Quaternion Averaging and a Deep Multi-Headed Network” by Peretroukhin et al. [https: / / arxiv.org / pdf / 1904.03182.pdf] can be followed — the entire content of which is incorporated herein by reference. In this method, a quaternion representation of the orientation is used. Noise is injected into the tangent space around the quaternion and may be mixed with the surrounding quaternions where the noise is injected.
[0508] Model the fake “negativeness” using a Bernoulli random variable.
[0509] Given the final network layer A04, the individual variable distributions may be conditionally independent, but dependencies / correlations can be induced by feeding noise as an additional input into the neural network to form a stochastic likelihood function. This is actually a mixture distribution.
[0510] The neural network A00 is trained using stochastic gradient descent (maximum likelihood — using the negative log pdf of the random variable as the location and scale variables (which can be multivariate normal if appropriate)).
[0511] The single object method requires a clear association to be established between the ground truth object and the actual perception output A06. This is because the predicted distribution for a given object needs to match the appropriate single object perception output actually produced by the perception slice 204. Identifying and encoding these associations for the purpose of PSPM training can be implemented as an additional step within the ground truth pipeline 802.
[0512] 5.4 Set-to-set method
[0513] In the broadest sense, the set-to-set method is a method that does not rely on a clear association between the ground truth object and the predicted object, i.e., during training, the PSPM does not need to be told which ground truth object corresponds to which predicted object.
[0514] Figure 32Shows the set-to-set PSPM D00, which takes the perceived ground truth {t0, t1} as input for a set of ground truth objects of any size (two ground truth objects in this example, indexed 0 and 1), and provides a realistic perceived output or distribution {e0, e1} for the set of predicted perceived objects (also two in this example - but note the discussion of FPs and FNs below).
[0515] The set-to-set approach has various benefits.
[0516] The primary benefit is a reduced annotation burden - there is no need to determine the association between the ground truth and the actual perceived targets for training purposes.
[0517] Another benefit is the ability to model the correlations between objects. In the Figure 32 example, the set-to-set neural network PRISM is shown. The set-to-set neural network PRISM takes the perceived ground truth obtained for any number of input objects at its input layer and outputs the distribution for each predicted object. Notably, the architecture of the network is such that the predicted perceived output distribution p(e m |t0, t1) for any given predicted object m can generally depend on the perceived ground truth outputs for all ground truth objects (t0, t1 in this example). More precisely, the architecture is flexible enough to be able to learn these dependencies to the extent that they are reflected in the training data. The set-to-set approach can also learn the degree to which the perceived slice 204 provides overlapping bounding boxes, and any tendency for it to "swap" objects, which are further examples of learnable object correlations.
[0518] More generally, the result of the set-to-set approach is to consider the joint distribution of all detections at once, i.e., p(e1, e2,...|t0, t1) (when the e m are independent of each other, it reduces only to the product of each p(e m |t0, t1) mentioned in the previous paragraph). The advantage of doing this is the ability to model the correlations between detections. For example, e1 can have instance identifier 0, and e2 can too, but not simultaneously. Although the previous paragraphs and Figure 32 assume that the e m are independent, this is not generally required - the output layer can alternatively be configured to more generally represent the joint distribution p(e1, e2|t0, t1).
[0519] Another benefit is the ability to model false positives using certain set-to-set architectures. This is because the number of ground truth objects does not necessarily have to match the number of predicted perceived objects - the set-to-set architecture is viable where the latter is less than, equal to, or greater than the former, depending on the input to the network.
[0520] 5.5 CNN Set-to-Set Architecture
[0521] For example, now refer to Figures 33A - 33D the set-to-set CNN architecture is described. The following is assumed that the actual perception output provided by the perception slice 204 includes 3D bounding boxes for any detected object, with defined position, orientation, and extent (size / dimensions). The CNN uses an input tensor and produces an output tensor, which is constructed as described below.
[0522] The CNN PSPM D00 jointly models the output detections based on all ground truth detections in a specific frame. An RNN architecture can be used to induce temporal dependence, which is a way to implement explicit temporal dependence on previous frames.
[0523] The ground truth t and the output prediction are spatially encoded in the "PIXOR" format. Briefly, the PIXOR format allows for efficient encoding of 3D spatial data based on a top-down (bird's-eye) view. For details, see: Yang et al.'s "PIXOR: Real-time 3D Object Detection from Point Clouds" [https: / / arxiv.org / abs / 1902.06326], the entire content of which is incorporated herein by reference.
[0524] As Figure 33A shown, to represent the actual perception output A06 for training purposes, a low-resolution (e.g., 800px square) bird's-eye view (classification layer D12, or more generally an object map) of the actual perception 3D bounding boxes is generated. The output objects are drawn in only one "color" - i.e., a classification detection image with a binary encoding of "detection-ness" (the "detected" pixels for a given object form the object region). This can be repeated for the ground truth objects, or generalized by drawing the colors of the input objects according to their occlusion status (one-hot encoding).
[0525] To spatially encode other characteristics of the objects, more bird's-eye view images are generated, which represent the position, extent, orientation, and any other important variables of the vehicles present in each pixel. These further images are referred to as regression layers or perception layers and are denoted by the reference numeral D14. This means that a single detection is represented multiple times in adjacent pixels and that some information is redundant, as shown. The images are stacked to produce a tensor of size (height x, width x, number of important variables (HEIGHT X WIDTH X NUMBER OF IMPORTANT VARIABLES)).
[0526] Note that it is the perception layer that encodes the 3D bounding box and there is redundancy. When "decoding" the output tensor, the actual numerical values of the regression layer D14 define the position, orientation, and extent of the bounding box. The purpose of the spatial encoding in the bird's-eye view images is to provide the information encoded within the perception layer D14 of the input tensor in a form that is conducive to CNN interpretation.
[0527] One advantage of this model is that it can learn the correlations between different object detections - for example, it can learn if the stack does not predict overlapping objects. Additionally, PSPM can learn if the stack is swapping object IDs between objects.
[0528] By feeding in additional input images, the CNN can be encouraged to predict false positives in physically meaningful places, as this provides the information needed for the CNN to determine the correlation between false positives and the map during training. The images can be, for example, a map of the scene (indicating the environmental structure).
[0529] The CNN can also have inputs to receive confounding factors in any suitable form. Object-specific confounding factors can be spatially encoded in the same way. Examples of such confounding factors include occlusion values, i.e., the degree to which an object is occluded by other objects and / or measurements of truncation (the degree to which an object is outside the sensor's field of view).
[0530] Figure 33B The training of the CNN is illustrated schematically.
[0531] The CNN D00 is trained to predict an output tensor D22 from an input tensor D20 using the classification (e.g., cross entropy) loss D32 of the classification layer and the regression (e.g., smooth L1) loss D34 of the regression layer of those tensors. The classification and regression layers of the input and output tensors D20, D22 are shown separately only for clarity. Typically, the information can be encoded in one or more tensors.
[0532] The regression layer of the input tensor D20 encodes the perceived ground truth t of the current frame for any number of ground truth objects.
[0533] The classification loss D32 is defined with respect to the target classification image D24A derived from the actual perceived output e of the current frame. The regression loss is defined with respect to the target perception layer D24B, which spatially encodes the actual perceived output of the current frame.
[0534] Each pixel of the classification layer of the output tensor D22 encodes the probability of detecting an object at that pixel (the probability of "detectability"). The corresponding pixels of the regression layer define the corresponding object position, extent, and orientation.
[0535] The classification layer of the output tensor D22 is thresholded to produce a binary output classification image. During training, the binary output classification image D23 is used to mask the regression layer, i.e., the regression loss only considers regions of the regression layer where an object exists within the threshold image D23 and ignores regions of the regression layer outside this.
[0536] Figure 33C Shows how the trained network is applied at test time or inference time.
[0537] At inference time, the input tensor D20 now encodes the perceived ground truth t provided by the simulator 814 (for any number of objects).
[0538] The classification layer on the output image is thresholded and used to mask the regression layer of the output tensor D22. Unlabeled pixels within the object regions of the threshold image contain perceived values, which can then be considered detections.
[0539] The predicted perceived 3D bounding boxes are decoded from the masked regression layers of the output tensor D22.
[0540] Recall that for any given pixel, the numerical value of that pixel in the regression layer defines the extent, position, and orientation of the bounding box, so the predicted 3D bounding box for each unmasked pixel can be obtained directly. As shown, this typically results in a large number of overlapping boxes (proposed boxes) because each pixel within an object region is activated by the binary image (i.e., as a valid bounding box proposal).
[0541] Apply non-maximal suppression (NMS) to the decoded bounding boxes to ensure that an object is not detected multiple times. As is well known, NMS provides a systematic way to discard proposed boxes based on the confidence of the boxes and their degree of overlap with other boxes. In this case, for a box corresponding to any given pixel of the output tensor D22, the detection probability of that pixel from the (non-thresholded) classification layer can be used as the confidence.
[0542] As an alternative, non-maximal suppression can be avoided by selecting an output classification image that only activates the center location of the object. Thus, only one detection will be obtained for each object, and NMS is not required. This can be combined with random likelihood (feeding noise as an additional input into the neural network) to mitigate the effect of activating the output classification image only at the center location of the object.
[0543] In addition to other losses, a GAN (generative adversarial network) can be used to obtain a more realistic network output.
[0544] The simple example described above does not provide a probability distribution for the output layer - that is, the probability distribution of the output layer is a one-to-one mapping between the perceived ground truth t and the set of predicted perceived outputs e directly encoded in the output tensor (the network is deterministic in this sense). This can be interpreted as the "average" response of the perceptual slice 204 given the ground truth t.
[0545] However, as Figure 33D shown, the architecture can be extended to predict the distribution at the output tensor D22, thus applying the exact same principle as described above with reference to Figure 29 In this case, the perceived values of the output tensor D22 encode distribution parameters, and the L1 regression loss D34 is replaced by a log PDF loss or other loss suitable for learning conditional distributions.
[0546] Another option is to train an ensemble of deterministic neural networks in the same way but on different subsets of the training data. In the case of M neural networks trained in this way, combining those networks will directly provide sampled perceived outputs (a total of M samples for each ground truth t). Using a sufficient number of appropriately configured deterministic networks, the distribution of their output samples can capture the statistical properties of the modeled perceptual slice 204 in a manner similar to a learned parametric distribution.
[0547] 5.6 Modeling Online Error Estimation
[0548] Figure 34Shows further expansion to accommodate online error (e.g., covariance) estimation within the perception slice 816 to be modeled. The online error estimator 816U within the stack provides an error estimate (or set of error estimates) associated with its perception output. Online error estimation is an estimate within the prediction system 816 of the error associated with its output. Note that this is an estimate by the prediction slice itself (potentially defective) of its output uncertainty, typically generated in real time using only information available on the runtime vehicle. This itself can be in error.
[0549] Such error estimation is important, for example, in the context of filtering or fusion, where multiple perception outputs (e.g., from different sensor modalities) can be fused in a way that respects their relative uncertainty levels. Incorrect online covariance can lead to fusion errors. Online error estimation can also be fed directly into prediction 104 and / or planning 106, for example, where planning is based on probabilistic prediction. Thus, errors in online error estimation can have a significant impact on stack performance and, in the worst case, may lead to unsafe decisions (especially if the error level of a given perception output is underestimated).
[0550] The method of modeling online covariance (or other online error estimation) is different from position, range, and orientation because there is no available ground truth covariance, i.e., the ground truth input t does not include any ground truth covariance.
[0551] Therefore, the only change is to add additional distribution parameters in the output layer A04 to additionally model the distribution p(E|t), i.e., the probability that the online error estimator 816U provides an error estimate of E given the perception ground truth t.
[0552] Note that this also treats the online error estimator 816U as a random function. Without loss of generality, this can be referred to as the neural network A00 learning the "covariance of the covariance". Modeling the online error estimation component 816U in this way can accommodate epistemic uncertainty about the online error estimator 816U in the same way as other such uncertainties regarding the perception system 204. This is particularly useful if the input to the online error estimator 816U is difficult to simulate or expensive to simulate. For example, if the online error estimator 816U is applied directly to sensor data, this would be a way to model the online error estimator 816 without having to simulate those sensor data inputs.
[0553] The covariance is fitted by performing a Cholesky decomposition E02 on the covariance matrix to produce a triangular matrix. This results in positive diagonal elements, allowing the logarithm of the diagonal elements to be computed. Then, a normal distribution can be fitted to each component of the matrix (a multivariate normal distribution can be used if desired). At test time, the process is inverted to produce the desired covariance matrix (the lower triangular scale matrices are multiplied together). This allows the loss function to be formulated as a straightforward numerical regression loss on the unconstrained space of the Cholesky decomposition. To "decode" the neural network, the inverse transformation can be applied.
[0554] At inference time, p(E|t) can be sampled in the same way as p(e|t) to obtain a realistic sampled online error estimate.
[0555] Figures 29 - 34 All architectures are depicted in. For example, time and / or confounding factor dependencies can be incorporated into Figure 34 the model such that the covariance of the covariance depends on one or both.
[0556] More generally, the network can be configured to learn a joint distribution of the form P(e,E|t), which reduces to the above when e and E are independent of each other (but both depend on the ground truth t).
[0557] 6. PSPM Applications
[0558] PSPM has many useful applications, some of which will now be described.
[0559] 6.1 Planning Under Uncertainty
[0560] The use cases listed above test planning under uncertainty. This means testing how the planner 106 performs in the presence of statistically representative perception errors. In that case, the benefit is being able to expose the planner 106 and the prediction stack 104 to realistic perception errors in a robust and efficient manner.
[0561] One benefit of the confounding factor approach is that when an instance of an unsafe behavior occurs in a particular scenario, the contribution of any confounding factor to that behavior can be explored by running the same scenario but using a different confounding factor c (which may have the effect of changing the perception uncertainty p(e|t,c)).
[0562] As previously mentioned, when sampling from PSPM, it is not necessary to sample in a uniform manner. It may be beneficial to deliberately bias the sampling towards outliers (i.e., PSPM samples with lower probabilities).
[0563] The manner of incorporating confounding factor c also facilitates testing more challenging scenarios. For example, if it is observed through simulation that planner 106 makes relatively more errors in the presence of occlusion, this may be a trigger to test more scenarios where external objects are occluded.
[0564] 6.2 Separating Perception and Planning / Prediction Errors
[0565] Another somewhat related but still independent application is the ability to isolate the reasons for the unsafe decisions of planner 106 in the runtime stack 100. In particular, it provides a convenient mechanism to infer whether the reason is a perception error rather than a prediction / planning error.
[0566] For example, consider a simulated scenario of an instance of unsafe behavior. This unsafe behavior may be caused by a perception error, but equally it may be caused by a prediction or planning error. To help isolate the reason, the same scenario can be run without PSPM, i.e., directly on the perfect perception ground truth, to see how planner 106 performs in the exact same scenario but with perfect perception output. If the unsafe behavior still occurs, this indicates that the unsafe behavior is at least partially attributable to errors outside the perception stack 102, which may indicate prediction and / or planning errors.
[0567] 6.3 Training
[0568] Simulation can also be used as a basis for training (e.g., reinforcement learning training). For example, simulation can be used as a training basis for components within the prediction stack 104, planner 106, or controller 108. In some cases, it may be beneficial to run training simulations based on the realistic perception output provided by PSPM.
[0569] 6.4 Testing Different Sensor Arrangements
[0570] One possible advantage of the PSPM method is the ability to simulate sensor types / locations that have not been actually tested. This can be used to make reasonable inferences about the impact of a group-specific set of sensors on a moving AV or the use of different types of sensors.
[0571] For example, a relatively simple way to test the impact of reducing the pixel resolution of an in-vehicle camera is to reduce the pixel resolution of the annotated images in the annotated ground truth database 804, reconstruct the PSPM, and re-run the appropriate simulations. As another example, the simulations can be re-run with a particular sensor modality (e.g., LiDAR) completely removed to test the possible impact.
[0572] As a more complex example, the impact of changing a particular sensor on the perception uncertainty can be inferred. This is unlikely to be used as a basis for proving safety, but can be used as a useful tool when considering, for example, camera placement.
[0573] 6.6 PSPM for Simulated Sensor Data
[0574] While the above considered the PSPM generated by applying the perception slice 204 to real sensor data, the actual perception output used to train the PSPM can alternatively be derived by applying the perception slice 204 to simulated sensor data in order to model the performance of the perception slice 204 on the simulated sensor data. Note that the trained PSPM does not require simulated sensor data—it still applies to the perception ground truth without the need for simulated sensor input. The simulated sensor data is only used to generate the actual perception output for training. This can be used as a way to model the performance of the perception slice 204 on simulated data.
[0575] 6.7 Online Application
[0576] Certain PSPMs can also be effectively deployed on an AV at runtime. That is, as part of the runtime stack 100 itself. This can in turn ultimately help the planner 106 take into account the knowledge of the perception uncertainty. The PSPM can be used in conjunction with existing online uncertainty models that form the basis of filtering / fusion.
[0577] Since the PSPM depends on confounding factors, in order to maximize the usefulness of the PSPM at runtime, the relevant confounding factors need to be measured in real time. This may not be applicable to all types of confounding factors, but when the appropriate confounding factors are measurable, the PSPM can still be effectively deployed.
[0578] For example, the uncertainty estimate made by the PSPM can be used as a prior in combination with an independent measurement of the uncertainty from one of the AV online uncertainty models at runtime. Collectively, these can provide a more reliable indication of the actual perception uncertainty.
[0579] Structural perception refers to a class of data processing algorithms that can meaningfully interpret the structures captured in perceptual inputs (sensor outputs or perceptual outputs from lower-level perceptual components). This processing can be applied across different forms of perceptual inputs. Perceptual inputs generally refer to any structural representation, i.e., any dataset that captures a structure. Structural perception can be applied in two-dimensional (2D) and three-dimensional (3D) spaces. The result of applying a structural perception algorithm to a given structural input is encoded as a structural perception output.
[0580] One form of perceptual input is a two-dimensional (2D) image; that is, an image with only color components (one or more color channels). The most basic form of structural perception is image classification, which simply classifies an image as a whole relative to a set of image categories. More complex forms of structural perception applied to 2D space include 2D object detection and / or localization (e.g., direction, pose, and / or distance estimation in 2D space), 2D instance segmentation, etc. Other forms of perceptual inputs include three-dimensional (3D) images, i.e., images with at least one depth component (depth channel); for example, 3D point clouds captured using radar or lidar or derived from 3D images; voxel- or mesh-based structural representations, or any other form of 3D structural representation. Examples of perceptual algorithms that can be applied in 3D space include 3D object detection and / or localization (e.g., distance, direction, or pose estimation in 3D space), etc. A single perceptual input can also be formed by multiple images. For example, stereo depth information can be captured in a stereo pair of 2D images, and this image pair can be used as the basis for 3D perception. 3D structural perception can also be applied to a single 2D image, an example being monocular depth extraction, which extracts depth information from a single 2D image (note that a 2D image can capture a certain degree of depth information within one or more of its color channels without any depth channel). This form of structural perception is an example of different "perception modalities" in the terms used herein. Structural perception applied to 2D or 3D images can be referred to as "computer vision".
[0581] Object detection refers to detecting any number of objects captured in a perceptual input and generally involves characterizing each such object as an instance of an object class. Such object detection may involve or be performed in combination with one or more forms of position estimation, e.g., 2D or 3D bounding box detection (a form of object localization where the aim is to define a region or volume in 2D or 3D space that encloses the object), distance estimation, pose estimation, etc.
[0582] In the case of machine learning (ML), the structure-aware component can include one or more trained perception models. For example, machine vision processing often uses convolutional neural networks (CNNs) to implement. Such networks require a large number of training images, which have been annotated with the information that the neural network needs to learn (in the form of supervised learning). During training, the network is presented with thousands, or preferably hundreds of thousands, of such annotated images, and learns on its own how the features captured in the images themselves relate to the associated annotations. Each image is annotated in the sense of being associated with annotation data. The image serves as the perception input, and the associated annotation data provides the "ground truth" for the image. CNNs and other forms of perception models can be constructed to receive and process other forms of perception input (e.g., point clouds, voxel tensors, etc.) and perceive structures in 2D and 3D spaces. In the case of general training, the perception input can be referred to as a "training example" or a "training input". In contrast, the training examples captured at runtime for processing by the trained perception component can be referred to as "runtime inputs". The annotation data associated with the training input provides the ground truth for that training input, as the annotation data encodes the expected perception output for that training input. During supervised training, the parameters of the perception component are systematically adjusted to minimize the overall measure of the difference between the perception output ("actual" perception output) produced by the perception component when applied to the training examples in the training set and the corresponding ground truth (expected perception output) provided by the associated annotation data within a defined range. In this way, the perception input "learns" from the training examples and is also able to "generalize" that learning, in the sense that it can be trained to provide a meaningful perception output for perception inputs that it did not encounter during training.
[0583] Such sensing components are the cornerstone of many mature and emerging technologies. For example, in the field of robotics, mobile robot systems capable of autonomously planning paths in complex environments are becoming increasingly common. An example of such a rapidly emerging technology is an autonomous vehicle (AV) that can navigate itself on urban roads. Such vehicles must not only perform complex maneuvers between people and other vehicles, but they must do so frequently while ensuring that the probability of adverse events occurring, such as collisions with these other entities in the environment, is strictly limited. For an AV to plan safely, it is crucial that it can observe its environment accurately and reliably. This includes the need to accurately and reliably detect real-world structures near the vehicle. An autonomous vehicle, also known as a driverless vehicle, is a vehicle that has a sensor system for monitoring its external environment and a control system that can automatically make and execute driving decisions using these sensors. This particularly includes the ability to automatically adapt the vehicle speed and driving direction based on the perceptual input from the sensor system. A fully autonomous or "driverless" vehicle has sufficient decision-making capabilities to operate without any input from a human driver. However, the term "autonomous vehicle" as used in this article also applies to semi-autonomous vehicles, which have more limited autonomous decision-making capabilities and thus still require a certain degree of supervision from a human driver. Other mobile robots are being developed, for example, to transport goods within and outside industrial areas. Such mobile robots will have no one on board and belong to a class of mobile robots known as UAVs (unmanned autonomous vehicles). Autonomous aerial mobile robots (drones) are also being developed.
[0584] Thus, in the more general context of autonomous driving and robotics, one or more sensing components may be required to interpret the perceptual input, i.e., one or more sensing components can determine the information about real-world structures captured in a given perceptual input.
[0585] Increasingly complex robotic systems (e.g., AVs) may need to implement multiple sensing modalities to accurately interpret multiple forms of perceptual input. For example, an AV can be equipped with one or more pairs of stereo optical sensors (cameras) from which relevant depth maps are extracted. In this case, the AV's data processing system can be configured to apply one or more forms of 2D structure perception to the image itself—e.g., 2D bounding box detection and / or other forms of 2D localization, instance segmentation, etc.—plus apply one or more forms of 3D structure perception to the data of the relevant depth maps—e.g., 3D bounding box detection and / or other forms of 3D localization. Such depth maps can also come from LiDAR, RADAR, etc., or be derived by combining multiple sensor modalities.
[0586] This technology can be used to simulate the behavior of various robotic systems for purposes such as testing / training. The runtime application can also be implemented in different robotic systems.
[0587] To train a perception component for a desired perception modality, the perception component is constructed such that it can receive perception inputs in a desired form and provide perception outputs in a desired form in response. Additionally, to train a perception component with an appropriate architecture based on supervised learning, annotations conforming to the required perception modality need to be provided. For example, to train a 2D bounding box detector, 2D bounding box annotations are required; similarly, to train a segmentation component that performs image segmentation (per-pixel classification of individual image pixels), the annotations need to encode suitable segmentation masks from which the model can learn; a 3D bounding box detector needs to be able to receive 3D structure data, as well as annotated 3D bounding boxes, etc.
[0588] A perception component can refer to any tangible embodiment (instance) of one or more underlying perception models of the perception component, which can be a software or hardware instance, or a combined software and hardware instance. Such an instance can be embodied using programmable hardware such as a general-purpose processor (e.g., a CPU, an accelerator such as a GPU, etc.) or a field programmable gate array (FPGA) or any other form of programmable computer. Thus, a computer program for programming a computer can take the form of program instructions executed on a general-purpose processor, circuit description code for programming an FPGA, etc. Instances of the perception component can also be implemented using non-programmable hardware, for example, an application specific integrated circuit (ASIC), and such hardware can be referred to herein as a non-programmable computer. Generally speaking, a perception component can be embodied in one or more computers, where the one or more computers can be programmable or not, and the one or more computers are programmed or otherwise configured to execute the perception component.
[0589] Refer to Figure 8 , the depicted pipeline component is a functional component of a computer system, which can be implemented in various ways at the hardware level: Although Figure 8Although not shown, the computer system includes one or more processors (computers), and one or more processors execute the functions of the above components. The processor may be in the form of a general-purpose processor (e.g., a CPU (Central Processing Unit) or an accelerator (e.g., a GPU), etc.), or in the form of a more specialized hardware processor (e.g., an FPGA (Field Programmable Gate Array), or an ASIC (Application-Specific Integrated Circuit)). Although not shown separately, the UI generally includes at least one display and at least one user input device for receiving user input to allow the user to interact with the system, such as a mouse / touchpad, a touch screen, a keyboard, etc.
[0590] The various aspects and exemplary embodiments of the present invention have been described above. Other aspects and exemplary embodiments of the present invention are described below.
[0591] On the other hand, a method for testing the performance of a robotic planner and a perception system is provided, the method comprising:
[0592] receiving at least one probability uncertainty distribution for modeling at least one perception component of the perception system, such as determining at least one probability uncertainty distribution based on a statistical analysis of the actual perception output derived by applying at least one perception component to the input directly or indirectly obtained from one or more sensor components; and
[0593] running a simulation scenario in a simulator, wherein the simulated robot state changes according to the autonomous decision made by the robotic planner based on the realistic perception output calculated for each simulation scenario;
[0594] wherein the realistic perception output models the actual perception output to be provided by at least one perception component in the simulation scenario, but is calculated without applying at least one perception component to the simulation scenario and without simulating one or more sensor components, but by:
[0595] (i) directly calculating the perception ground truth of at least one perception component based on the simulation scenario and the simulated robot state, and
[0596] (ii) modifying the perception ground truth according to at least one probability uncertainty distribution to calculate the realistic perception output.
[0597] Note that the terms "perception pipeline", "perception stack", and "perception system" are used synonymously herein. The term "perception slice" is used to refer to all or part of a perception stack (including one or more perception components) modeled by a single PSPM. As described later, during simulation safety testing, the perception stack can be replaced in whole or in part by one or more PSPMs. The term slice can also be used to refer to the part of the prediction stack that is not modeled by a PSPM or is replaced by a PSPM, and the meaning will be clear from the context.
[0598] In a preferred embodiment of the present invention, the realistic perception output depends not only on the perception ground truth but also on one or more "confounders". That is, the influence of the confounders on the perception output modeled by the PSPM. Confounders represent real-world conditions that may affect the accuracy of the perception output (e.g., weather, lighting, the speed of another vehicle, the distance to another vehicle, etc.; examples of other types of confounders will be given later). It is said that the PSPM is mapped to a "confounder space" that represents all possible confounders or combinations of confounders that the PSPM can consider. This allows the PSPM to accurately model different real-world conditions represented by different points in the confounder space in an efficient manner, because the PSPM eliminates the need to simulate sensor data for those different conditions and does not require the application of the perception components themselves as part of the simulation.
[0599] The term "confounder" is sometimes used in statistics to refer to a variable that has a causal effect on both the dependent variable and the independent variable. However, in this document, the term is used in a more general sense to represent a variable of a perception error model (PSPM) that represents some physical condition.
[0600] In an embodiment, at least one probability uncertainty distribution can be used to model multiple collaborative perception components of the perception system.
[0601] In an embodiment, only a part of the perception system can be modeled, and at least a second perception component of the perception system can be applied to the realistic perception output to provide a second perception output for making the decision.
[0602] The second perception component can be a fusion component, such as a Bayesian filter or a non-Bayesian filter.
[0603] The modelled perception component can be a sensor data processing component that is highly sensitive to artefacts in the simulated data. In this case, the above method avoids the need to simulate high-quality sensor data for this component. For example, the perception component can be a Convolutional Neural Network (CNN) or other form of neural network.
[0604] Alternatively or additionally, the modelled perception component can be a sensor data processing component that processes sensor data that is inherently difficult to simulate. For example, a RADAR processing component.
[0605] The method can include the step of analyzing changes in the simulated robot state to detect instances of unsafe behavior of the simulated robot state and determining the cause of the unsafe behavior.
[0606] Instances of unsafe behavior can be detected based on a predefined set of acceptable behavior rules applied to the simulated scenario and the simulated robot state.
[0607] The rules for such acceptable behavior can take the form of a 'Digital Highway Code' (DHC).
[0608] The PSPM combined with the DHC allows for many realistic simulations to be efficiently run without knowing which will lead to unsafe / unacceptable behavior (as opposed to movement variations in scenarios known to be unsafe from real-world test drives), where the predefined rules of the DHC are used to automatically detect instances of such behavior.
[0609] The perception component and / or the planner can be modified to mitigate the cause of the unsafe behavior.
[0610] The sensor output obtained from one or more sensors and the corresponding perception ground truth associated with the sensor output can be used to determine a probability uncertainty distribution.
[0611] The probability uncertainty distribution can vary with one or more confounding factors, where a set of one or more confounding factors selected for the simulated scenario can be used to modify the perception ground truth according to the probability uncertainty distribution, where each confounding factor represents a physical property.
[0612] The above one or more confounding factors can include one or more of the following:
[0613] - Occlusion level
[0614] - One or more lighting conditions
[0615] - Indication of the time of day
[0616] - One or more weather conditions
[0617] - Indication of season
[0618] - Physical characteristics of at least one external object
[0619] - Sensor conditions (e.g., position of object in field of view)
[0620] In a time-dependent model, another variable input on which the PSPM depends can be the previous ground truth calculated therefrom and / or at least one previous real-world perception output.
[0621] The simulated scenario can be derived from an observed real scenario.
[0622] The simulated scenario can be a fuzzed scenario determined by fuzzing an observed real-world scenario.
[0623] That is, as a result of the perception error, in addition to causing changes in the inputs that occur in the prediction and planning system, it can also be combined with a method of generating additional test scenarios by making (small or large) changes to the situation of the test scenario (e.g., other vehicles that accelerate or decelerate slightly in the scenario, e.g., slightly changing the initial position and orientation of the ego car and other vehicles, etc.). These two types of changes to the known real scenario will together have a higher chance of hitting a dangerous situation, and the system needs to be able to handle it.
[0624] On the other hand, there is provided a computer-implemented method for training a perception statistical performance model (PSPM), wherein the PSPM models the uncertainty in the perception output calculated from perception slices, and the method includes:
[0625] Applying perception slices to a plurality of training sensor outputs to calculate a training perception output for each sensor output, wherein each training sensor output is associated with a perception ground truth;
[0626] Comparing each perception output with the associated perception ground truth to calculate a set of perception errors Δ;
[0627] Using the set of perception errors Δ to train the PSPM, wherein the trained PSPM provides a probability perception uncertainty distribution of the form p(e|t), where p(e|t) represents the probability that the perception slice calculates a specific perception output e given the perception ground truth t.
[0628] On the other hand, there is provided a perception statistical performance model (PSPM) implemented in a computer system, the PSPM being used to model perception slices and configured to:
[0629] Receive the calculated perception ground truth t;
[0630] Determine a probabilistic perception uncertainty distribution of the form p(e|t) from the sensed ground truth t based on a set of learning parameters θ, where p(e|t) represents the probability of computing a particular sensed output e for a sensed slice given the computed sensed ground truth t, and the probabilistic perception uncertainty distribution is defined over a range of possible sensed outputs, and the parameter θ is learned from a set of actual sensed outputs generated using the sensed slices to be modeled.
[0631] In a preferred embodiment, the PSPM can vary according to one or more confounding factors c, where each confounding factor characterizes a physical condition. In this case, the probabilistic perception uncertainty distribution takes the form p(e|t,c).
[0632] In an embodiment, the PSPM can take the form of a parametric distribution defined by a set of parameters θ learned from a set of sensing errors Δ and varying with a given sensed ground truth t.
[0633] To train the PSPM according to confounding factors, each training sensed output can also be associated with a set of one or more confounding factors that characterize one or more physical conditions in which the training sensor outputs are captured.
[0634] The ground truth for training the PSPM can be generated offline because more accurate and thus generally more computationally intensive algorithms can be used compared to the online case. These only need to be generated once.
[0635] Note that the term parameter includes hyperparameters, e.g., hyperparameters learned by variational inference.
[0636] In an embodiment, the sensed ground truth associated with the sensor output can be derived from the sensor output using offline processing (e.g., processing that cannot be performed in real time due to hardware limitations or because offline processing is inherently non-real-time).
[0637] Models suitable for the PSPM typically draw attention to confounding factors in the data used for modeling, which may or may not be initially obvious. The advantage of this is that only significant confounding factors need to be modeled separately, and their importance is determined by the degree to which the data deviates from the model.
[0638] The confounding factor c is a variable on which the trained PSPM depends. At runtime, by changing the value of the confounding factor c, realistic sensed outputs (i.e., with realistic errors) can be obtained for different physical situations. The variable can be numerical (e.g., continuous / pseudo-continuous) or categorical (e.g., binary value or non-binary categorical value).
[0639] It is possible that the training of the PSPM reveals a statistically significant dependence on one or more physical characteristics that are not currently characterized by the existing confounding factor c. For example, when the trained PSPM is validated, its performance on certain types of data may be worse than expected, and the analysis can attribute this to a dependence on physical conditions that are not explicitly modeled in the PSPM.
[0640] Accordingly, in an embodiment, the method may include the steps of analyzing the trained PSPM for the confounding factor c (e.g., validating the PSPM using a validation-aware error dataset), and in response thereto, retraining the PSPM for a new set of one or more confounding factors c', whereby the probability-aware uncertainty distribution of the retrained PSPM takes the form p(e|t,c').
[0641] For example, c' can be determined by adding or removing the confounding factor c. For example, if a confounding factor is determined to be statistically significant, it can be added, or if the analysis indicates that a confounding factor is not actually statistically significant, it can be removed.
[0642] By modeling the PSPM in this way, it is possible to determine which confounding factors are statistically significant and need to be modeled, and which confounding factors are not statistically significant and do not need to be modeled.
[0643] For example, the one or more confounding factors described above may include one or more of the following:
[0644] - The occlusion level of at least one external object (indicating the degree of occlusion of the object relative to the subject. The external object can be a moving actor or a static object)
[0645] - One or more lighting conditions
[0646] - An indication of the time of day
[0647] - One or more weather conditions
[0648] - An indication of the season
[0649] - The physical characteristics of at least one external object (e.g., position / distance from the subject, speed / rate / acceleration relative to the subject, etc.)
[0650] - The position of the external object in the subject's field of view (e.g., the angle from the center of the image in the case of a camera)
[0651] Another aspect of the present disclosure provides a computer system for testing and / or training a runtime stack of a robotic system, the computer system including:
[0652] A simulator configured to run a simulation scenario, where a simulation subject interacts with one or more external objects;
[0653] A runtime stack including an input configured to receive a time series of perceptual outputs of each simulation scenario, a planner configured to make autonomous decisions based on the perceptual outputs, and a controller configured to generate a series of control signals to cause the simulation subject to execute the decisions as the simulation scenario progresses;
[0654] Wherein the computer system is configured to calculate each perceptual output of the time series by:
[0655] Calculating a perceptual ground truth based on the current state of the simulation scenario;
[0656] Applying the above PSPM to the perceptual ground truth to determine a probabilistic perceptual uncertainty distribution; and
[0657] Sampling a perceptual output from the probabilistic perceptual uncertainty distribution.
[0658] Preferably, the PSPM is applied to the perceptual ground truth and one or more sets of confounding factors associated with the simulation scenario.
[0659] Ray tracing can be used to calculate the perceptual ground truth for each external object.
[0660] Each external object can be a moving actor or a static object.
[0661] The same simulation scenario can be run multiple times.
[0662] The same simulation scenario can be run multiple times with different confounding factors.
[0663] The runtime stack may include a prediction stack configured to predict the behavior of external actors based on the perceptual outputs, wherein the controller may be configured to make decisions based on the predicted behavior.
[0664] The computer system may be configured to record details of each simulation scenario in a test database, where the details include the decisions made by the planner, the perceptual outputs on which those decisions are based, and the behavior of the simulation subject when executing those decisions.
[0665] The computer system may include a scenario evaluation component configured to analyze the behavior of the simulation subject in each simulation scenario related to a set of predetermined behavior rules in order to classify the behavior of the subject.
[0666] The analysis results of the scenario assessment component can be used to develop simulation strategies. For example, the scenario can be "fuzzified" (see below) based on the analysis results.
[0667] The behavior of the subject can be classified as safe or unsafe.
[0668] To model false negative detections, a probability-aware uncertainty distribution can provide the probability of successfully detecting a visible object, which is used to determine whether to provide an object detection output for that object. (In this case, a visible object means an object that is in the field of view of the subject's sensor in the simulation scenario but the subject may still fail to detect it).
[0669] A time-dependent PSPM (e.g., a hidden Markov model) can be used in any of the above.
[0670] In the case of modeling false negatives, a time-dependent PSPM can be used such that the probability of detecting a visible object depends on at least one earlier determination regarding whether to provide an object detection output for the visible object.
[0671] To model false positive detections, a probability uncertainty distribution can provide the probability of false object detections, which is used to determine whether to provide a perception output for a non-existent object.
[0672] Once the "ground truth" is determined, if the scenario is run in a loop without a PSPM, potential errors in the planner can be explored. This can be extended to automatically classify the data to indicate perception problems or planner problems.
[0673] In an embodiment, the simulation scenario in which the simulated subject exhibits unsafe behavior can be re-run without applying a PSPM but by directly providing the perception ground truth to the runtime stack.
[0674] Then, an analysis can be performed to determine whether the simulated subject still exhibits unsafe behavior in the re-run scenario.
[0675] Another aspect of the present invention provides a method for testing a robot planner that is used to make autonomous decisions using the perception output of at least one perception component, the method comprising:
[0676] Running a simulation scenario in a computer system, wherein the simulated robot state changes according to the autonomous decisions made by the robot planner using the realistic perception output calculated for each simulation scenario;
[0677] For each simulation scenario, determining an ontological representation of the simulation scenario and the simulated robot state; and
[0678] Apply a set of predefined acceptable behavior rules [e.g., DHC] to the ontology representation of each simulated scenario in order to record and flag violations of the predefined acceptable behavior rules in one or more simulated scenarios.
[0679] Another aspect of the present invention provides a computer-implemented method that includes steps of implementing any of the above programs, systems, or PSPM functions.
[0680] A further aspect provides a computer system that includes one or more computers programmed or otherwise configured to perform any of the functions disclosed herein, and one or more computer programs for programming the computer system to perform the functions.
[0681] It should be understood that the various embodiments of the present invention have been described by way of example only. The scope of the present invention is not limited by the described embodiments, but is defined only by the appended claims.
Claims
1. A computer system for testing and / or training a runtime stack of a robotic system, the computer system comprising: A simulator configured to run a simulation scenario, wherein a simulated agent interacts with one or more external objects; A planner of the runtime stack, the planner of the runtime stack being configured to make autonomous decisions for each simulation scenario based on a time series of perceptual outputs computed for the simulation scenario; and A controller of the runtime stack, the controller of the runtime stack being configured to generate a series of control signals to cause the simulated agent to execute the autonomous decisions as the simulation scenario progresses; Wherein the computer system is configured to compute each perceptual output by: Computing a perceptual ground truth based on the current state of the simulation scenario; Applying a perceptual statistical performance model PSPM to the perceptual ground truth to determine a probabilistic perceptual uncertainty distribution; and Sampling the perceptual output from the probabilistic perceptual uncertainty distribution; Wherein the PSPM is used to model a perceptual slice of the runtime stack and is configured to determine the probabilistic perceptual uncertainty distribution based on a set of parameters learned from a set of actual perceptual outputs generated using the perceptual slice to be modeled; Wherein the perceptual slice includes an online error estimator, and the computer system is configured to use the PSPM to obtain a predicted online error estimate of the perceptual output in response to the perceptual ground truth.
2. The computer system according to claim 1, wherein Sampling the predicted online error estimate from the probabilistic perceptual uncertainty distribution.
3. The computer system according to claim 2, wherein, The PSPM takes the form of a function approximator that receives a perceptual ground truth t and outputs parameters of a probabilistic perceptual uncertainty distribution, from which the perceptual output and the predicted online error estimate are sampled.
4. The computer system according to claim 3, wherein, The PSPM has a neural network architecture.
5. The computer system according to any one of the preceding claims, wherein, The PSPM is applied to the perceptual ground truth and one or more confounding factors associated with the simulation scenario, wherein each confounding factor is a variable of the PSPM, the value of the variable characterizing the physical conditions applicable to the simulation scenario, and the probabilistic perceptual uncertainty distribution depends on the variable, and the predicted online error estimate depends on the confounding factor.
6. The computer system according to claim 5, wherein, The one or more confounding factors include one or more of the following confounding factors, which at least partially determine the probabilistic uncertainty distribution from which the perceptual output is sampled: The occlusion level of at least one of the external objects; One or more lighting conditions; An indication of the time of day; One or more weather conditions; An indication of the season; The physical characteristics of at least one of the external objects; Sensor conditions, the position of at least one of the external objects in the sensor field of view of the agent; The number or density of the external objects; The distance between two of the external objects; The truncation level of at least one of the external objects; The type of at least one of the objects, and An indication of whether at least one of the external objects corresponds to any external object from an earlier moment of the simulation scenario.
7. The computer system according to any one of claims 1-4 and 6, wherein, the PSPM includes a time-dependent model such that the sampled sensed output sampled at the predicted online error estimate depends on at least one of the following: the earlier one of the sensed outputs sampled at the previous moment, and the earlier one of the sensed ground truths calculated for the previous moment.
8. The computer system according to any one of claims 1-4 and 6, comprising: a scenario evaluation component configured to evaluate the behavior of the simulated entity in each of the simulated scenarios by applying a set of predetermined rules.
9. The computer system according to claim 8, wherein, At least some of the predetermined rules are related to security, and the scenario evaluation component is configured to evaluate the security of the behavior of the simulated entity in each of the simulated scenarios.
10. The computer system according to any one of claims 1-4, 6, and 9, the computer system being configured to record details of each simulation scenario in a test database, wherein, The details include the decisions made by the planner, the sensed outputs on which those decisions are based, and the behavior of the simulated entity when executing those decisions.
11. The computer system according to any one of claims 1-4, 6, and 9, wherein, The sampling from the probabilistic sensed uncertainty distribution is non-uniform and biased towards lower-probability sensed outputs.
12. The computer system according to any one of claims 1-4, 6 and 9, comprising a scenario fuzzification component configured to generate at least one fuzzified scenario for running in the simulator by fuzzifying at least one existing scenario.
13. The computer system according to any one of claims 1-4, 6, and 9, wherein, To model false negative detections, the probabilistic sensed uncertainty distribution provides the probability of visible objects among successfully detected objects, which is used to determine whether to provide an object detection output for the object. The object is visible when it is within the sensor field of view of the simulated entity in the simulated scenario, so the detection of the visible object is not guaranteed.
14. The computer system according to any one of claims 1-4, 6 and 9, wherein, The sensed ground truth is calculated for the one or more external objects using ray tracing.
15. The computer system according to any one of claims 1-4, 6, and 9, wherein, At least one of the external objects is a moving actor, and the computer system includes a prediction stack of the runtime stack, the prediction stack of the runtime stack being configured to predict the behavior of the external actor based on the sensed output, and the planner being configured to make autonomous decisions according to the predicted behavior.
16. A computer-implemented method for performance testing a runtime stack of a robotic system, the method comprising: running a simulated scenario in a simulator, wherein a simulated entity interacts with one or more external objects, wherein a planner of the runtime stack makes autonomous decisions for the simulated scenario according to a time series of sensed outputs calculated for the simulated scenario, and a controller of the runtime stack generates a series of control signals to cause the simulated entity to execute the autonomous decisions as the simulated scenario progresses; wherein each sensed output is calculated by: calculating a sensed ground truth based on the current state of the simulated scenario; applying a probabilistic sensed performance model PSPM to the sensed ground truth to determine a probabilistic sensed uncertainty distribution; and sampling the sensed output from the probabilistic sensed uncertainty distribution; Wherein, the PSPM is used to model the perceptual slices of the runtime stack and determine the probabilistic perceptual uncertainty distribution based on a set of parameters learned from an actual set of perceptual outputs generated using the perceptual slices to be modeled; Wherein, the perceptual slice includes an online error estimator, and the PSPM is used to obtain an online error estimate of the prediction of the perceptual output in response to the perceptual ground truth.
17. A computer - implemented method for training a Perceptual Statistical Performance Model (PSPM), wherein, The PSPM models the uncertainty in the perceptual output calculated from the perceptual slices of the runtime stack of a robotic system, the method comprising: Applying the perceptual slice to a plurality of training sensor outputs, thereby calculating a training perceptual output for each sensor output, wherein each training sensor output is associated with a perceptual ground truth, and wherein the perceptual slice includes an online error estimator that provides an online perceptual error estimate for each training perceptual output; The PSPM is trained using the training perception output and the online perception error estimation of the training perception output, wherein the trained PSPM provides a probability perception uncertainty distribution in the form of p(e, E |t), where p(e, E |t) represents the probability of calculating a specific perception output e and a specific online perception error estimation E for the perception slice given the perception ground truth t.
18. A computer system comprising a Perceptual Statistical Performance Model (PSPM) for modeling perceptual slices of a runtime stack of a robotic system and configured to: Receive a calculated perceptual ground truth t; Determine a probabilistic perceptual uncertainty distribution of the form p(e, E | t) from the set of learned parameters from the perceived ground truth t, where p(e, E |t) represents the probability of computing a particular perceptual output e and a particular online perceptual error estimate for a perceptual slice given the perceptual ground truth t; and the probability perceptual uncertainty distribution is defined over a range of possible perceptual outputs and online perceptual error estimates, and the parameters are learned from a set of actual perceptual outputs generated using the perceptual slice to be modeled. E 19. A computer program product for programming one or more computers to implement the computer system of any one of claims 1-15, or to implement the computer-implemented method of claim 16 or 17, or to implement the computer system of claim 18.
Citation Information
Patent Citations
Autonomous vehicle manoeuvres
GB201816852D0