Performance Testing of Robot System
Through the perceived statistical performance model (PSPM), the problem that simulated reality perception errors cannot be effectively simulated by simulated scenarios in autonomous driving vehicle tests is solved, which improves test efficiency and accuracy and reduces costs.
Patent Information
- Application Number
- CN202080059350.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-06
- Filing Date
- 2020-08-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-08-21
AI Technical Summary
The existing autonomous vehicle testing methods cannot effectively simulate realistic perception output in simulation scenarios, resulting in inefficient safety testing and high-fidelity simulation cost, making it difficult to reach the same safety level as human drivers within a limited mileage.
Perceived statistical performance model (PSPM) is used to model the probability uncertainty distribution of perceived output, simulate reality perception errors, and combine the time-dependent model to generate reality perception outputs for use by the test system.
It improves the efficiency and accuracy of autonomous driving vehicle testing, reduces the testing cost, and can effectively evaluate the safety performance of the vehicle in a simulated environment, approaching the real-world test effect.
Smart Images

Figure CN114270368B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to performance testing of autonomous vehicles and other robotic systems. Performance testing is crucial for ensuring that such systems can operate at a guaranteed level of safety. Background Art
[0002] It is estimated that in order for an autonomous vehicle (AV) to reach a safety level comparable to that of a human driver, at most 1 error must occur in every 10^7 autonomous driving decisions across the entire Operational Design Domain (ODD) of the AV.
[0003] Given the complexity of the AV and the ODD itself, this poses a huge challenge. A self-driving system is an extremely complex assembly composed of interdependent and interacting software and hardware components (each component is vulnerable to limitations or errors). Several components use neural networks for object detection, type classification, action prediction, and other critical tasks. The system needs to operate safely within the ODD. In this context, the ODD describes all possible driving scenarios that an AV may encounter, and thus, it has infinite possibilities in itself, with variables including road topologies, users, appearances, lighting, weather, behaviors, seasons, speeds, randomness, and intentional behaviors.
[0004] The industry standard method for safety testing is based on actual driving test mileage. An autonomous vehicle fleet is driven by test drivers, and when test driver intervention is required, the decision is characterized as unsafe. Once an instance of test driver intervention occurs in a specific real-world driving scenario, the circumstances of that driving scenario can be explored to isolate the factors that led to the unsafe behavior of the AV and take appropriate mitigation actions. Summary of the Invention
[0005] Simulation has been used for safety testing, but it is only useful if the simulation scenario is realistic enough (if an AV planner makes an unsafe decision in a completely unrealistic simulation scenario, then it is far less useful in the context of safety testing than an instance of unsafe behavior in a realistic scenario).
[0006] One approach is to run simulations based on real-world scenarios that require testing driver intervention. Sensor outputs from the AV are collected, and the sensor outputs can be used to reconstruct driving scenarios in the simulator that require testing driver intervention. The variables of the scenarios can be "fuzzified" at the planning level so as to test variations of real-world scenarios that are still realistic. In this way, more information about the causes of unsafe behavior can be obtained, analyzed, and used to improve the prediction and planning models. However, significant problems arise because as the number of errors per decision decreases, the number of test miles needed to find a sufficient number of instances of unsafe behavior increases. A typical AV planner may make approximately 1 decision every two seconds on average. At an average speed of 20 miles per hour, this amounts to approximately 90 decisions per mile driven. This in turn implies that there is less than one error per 10^5 miles driven in order to match the human safety level. Robust safety testing requires multiple tests to fully test the AV across its ODD. This situation will deteriorate further with the development of the perception stack because with each change in the perception stack, more test miles are required. For these reasons, this approach is simply not feasible when testing at a safety level close to that of humans.
[0007] There are other problems with existing simulation methods.
[0008] One approach is simulation at the planning level, but this cannot fully account for the effects of perception errors. Many factors can affect perception errors, such as weather, lighting, distance to another vehicle or the speed of another vehicle, occlusion, etc.
[0009] An alternative would be a full "fidelity" simulation, where the entire hardware and software stack of the simulated AV is replicated. However, this is a huge challenge in itself. The AV perception pipeline typically consists of multiple perception components that work together to interpret the sensor outputs of the AV.
[0010] One problem is that some perception components (such as Convolutional Neural Networks (CNNs)) are particularly sensitive to the quality of the simulated data. Although high-quality simulated image data can be generated, CNNs in perception are extremely sensitive even to minor deviations from real data. Therefore, these would require extremely high-quality simulated image data that covers all possible conditions that the AV might encounter in the real world (e.g., different combinations of simulated weather conditions, lighting conditions, etc.) - otherwise their behavior in simulated scenarios will not fully reflect their behavior in the real world.
[0011] The second problem is that certain types of sensor data are particularly difficult to model (simulate). As a result, even perception systems that are not particularly sensitive to the quality of the input data will give poor results. For example, radar belongs to the category of sensor data that is extremely difficult to simulate. This is because the physical characteristics of radar are inherently difficult to model.
[0012] The third overarching problem is the issue of computational efficiency. Based on current hardware constraints, it is estimated that it may be possible to achieve photorealistic simulation in real time as much as possible (even if other problems can be overcome).
[0013] The present disclosure provides a distinct approach for simulation-based safety testing using a model herein referred to as a "Perceptual Statistical Performance Model" (PSPM). The core problem addressed in the present disclosure is to simulate realistic perception outputs - that is, perception outputs with realistic errors - in a manner that is not only more robust than photorealistic simulation but also significantly more efficient.
[0014] The PSPM models perceptual errors in terms of a probabilistic uncertainty distribution based on a robust statistical analysis of the actual perceptual outputs computed by one or more perceptual components being modeled. The unique aspect of the PSPM is that, given a perceptual ground truth (i.e., the "perfect" perceptual output that would be computed by a perfect but unrealistic perceptual component), the PSPM provides a probabilistic uncertainty distribution that represents the realistic perceptual component that can be provided by the perceptual component it is modeling. For example, given a ground truth 3D bounding box, the PSPM modeling the PSPM of a simulated 3D bounding box detector will provide an uncertainty distribution representing the realistic 3D object detection output. Even when the perception system is deterministic, it can be effectively modeled as stochastic to account for the epistemic uncertainty of the many hidden variables it depends on in practice.
[0015] Of course, the perceptual ground truth will not be available at runtime in the real world AV (which is why complex perceptual components are needed to reliably interpret imperfect sensor outputs). However, the perceptual ground truth can be directly derived from the simulated scenarios running in the simulator. For example, in the case of a 3D simulation of a driving scenario with an ego vehicle (the simulated AV being tested) given the presence of external actors, the ground truth 3D bounding box can be directly computed from the simulated scenario of the external actors based on the size and pose (position and orientation) of the external actors relative to the ego vehicle. Then, the PSPM can be used to derive realistic 3D bounding object detection outputs from these ground truths, and the realistic 3D bounding object detection outputs can in turn be processed by the remaining AV stack as if they were at runtime.
[0016] A first aspect of the present disclosure provides a computer system for testing and / or training a runtime stack of a robotic system, the computer system comprising:
[0017] A simulator configured to run a simulation scenario in which a simulation subject interacts with one or more external objects;
[0018] A planner for a runtime stack, the planner for the runtime stack being configured to make autonomous decisions for each simulation scenario based on a time series of sensed outputs computed for the simulation scenario; and
[0019] A controller for the runtime stack, the controller for the runtime stack being configured to generate a series of control signals to cause the simulation subject to execute the autonomous decisions as the simulation scenario progresses;
[0020] wherein the computer system is configured to compute each sensed output by:
[0021] Computing a sensed ground truth at the current moment based on the current state of the simulation scenario;
[0022] Applying a Perceptual Statistical Performance Model (PSPM) to the sensed ground truth to determine a probabilistic sensed uncertainty distribution at the current moment; and
[0023] Sampling a sensed output from the probabilistic sensed uncertainty distribution;
[0024] wherein the PSPM is used to model a sensing slice of the runtime stack and is configured to determine the probabilistic sensed uncertainty distribution based on a set of parameters learned from a set of actual sensed outputs generated using the sensing slice to be modeled;
[0025] wherein the PSPM includes a time-dependent model such that the sensed output sampled at the current moment depends on at least one of: an earlier one of the sensed outputs sampled at the previous moment, and an earlier one of the sensed ground truths computed for the previous moment.
[0026] Incorporating an explicit time dependency between sensed outputs at different times can provide a more statistically representative model of the sensing slice. Generally, this allows for the explicit modeling of temporal correlations within the PSPM framework. This has many applications, one example being when the sensing slice being modeled includes components (such as filters) with explicit time dependencies.
[0027] The time-dependent model can be a hidden Markov model.
[0028] The modeled sensing slice can include at least one filtering component, and the time-dependent model is used to model the time dependency of the filtering component.
[0029] The PSPM can have a time-dependent neural network architecture.
[0030] For example, the PSPM can be configured to receive earlier perception outputs and / or earlier perception ground truth as inputs.
[0031] As another example, the PSPM can have a Recurrent Neural Network architecture.
[0032] The PSPM can be applied to perception ground truth and one or more confounding factors related to a simulation scenario. Each confounding factor is a variable of the PSPM, the value of the variable characterizes the physical conditions applicable to the simulation scenario, and the probability perception uncertainty distribution depends on the variable.
[0033] One or more confounding factors can include one or more of the following confounding factors, and the confounding factors at least partially determine the probability uncertainty distribution from which the perception output is sampled:
[0034] The occlusion level of at least one of the external objects;
[0035] One or more lighting conditions;
[0036] An indication of the time of day;
[0037] One or more weather conditions;
[0038] An indication of the season;
[0039] The physical characteristics of at least one of the external objects;
[0040] Sensor conditions, e.g., the position of at least one of the external objects in the sensor field of view of the subject;
[0041] The number or density of external objects;
[0042] The distance between two external objects;
[0043] The truncation level of at least one of the external objects;
[0044] The type of at least one of the objects, and
[0045] An indication of whether at least one of the external objects corresponds to any external object from an earlier moment of the simulation scenario.
[0046] The computer system can include a scene evaluation component configured to evaluate the behavior of an external subject in each simulation scenario by applying a predetermined set of rules.
[0047] At least some of the predefined rules may be safety related, and the scenario evaluation component may be configured to evaluate the safety of the subject's behavior in each simulated scenario.
[0048] The scenario evaluation component may be configured to automatically flag instances of unsafe behavior by the subject for further analysis and testing.
[0049] The computer system may be configured to re-run the simulated scenario in which the subject initially exhibited unsafe behavior based on a time series of the perceived ground truth determined for re-running the scenario, without applying the PSPM to those perceived ground truths and thus without perception errors; and to evaluate whether the subject still exhibits unsafe behavior in the re-run scenario.
[0050] Sampling from the probabilistic perception uncertainty distribution may be non-uniform and biased towards lower probability perception outputs.
[0051] The planner may be configured to make the autonomous decision based on a second time series of perception outputs, where the PSPM may be a first PSPM, and the computer system may be configured to compute the second time series of perception outputs using a second PSPM that models a second perception slice of the runtime stack, the first PSPM learning from a first sensor modality of the perception slice and data of the time series, and the second PSPM independently learning from a second sensor modality of the second perception slice and data of the second time series.
[0052] The planner may be configured to make the autonomous decision based on a second time series of perception outputs that correspond to a first sensor modality and a second sensor modality respectively, where the modeled perception slice may be configured to process sensor inputs of both sensor modalities and to compute the two time series by applying a PSPM that learns from data of both sensor modalities for modeling any correlations between within the perception slice.
[0053] The computer system may be configured to apply at least one unmodeled perception component of the runtime stack to the perception output to compute a processed perception output, and the planner is configured to make the autonomous decision based on the processed perception output.
[0054] The unmodeled perception component may be a filtering component applied to the time series of the perception output, and the processed perception output is the filtered perception output.
[0055] The computer system may include a scenario fuzzification component configured to generate at least one fuzzified scenario for running in the simulator by fuzzifying at least one existing scenario.
[0056] To model false negative detections, a probability-aware uncertainty distribution can provide the probability of visible objects in a successfully detected object, which is used to determine whether to provide an object detection output for the object. An object is visible when it is within the sensor's field of view of the subject in a simulated scenario, and thus the detection of such a visible object is not guaranteed.
[0057] To model false positive detections, a probability uncertainty distribution can provide the probability of false object detections, which is used to determine whether to provide a perception output for a non-existent object.
[0058] Ray tracing can be used to compute the perception ground truth for one or more external objects.
[0059] At least one of the external objects can be a moving actor, and the computer system can include a prediction stack of a runtime stack, the prediction stack of the runtime stack being configured to predict the behavior of the external actor based on the perception output, and a planner being configured to make autonomous decisions according to the predicted behavior.
[0060] The computer system can be configured to record details of each simulated scenario in a test database, where the details include the decisions made by the planner, the perception outputs on which those decisions are based, and the behavior of the simulated subject when executing those decisions.
[0061] A second aspect of the present disclosure provides a computer-implemented method for performance testing a runtime stack of a robotic system, the method comprising:
[0062] Running a simulated scenario in a simulator, where a simulated subject interacts with one or more external objects, where a planner of the runtime stack makes autonomous decisions for the simulated scenario based on a time series of perception outputs computed for the simulated scenario, and a controller of the runtime stack generates a series of control signals to cause the simulated subject to execute the autonomous decisions as the simulated scenario progresses;
[0063] where each perception output is computed by:
[0064] Computing the perception ground truth at the current moment based on the current state of the simulated scenario;
[0065] Applying a Perception Statistical Performance Model (PSPM) to the perception ground truth to determine the probability-aware uncertainty distribution at the current moment; and
[0066] Sampling a perception output from the probability-aware uncertainty distribution;
[0067] where the PSPM is used to model the perception slice of the runtime stack and determines the probability-aware uncertainty distribution based on a set of parameters learned from a set of actual perception outputs generated using the perception slice to be modeled.
[0068] Among them, the PSPM includes a time-dependent model such that the perceptual output sampled at the current moment depends on at least one of the following: an earlier one of the perceptual outputs sampled at the previous moment, and an earlier one of the perceptual ground truths calculated for the previous moment.
[0069] A third aspect of the present invention provides a Perception Statistical Performance Model (PSPM) implemented in a computer system, the PSPM being used to model a perception slice of a runtime stack of a robotic system and being configured to:
[0070] Receive the computed perceptual ground truth t at the current time t t ;
[0071] According to the set of learned parameters θ, the ground truth t t Determine the form p(e t |t t , t t-1 ) or p(e t |t t , e t-1 ) is a probability-aware uncertainty distribution, where p(e t |t t , t t-1 ) represents the calculated perceptual ground truth t given the current time t and the previous time t-1 t , t t-1 In the case of perceptual slices, a specific perceptual output e is calculated. t The probability of t |t t , e t-1 ) represents the calculated perceptual ground truth at a given current time t t and the sampled perceptual output e at the previous time t-1 t-1 In the case of perceptual slices, a specific perceptual output e is calculated. t where the probabilistic perceptual uncertainty distribution is defined over the range of possible perceptual outputs and the parameters θ are learned from the set of actual perceptual outputs generated using the perceptual slice to be modeled.
[0072] Another aspect of the present invention provides a computer program for programming one or more computers to implement any of the methods or functions described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] For a better understanding of the present disclosure, and to illustrate how embodiments of the present disclosure may be implemented, reference is made to the accompanying drawings, in which:
[0074] Figure 1Shows a schematic block diagram of the stack during the operation of an autonomous vehicle;
[0075] Figure 2 Shows an example of a real - world driving scenario;
[0076] Figure 3 Shows a test pipeline using photorealistic simulation;
[0077] Figure 4 Shows an alternative PSPM - based test pipeline according to the present disclosure;
[0078] Figure 5 Shows how perception performance is affected by confounding factors;
[0079] Figure 6 Provides a high - level overview of certain principles of PSPM - based safety testing;
[0080] Figure 7 Shows a perception error dataset that can be used to train PSPM;
[0081] Figure 7A Shows applied to Figure 7 the results of a trained PSPM on a perception error dataset;
[0082] Figure 8 Shows an engineering pipeline incorporating PSPM;
[0083] Figure 9 Shows an example of a perception stack;
[0084] Figures 9A - 9C Shows different ways in which a perception stack can be modeled using one or more PSPMs Figure 9 ;
[0085] Figure 10 Provides a schematic overview of factors that can lead to perception uncertainty;
[0086] Figure 11 Shows an example of simulated image data to which certain forms of perception components are highly sensitive;
[0087] Figure 12 And Figure 13 Shows an aerial view and a driver's view of a roundabout scenario;
[0088] Figure 14 Schematically depicts stereo imaging geometry;
[0089] Figure 15 Shows an example time series of the additive error of the position component;
[0090] Figure 16 Shows a hysteresis plot of the position error;
[0091] Figure 17 Shows a graphical representation of the time-dependent position error model;
[0092] Figure 18 Shows an example binning scheme for the confounding factor azimuth and distance;
[0093] Figure 19 Shows a hysteresis plot of the position error increment;
[0094] Figure 20 Shows a histogram of the position error increment for the X, Y, and Z components;
[0095] Figure 21 Shows the PDF-fitted position error increment for the X, Y, and Z components;
[0096] Figure 22 Shows an example mean of the error increment distribution in the training data (based on a single object being tracked);
[0097] Figure 23 Shows a time series plot of the true perception error and the simulated error;
[0098] Figure 24 Shows a hysteresis plot for the true perception error and the simulated error;
[0099] Figure 25 Graphically depicts the relative importance of certain confounding factors in a particular (left to right) for the target association state as determined by MultiSURF Relief analysis;
[0100] Figure 26 Graphically depicts the relative importance of the confounding factors (left to right) for the target transition as determined by MultiSURF Relief analysis;
[0101] Figure 27 Shows an example node in a neural network;
[0102] Figure 28 Shows a high-level overview of a convolutional neural network architecture;
[0103] Figure 29 Shows the PSPM implemented as a neural network during training and inference;
[0104] Figure 30 Shows the neural network PSPM with one or more confounding factor inputs at the input layer;
[0105] Figure 31Shows an example of a time-dependent neural network architecture;
[0106] Figure 32 Shows a "set-to-set" PSPM implemented as a neural network;
[0107] Figure 33A Schematically depicts the spatial encoding of perceptual outputs that facilitate processing in a convolutional neural network (CNN);
[0108] Figure 33B Schematically depicts the training phase of a CNN PSPM;
[0109] Figure 33C Schematically depicts a trained CNN PSPM at inference time;
[0110] Figure 33D Shows how a CNN PSPM can be architected to encode a perceptual output distribution in an output tensor from which a realistic perceptual output can be sampled; and
[0111] Figure 34 Shows how a PSPM can be configured to model a perceptual slice that includes an online error estimation component. Detailed Description
[0112] 1. Overview
[0113] The terms "PSPM" and "PRISM" are used interchangeably in the following description.
[0114] When making a safety case for an autonomous vehicle, it is impractical to perform all the required tests in the real world. However, building simulations in such high fidelity to make the vehicle's perception system perform equivalently on real and simulated data is an unsolved problem. The method referred to herein as "PRISM" solves this problem by building a surrogate model of the perception system, where the perception system includes both sensors and one or more perception components that interpret the sensor data captured by the sensors. PRISM is the distribution of credible perceptual outputs given some low-fidelity scene representation (perceptual ground truth).
[0115] Building on the above, to ensure that autonomous driving technology can be proven safe, it is necessary to test autonomous driving technology in a large number of scenarios. Conducting such tests with real vehicles is both expensive and time-consuming. In natural scenarios, most of the miles driven will be uneventful - in the UK in 2016, 136,621 people were injured and 1,792 people died in road accidents, and all motor vehicles traveled 323.7 billion miles, with only one accident per 2.4 million miles driven. Simulation must become part of the autonomous driving technology test strategy. Simulation miles are much cheaper than real miles, and it is easier and safer to increase the number of hazards per mile in simulation than in the real world.
[0116] One way to generate realistic perception outputs is via a high-fidelity simulation of the world, including sensor measurements. In this method, realistic sensor readings are generated, and the sensor readings are fed into the car's software in place of real sensor readings. For example, a realistic twin of the real world is rendered as an image for perception input. This rendering is as Figure 11 shown. The car's software outputs control signals for the car's actuators, and the control signals are fed into a physics simulation. New sensor readings are generated based on the output of the physics simulation, thus closing the loop. This method requires accurate models to be generated for tasks ranging from challenging to unresolved:
[0117] · It is possible to simulate road surfaces, vehicle dynamics, and other physical characteristics with current technology, but they are not well understood.
[0118] · It is possible to simulate GPS, IMU, and wheel-encoding, but it is important to obtain their error statistics correctly.
[0119] · Visual appearance, camera lenses, and image sensor modeling are reasonably well understood, but high-fidelity rendering is slow.
[0120] · Lidar modelling is similar to camera modelling, but has different material reflection characteristics. The scanning characteristics of lidar are an additional challenge.
[0121] · It is difficult to accurately model radar echoes with current technology because it is difficult to model the relevant material characteristics, the detailed dependence on shape and multiple reflections.
[0122] ·Worst of all, state-of-the-art neural networks for visual object detection are extremely sensitive to detailed image statistics, and constructing synthetic images that elicit the same network behavior as equivalent real images is an unsolved problem.
[0123] Inaccurate models of the above sensors will affect the output of the perception module in the simulation, resulting in potentially different ego behavior. Such differences in behavior limit the usefulness of these simulations in evaluating real-world performance. Additionally, running the many miles of photorealistic simulation required to verify the safety behavior of an autonomous vehicle is expensive. This is because rendering photorealistic scenes is a slow, computationally intensive task that requires a GPU. High-fidelity simulation is both difficult and expensive, and conclusions from testing using high-fidelity simulation are unlikely to generalize to the real world.
[0124] Figure 1 A data flow diagram of an autonomous vehicle stack 100 through decomposition is shown. A perception system 102 receives sensor readings from the world and outputs a scene representation. A planning and prediction system (represented by reference numerals 104 and 106 respectively) employs the scene representation and plans a trajectory through the scene. A control system 108 outputs control signals to the world, and the control signals will cause the vehicle to follow the trajectory.
[0125] The perception system 102, the planning and prediction systems 104, 106, and the control system 108 communicate with each other using well-defined interfaces. The perception system 102 uses raw sensor data and processes the raw sensor data into a more abstract scene representation. This representation includes dynamic object pose, extent, motion, and detection confidence. The planning and prediction system predicts the possible trajectories of other agents in the scene and plans a safe, legal, and comfortable path through the scene. The control system uses the desired trajectory from the planning and prediction system and outputs control signals for the actuators.
[0126] In many cases, especially at the interface between perception and planning, these internal interfaces are easier to simulate than sensor readings. These interfaces can be used for a second type of simulation called low-fidelity simulation. One can simulate only those aspects of the world necessary to reconstruct the abstract scenario representation used by the planner and provide that abstract scenario representation directly to the planner, thus taking the perception system out of the loop. While this avoids some of the burden of high-fidelity simulation, it poses new challenges: replicating the behavior of the perception system. It is well known that perception systems are not perfect, and errors in the perception system can affect prediction, planning, and control systems in meaningful ways. Since the results tested in simulation should generalize to the real world, it must be possible to simulate realistic perception outputs.
[0127] A method is proposed for simulating realistic perception outputs using a model called PRISM. PRISM is a distribution of plausible perception outputs given some low-fidelity scenario representations. The mathematical framework guiding the creation of PRISM is outlined, a prototype is created, and the modelling choices are documented. Doing so demonstrates that the modelling approach is sound.
[0128] Broadly speaking, in high-fidelity simulation, the simulator replaces the world and the entire vehicle stack is treated as a black box. In low-fidelity simulation, the world and the perception system 102 are replaced (see Figure 4 and the description below).
[0129] Figure 1 A highly schematic block diagram of a runtime stack 100 for an autonomous vehicle (AV) is shown. The runtime stack 100 is shown as including a perception stack 102, a prediction stack 104, a planner 106, and a controller 108.
[0130] The perception stack 102 receives sensor outputs from the on-board sensor system 110 of the AV.
[0131] The vehicle-mounted sensor system 110 can take different forms, but generally includes various sensors, such as, an image capture device (camera / optical sensor), a Light Detection and Ranging (LiDAR) and / or a Radio Detection and Ranging (RADAR) unit, a satellite positioning sensor (GPS, etc.), a motion sensor (accelerometer, gyroscope, etc.), etc. The various sensors jointly provide rich sensor data, from which detailed information about the surrounding environment and the state of the AV and any external actors (vehicles, pedestrians, cyclists, etc.) within that environment can be extracted.
[0132] Thus, the sensor output typically includes sensor data of multiple sensor modalities, such as, stereo images from one or more stereo optical sensors, LiDAR, RADAR, etc.
[0133] The perception stack 102 includes multiple perception components that cooperate to interpret the sensor output and thereby provide a perception output to the prediction stack 104.
[0134] The perception output from the perception stack 102 is used by the prediction stack 104 to predict the future behavior of external actors.
[0135] The prediction calculated by the prediction stack 104 is provided to the planner 106, which uses the prediction to make autonomous driving decisions to be executed by the AV in a way that takes into account the predicted behavior of external actors.
[0136] The controller 108 executes the decisions made by the planner 106 by providing appropriate control signals to the vehicle-mounted motor 112 of the AV. In particular, the planner 106 plans the manoeuvres to be taken by the AV, and the controller 108 generates control signals to execute these manoeuvres.
[0137] Figure 2 Examples of certain perception components that can form part of the perception stack 102 are shown, namely, a 3D object detector 204 and a Kalman filter 206.
[0138] The depth estimator 202 captures stereo image pairs and applies stereo imaging, such as Semi-Global Matching, to extract depth estimates from the stereo image pairs. Each depth estimate is in the form of a depth map that assigns depth values to the pixels of one image of the stereo image pair derived from the depth map (the other image being used as a reference). The depth estimator 202 includes a stereo pair of optical sensors and a stereo processing component (hardware and / or software), which are not shown separately in the figure. In accordance with the terminology used herein, both the optical sensors and the stereo processing component of the depth estimator 202 are considered to be part of the vehicle sensor system 110 (not the perception stack 102). The depth map is a form of sensor output provided to the perception stack 102.
[0139] The 3D object detector 204 receives the depth estimates and uses them to estimate the poses of external actors near the AV (ego vehicle). Two such external actors are shown in the form of two other vehicles. The pose representation in this context is a 6D pose, i.e., (x, y, z, pitch, roll, yaw), representing the position and orientation of each external actor in 3D space.
[0140] For illustrative purposes, Figure 2 is highly simplified. For example, the 3D object detector can be formed by multiple collaborative perception components that jointly operate on the sensor outputs of multiple sensor modalities. The application of PSPM in a more complex stack will be described later. For now, to illustrate some of the core principles of PSPM, consider a simplified example in which it is assumed that the 3D object detector operates on the sensor output of a single modality (stereo depth).
[0141] In real-world scenarios, multiple physical conditions can affect the performance of the perception stack 102. As noted, the physical conditions that are considered variables with respect to a particular PSPM are referred to as "confounders". This allows for the consideration of variable physical conditions that are statistically relevant to a particular perception slice.
[0142] As described above, one simulation method would be to not only attempt to Figure 1perform a realistic simulation of the entire runtime stack 100, and also attempt to realistically simulate the vehicle sensor system 110 and the motor 112. This is described in Figure 3 In this scenario, the challenge is the simulation of sensor data: certain types of sensor data (e.g., radar (RADAR)) are inherently difficult to simulate well, while other types of sensor data (image data) are relatively easier to simulate, and certain perception components (such as, CNN) are highly sensitive to even minor deviations from actual sensor data. Another challenge is the simulation of sensors and the large amount of computing resources required to run complex perception components (such as, CNN).
[0143] For example, for Figure 3 the arrangement, it would be necessary to simulate extremely high-quality depth maps and run the 3D object detector 204 on those simulated depth maps. During the simulation, even a very small deviation in the simulated depth map (compared to the ground truth depth map provided by the stereo depth estimator 202) would significantly affect the performance of the 3D object detector 204.
[0144] Figure 4 An advanced schematic overview of the PSPM-based simulation is provided. In this case, a "headless" simulator setup is used, where there is no need to create simulated sensor data (e.g., simulated images, depth maps, lidar, and / or radar measurements, etc.), and there is no need to apply the perception stack 102 (or at least not fully apply the perception stack 102 - see below). Instead, one or more PSPMs are used to efficiently compute the realistic perception output, which is then fed into the high-level components of the runtime stack 100 and processed as they would be at runtime.
[0145] It is said that the PSPM models a "perception slice", which can be all or part of the perception stack 102. The perception slice can be a single perception component or multiple collaborating perception components.
[0146] Mathematically, the perception slice can be represented as a function F, where:
[0147] e = F(x),
[0148] e is the perception output of the perception slice, and x is the set of sensor outputs that are run on the perception component.
[0149] On the AV at runtime, e is determined by applying F to x, and x is given by the sensors.
[0150] The PSPM mapped to the confounder space C can be represented as a function p, where:
[0151] p(e|t, c) represents a probabilistic uncertainty distribution that provides the probability of F computing a sensed output e given a sensed ground truth t and a set of one or more confounding factors c (i.e., given a particular set of possible real-world conditions represented by a point c in the confounding space C).
[0152] For example, for 2D bounding box detection:
[0153] · F can be a CNN
[0154] · x can be an RGB image
[0155] · t can be a ground truth bounding box that can be directly computed from the simulation using ray tracing (without simulating x and without applying F), or a set of multiple such boundaries bounding multiple ground truth objects (set-to-set approach)
[0156] · c can be distance and / or weather, etc.
[0157] In the Figure 3 example, e represents one or more 6D pose vectors computed by a 3D object detector 204, and x represents a depth map from which e is derived provided by a stereo depth estimator 202.
[0158] Figure 5 Shows how to use PSPM to simulate Figure 3 the realistic sensed output of a scene. In this case, the sensing slice is the 3D object detector 204.
[0159] "Realistic" in the current context means a sensed output that is more realistic than the sensed ground truth.
[0160] Provides a PSPM 500 that essentially models the sensing slice 204 as a noisy "channel" affected by both the characteristics of the stereo depth estimator 202 and the physical environment. The physical environment is characterized by a set of confounding factors c, which in this example are: lighting, weather, occlusion, and distance to each external actor.
[0161] To apply the PSPM 500, the sensed ground truth can be directly computed from the simulated scene under consideration. For example, in a simulated scene of an AV (ego vehicle) with many external actors in its vicinity, the 6D pose ground truth can be determined by directly computing the 6D pose of each external actor in the ego vehicle's reference frame.
[0162] Then, the PSPM 500 uses the computed ground truth t to compute the distribution p(e|t, c). Continuing with the above example, this will provide, for each simulated external actor, the probability that the actual 3D object detector 204 computes the perceptual output e [estimated 3D pose of the external actor] given the perceptual ground truth t ["actual" 6D pose] in a real-world scenario characterized by the same confounding factors c.
[0163] After computing p(e|t, c), it can be used to run multiple simulations on a series of realistic perceptual outputs (PSPM samples) obtained by sampling p(e|t, c). Depending on the realistic means of p(e|t, c) with a high enough probability—note that if relatively low-probability perceptual outputs (outliers) are still realistic, it may be highly desirable to test them. The degree of testing outliers will depend on the safety level that the AV needs to meet.
[0164] In Figure 5 , three realistic perceptual outputs e1, e2, and e3 are shown by way of example. These are sampled from p(e|t, c).
[0165] One way is to sample the perceptual output from p(e|t, c) in a way that favors the most likely perceptual outputs, for example, using Monte Carlo sampling. Broadly speaking, this will test a greater number of the most likely perceptual outputs and fewer of the less likely ones.
[0166] However, while this may be useful in some contexts, in other contexts, it may be more useful to deliberately test a greater number of "outliers" (i.e., less likely but still realistic perceptual outputs) because it is possible that outliers are more likely to cause or contribute to unsafe behavior. That is, p(e|t, c) can be sampled in a way that deliberately biases towards outliers to deliberately make a particular scenario more "challenging" or "interesting" as it progresses. This can be achieved by transforming the distribution of the PSPM and sampling from the transformed distribution.
[0167] Figure 6 An overview of the process for constructing the PSPM is provided. A large number of real sensor outputs x are collected and annotated with the perceptual ground truth t. This is exactly the same process as generating the training data for the perceptual components of the perceptual stack 102 (represented by box 602)—and a first subset of the annotated sensor outputs is used for this proposal. The trained perceptual slice 204 is shown, which (in the real world) is run and used to construct the PSPM 500, which will model this perceptual slice 204 during simulation.
[0168] Continue Figure 3 As an example, the real sensor output would be a depth map, and the ground truth would be the ground truth 6D pose of any object captured in the depth map. This annotated data is used not only to train the 3D object detector 204 (the perception slice in this example), but also to build the PSPM 500, which models the 3D object detector 204 during simulation.
[0169] Box 604 represents PSPM construction (training), and a second subset of the annotated sensor output is used for this purpose. Each sensor output is additionally annotated with a set of confounding factors c, which characterize the physical conditions under which it was captured in the real world. Each set of confounding factors c that the PSPM needs to be able to adapt to requires a large amount of sensor output. For full "Level 4" autonomous driving, this means capturing and annotating sensor output across the entire ODD.
[0170] The PSPM 500 can take the form of a parametric distribution:
[0171] Dist(t, c; θ)
[0172] where t and c are the variables on which the distribution depends, and θ is the set of learning parameters.
[0173] The parameters θ are learned as follows:
[0174] 1) Apply the trained perception slice 204 to each sensor output x to compute the corresponding perception output e;
[0175] 2) For each perception output e, determine the deviation (error) Δ between e and the corresponding ground truth t;
[0176] 3) Each error Δ is associated with the ground truth t and the set of confounding factors c associated with the corresponding sensor output x;
[0177] 4) Considering the associated ground truth and the variable confounding factors c, adjust the parameters θ to fit the distribution to the error Δ.
[0178] As will be apparent, various known forms of parametric distributions / models can be applied in this context. Therefore, they will not be elaborated further.
[0179] More generally, the training set for PSPM training consists of the perceived ground truth (from manual, automatic, or semi-automatic annotation) and the corresponding actual perceived output generated by the perceived slice 204 to be modeled. The purpose of training is to learn the mapping between the perceived ground truth and the perceived output distribution that captures the statistics of the actual perceived output. Thus, for a given ground truth t, the perceived output sampled from the distribution p(e|t) will be statistically similar to the actual perceived output used for training.
[0180] As an example, the perceived slice 204 can be modeled as having zero-mean Gaussian noise. However, it should be emphasized that the present disclosure is not limited to this aspect. The PSPM can well take the form of more complex non-Gaussian models. As an example, the PSPM can take the form of a hidden Markov model, which will allow explicitly modeling the temporal dependence between the perceived outputs at different times.
[0181] In the Gaussian case, for example, the PSPM 500 can be characterized as:
[0182] e = t + ε
[0183] ε ∼ N(0, Σ(c)),
[0184] where N(0, Σ(c)) represents a Gaussian distribution with zero mean and covariance Σ(c), and the covariance Σ(c) varies according to the confounding factor c. During simulation, then Gaussian sampling noise will be drawn and added to the perceived ground truth. This will depend on the variance of the Gaussian and thus on the confounding factor applicable to the simulation scenario.
[0185] Exemplary PSPM Error Dataset
[0186] Figure 7 An example of a raw error plot of a two-dimensional prediction space is shown - for example, each point can correspond to (x, y) coordinates, which can be estimated by a 2D object detector. Each prediction e is represented by a circle, and each ground truth t is represented by a star. Each error Δ is represented by a line segment between the corresponding prediction e and the corresponding ground truth t (the longer the line segment, the larger the error).
[0187] To construct the PSPM, the aim is to take into account the variable confounding factor c to accurately capture Figure 7 the error relationship between the data points (in this context, the data points are the errors Δ) in a way that adjusts the parameter distribution.
[0188] Figure 7A Shown is applied to Figure 7Results of the trained PSPM for the error dataset.
[0189] Selected Confounding Factors
[0190] The decision on which confounding factors to incorporate is observation-driven: when it can be seen that a particular physical property / condition has a material impact on the perceived uncertainty, this may trigger its introduction as a confounding variable into the applicable PSPM. Only statistically relevant confounding factors should be introduced.
[0191] One way to handle confounding factors is to partition the error dataset according to the confounding factors and train separate models for each partition of the dataset. To give a very simple example, two confounding factors could be "lighting" and "weather", each of which could take a binary "good / poor" value. In this case, the dataset could be divided into four subsets with (lighting, weather) = (good, good), (good, bad), (bad, good), and (bad, bad) respectively, and four separate models could be trained for each subset. In this case, the PSPM consists of four models, where the confounding factor variable c = (lighting, weather) is used as an index to determine the selection of the model.
[0192] Engineering Pipeline Architecture
[0193] Figure 8 A highly schematic overview of the engineering pipeline incorporating the PSPM is shown. The entire pipeline encompasses everything from data collection, annotation, and extraction; training of the perception components; PSPM characterisation and simulation-based testing.
[0194] A large number of sensor outputs (such as stereo images, depth maps, lidar measurements, and radar measurements) are collected using a fleet of vehicles, each equipped with a sensor system 110 of the type described above. These are collected in the types of environments and driving scenarios that will need to be able to be handled in AV practice, for example, in the target urban areas where AV deployment is desired. The collection vehicles themselves can be AVs or manually driven vehicles equipped with similar sensor systems.
[0195] For the purpose of annotating the captured sensor outputs with ground truth, a ground truth pipeline 802 is provided. This includes annotating the sensor outputs with the type of perception ground truth described above. The sensor outputs annotated with perception ground truth are stored in an annotated ground truth database 804. Further details are described below.
[0196] In addition, the sensor outputs captured by the vehicle fleet are also used to extract driving scenarios, which can then be recreated in the simulator. A high-level structured scenario description language is used to capture the driving scenarios and store them in the scenario database 806.
[0197] The sensor outputs captured from the vehicle fleet are not the only source of information from which driving scenarios can be extracted. Additionally, CCTV (closed circuit television) data 800 is used as the basis for scenario extraction. Typically, CCTV data captured in an urban environment, such as an urban environment that shows challenging urban driving scenarios, such as complex roundabouts, provides a rich source of challenging driving scenarios and thus an excellent basis for safety testing. A collection of backend perception components 808 is used to process the CCTV data 800 to assist in the process of extracting driving scenarios from the CCTV data 800, and the driving scenarios are also stored in the scenario database 806 in the format of a scenario description language.
[0198] Further details of the scenario description language and the process of extracting scenarios from CCTV data and other data can be found in UK Patent Application No. 1816852.6, which is incorporated herein by reference in its entirety.
[0199] The driving scenarios captured in the format of the scenario description language are high-level descriptions of driving scenarios. The driving scenarios have both a static layout (such as road layout (lanes, markings, etc.), buildings, road infrastructure, etc.) and dynamic elements. Figure 8 In the
[0200] pipeline, the static layout is captured in the scenario description as a pointer to an HD (high definition) map stored in the map database 826. The HD map itself can be derived from the annotated sensor outputs collected by the AV vehicle fleet and / or from CCTV.
[0201] Running Simulation
[0202] The test suite orchestration component 810 uses the captured driving scenarios to formulate test instance specifications 812, which in turn can be run in the 3D simulator 814 as 3D multibody simulations. The purpose of these simulations is to enable the derivation of accurate perception ground truth, and then apply PSPM. Therefore, they contain a sufficient level of 3D geometric details to be able to derive, for example, ground truth 3D bounding boxes (dimensions, 6D poses of external actors in the ego-vehicle reference frame), odometry, and ego-localization outputs, etc. However, they are not photorealistic simulations because that level of detail is not required. They also do not attempt to simulate conditions such as rain, lighting, etc., as they are modeled as confounding factors c.
[0203] To provide greater scene variation, a scene "fuzzer" 820 is provided that can fuzz the scene in the above sense. Fuzzing the scene means changing one or more variables of the scene to create a new scene that is still realistic.
[0204] Typically, this will involve fuzzing dynamic elements into a static scene, for example, changing the motion of external actors, removing or adding external actors, etc.
[0205] However, the static layout may also be fuzzed, for example, to change the curvature of the road, change the position of static objects, change road / lane markings, etc.
[0206] Figure 8 The training block 602 is shown as being able to access the annotated ground truth data database 804, which, as described above, is used to train the perception slice 204 of the runtime stack 100.
[0207] As described above and as Figure 8 shown, the perception slice 204 is not necessarily the entire perception stack 102. In this example, the perception stack 102 is "sliced" before the set of final fusion components (filters), and the set of final fusion components (filters) cooperate to fuse the perception outputs from below the perception stack 102. These form part of one or more remaining prediction slices 205, which do not use PSPM modeling but are applied to PSPM samples. The output of the final (unmodeled) prediction slice 205 is directly fed into the prediction stack 104.
[0208] The PSPM is shown as stored in the PSPM database 820.
[0209] Running Simulation
[0210] The PSPM sampling orchestration component 816 uses 3D multibody simulation in the 3D simulator 814 to derive ground truth, which in turn forms the input for one or more PSPMs for modeling the sensed slice 104PSPM and provides PSPM samples 818 for each simulation. The PSPM samples 818 are fed into the remainder of the runtime stack 100 (i.e., in this example the PSPM samples 818 are fed into the final set of filters 205) and used as the basis for planning and prediction, finally causing the controller 108 to generate control signals that are provided to the set of simulated AV motors.
[0211] The simulated motors are not shown in Figure 8 but are shown in Figure 4 and are denoted by the reference numeral 412. As Figure 4 shown, the 3D multibody simulation in the 3D simulator is driven in part by the simulated motors. These determine how the (simulated in this case) body moves within the static layout (i.e., they determine the changes in the state of the body (which can be referred to herein as the simulated robot state)). In turn, the behavior of the body can also affect the behavior of the simulated external actors in response to the movement of the AV. As the 3D simulation progresses, new sensed ground truth continues to be derived and fed into the PSPM 500 iteratively until the simulation is complete.
[0212] Each completed simulation is recorded as a set of test results stored in the test database 822.
[0213] Note that the same scenario can be run multiple times, but not necessarily produce the same results. This is due to the probabilistic nature of the PSPM: each time the scenario is run, different PSPM samples may be obtained from the PSPM. Therefore, a large amount of information is obtained by running the same simulation scenario on multiple occasions and observing, for example, the degree to which the behavior of the simulated body is different in each instance of the scenario (a large difference in body behavior indicates that the impact of sensing errors is significant) or the proportion of scenario instances in which the behavior of the body is unsafe. If the same scenario is run a large number of times and the body behaves safely and very similarly in each scenario, it indicates that the planner 106 is able to plan correctly under uncertain conditions in that scenario.
[0214] Test Oracle
[0215] Driving scenarios used as a simulation basis are typically based on real-world scenarios or blurred real-world scenarios. This ensures that real-world scenarios are being tested. However, note that these are usually driving scenarios that do not involve any actual autonomous vehicles, i.e., at least in most cases, the driving scenarios being tested are derived from real-life instances of human driving. As a result, it is not known which scenarios might lead to failure.
[0216] To this end, a scenario evaluation component 824 (referred to herein as the "test oracle") is provided, which has the function of evaluating whether the behavior of the simulated AV in the scenario is acceptable once the simulation is complete. The output of the test oracle 824 can include a simple binary (yes / no) output to flag whether the AV is behaving safely, or it can be a more complex output. For example, it can include a risk score.
[0217] To do this, the test oracle 824 applies a set of predefined rules, which can be referred to herein as the "Digital Highway Code (DHC)". Essentially, rules defining safe driving behavior are hard-coded. If the scenario is completed without violating those rules, the AV is considered to have passed. However, if any of those rules are violated, the AV is considered to have failed and is marked as an instance of unsafe behavior that requires further testing and analysis. Those rules are encoded at the ontological level such that they can be applied to the ontological description of the scenario. The concept of ontology is well-known in the robotics field and, in this context, is intended to represent the driving scenario and the behavior of the simulated AV in that scenario at the same level of abstraction, such that the DHC rules can be applied by the test oracle 824. The results of the analysis can quantify how well the subject is performing relative to the DHC, e.g., the degree of rule violation (e.g., the rule can specify maintaining a certain distance from a cyclist at all times, and the result can indicate the degree of violation of that rule and the circumstances of the violation).
[0218] Instances of unsafe behavior can also be marked as instances that require "disengagement". For example, this can be the case where a failover mechanism within the runtime stack 100 is activated to prevent a crash or some other critical failure (as would be the case in that real-world scenario).
[0219] The present technique is not limited to detecting unsafe behavior. Behavior can be evaluated according to other metrics, such as comfort, progress, etc.
[0220] Example Perception Stack
[0221] Figure 9 A schematic block diagram showing a portion of an example perception stack. A 3D object detector is shown and denoted by reference numeral 204, and the 3D object detector is further shown as including a 2D object detector 902, a 2D tracker filter 904, a size estimation component 906, an orientation estimation component 908, a depth segmentation component 910, and a template fitting component 912. This represents an example architecture of the 3D object detector 204 mentioned above and shown in earlier figures.
[0222] The 2D object detector receives one image (the right image R in this example) of each captured stereo image pair and applies 2D object detection to the image. The output is the 2D bounding boxes of each object detected in the image. This provides the 2D (x, y) position of each object in the image plane and the bounding boxes indicating the size of the projection of the object into the image plane. The 2D tracking filter 904 receives the 2D bounding box output and applies filtering to them to refine the 2D bounding box estimates. For example, based on an object behaviour model, the filtering can take into account previous 2D detected bounding boxes and the expected behaviour of the detected objects. The filtered 2D bounding boxes and the image data of the original images contained therein are then used for many different purposes. The 2D object detector 902 can take the form of a trained CNN.
[0223] The depth segmentation component 910 receives the filtered 2D bounding boxes and also receives the depth map extracted from the original stereo image pair by the stereo estimator 202. It uses the filtered 2D boxes to separate the depth points belonging to each object within the depth map. This is a form of depth segmentation.
[0224] The size estimation component 906 also receives the filtered 2D bounding boxes and uses them to estimate the 3D size of each detected object based on the image data of the right image contained within the 2D bounding boxes.
[0225] The direction estimation component 908 similarly receives the filtered 2D bounding boxes and uses them to determine the 3D direction of each detected object using the image data of the right-side image contained within the applied 2D bounding boxes. The size estimation component 906 and the direction estimation component 908 can take the form of a trained CNN.
[0226] For each detected object, the 3D template fitting component 912 receives the separated depth points of the object from the depth segmentation component 910, the 3D size of the object from the size estimation component 906, and the 3D direction of the detected object from the direction component 908. The 3D template fitting component 902 uses that three pieces of information to fit a template in the form of a 3D bounding box to the depth points belonging to the object. Both the 3D size and the 3D direction of the 3D bounding box are known from the size and direction estimation components 906, 908 respectively, and the points to which the bounding box must be fitted are also known. Therefore, this is simply a matter of finding the best 3D position of the 3D bounding box. Once this is done for each object, the 3D size and 6D pose (3D position and 3D direction) of each detected object at a given moment can be known.
[0227] Shows the input to output from the 3D template fitting component 912 to the final filter 205. Additionally, the final filter 205 is shown as having inputs receiving the perception outputs from the lidar and the radar respectively. The lidar and radar perception components are shown and are denoted by reference numerals 914 and 916 respectively. Each of these provides a perception output that can be fused with the perception output (such as the 6D pose) from the 3D object detector 204. This fusion occurs in the final filter 205, and the output of the final filter is shown as being connected to the input of the prediction stack 104. For example, this can be the filtered (refined) 6D pose that takes into account all these stereo, lidar, and radar measurements. It can also take into account the expected object behavior in 3D space as captured in the expected behavior model of the 3D object.
[0228] Slices of the Perception Stack
[0229] Figure 9A Shows how to Figure 9An example of a "slice" of the perception stack, which is modeled as a PSPM. The perception stack 102 is considered to be sliced after the final perception component modeled by the PSPM, and the perception output of that perception component can be referred to as the "final output" for the PSPM. The distribution of the PSPM will be defined over those final outputs, i.e., the e in p(e|t, c) corresponds to those final outputs of the component after which the perception stack 102 is sliced. The PSPM models all the perception components and the sensors that provide input to that component (directly or indirectly) based on the influence of all perception components and sensors on the uncertainty in the final output e (and is said to "wrap" those perception components and sensors).
[0230] In this case, a single PSPM is provided for each sensor modality, i.e., one for stereo imaging, a second for LiDAR, and a third for RADAR. The three PSPMs are denoted by reference numerals 500a, 500b, and 500c, respectively. To construct the first PSPM 500a, the perception stack 102 is sliced after the 3D template fitting component 912. Thus, the distribution of the first PSPM 500a is defined over the perception output of the template fitting component 912. All the perception components and sensors fed into the 3D template fitting component 912 are wrapped in the first PSPM 500a. The second and third PSPMs 914, 916 are sliced after the LiDAR and RADAR perception components 914, 916, respectively.
[0231] The final filter 205 is not modeled as a PSPM but is applied to the PSPM samples obtained from the three PSPMs 500a, 500b, and 500c during testing.
[0232] Figure 9B A second example slice is shown, where a single PSPM 500d is used to model all three sensor modalities. In this case, the distribution p(e|t, c) is defined over all three sensor modalities, i.e., e = (e stereo , e lidar e lidar ). Thus, each PSPM sample will include the perception outputs for all three sensor modalities. In this example, the final filter is still not modeled as a PSPM and will be applied during testing to the sampled PSPMs obtained using the single PSPM 500d.
[0233] Figure 9CShows a third example slice where all three sensor modalities along with the final filter 205 are modeled as a single PSPM 500e. In this case, a distribution p(e|t, c) is defined over the filtered perceptual output of the final filter 205. During testing, the PSPM 500e will be applied to the ground truth derived from the simulation and the resulting PSPM samples will be fed directly into the prediction stack 104.
[0234] Slice Considerations
[0235] A factor when deciding where to "slice" the perception stack is the complexity of the ground truth required (the required ground truth will correspond to the perceptual components after the stack is sliced): The potential motivation for the PSPM approach is to have a ground truth that is relatively easy to measure. The lowest part of the perception stack 102 operates directly on the sensor data, but the information level required for planning and prediction is much higher. In the PSPM approach, the idea is to "bypass" the lower-level details while still providing a statistically representative perceptual output for prediction and planning during testing. Broadly speaking, the higher the perception stack 102 is sliced, the simpler the ground truth generally is.
[0236] Another consideration is the complexity of the perceptual components themselves, since any perceptual components not wrapped within a PSPM must be executed during testing.
[0237] It is generally expected that the slicing will always occur after the CNN in the perception stack, thus avoiding the need to simulate the input to the CNN and avoiding using the computational resources to run the CNN during testing.
[0238] In a sense, it is beneficial to wrap as much of the perception stack 102 into a single PSPM as possible. In the extreme case, this means modeling the entire perception stack 102 as a single PSPM. The benefit of doing this is the ability to model any correlations between different sensors and / or perceptual components without having to know about these correlations. However, as more and more of the perception stack 102 is wrapped within a single PSPM, this significantly increases the complexity of the system being modeled.
[0239] For Figure 9A , each individual PSPM 500a, 500b, 500c can be constructed independently of the data of a single sensor modality. This has the benefit of modulation - existing PSPMs can be rearranged to test different configurations of the perception slice 204 without having to retrain. Finally, the optimal PSPM architecture will depend on the context.
[0240] Especially in Figure 9CIn such cases, it may also be necessary to use a time-dependent model to fully capture the dependence on the previous measurement / perception output introduced by the final filter 205. For example, Figure 9C The PSPM500e of
[0241] For Figure 9A and Figure 9B , cutting before the final filter 205 has the benefit that it may not be necessary to introduce explicit time dependence, i.e., a form of PSPM can be used that has no explicit dependence on the PSPM samples previously obtained from the PSPM.
[0242] Example of PSPM
[0243] The above description has mainly focused on dynamic objects, but PSPMs can also be used in the same way for static scene detectors, classifiers, and other static scene perception components (e.g., traffic light detectors, lane offset correction, etc.).
[0244] In fact, PSPMs can be built for any part of the perception stack 102, including:
[0245] - Odometry, such as:
[0246] o IMU,
[0247] o Visual-odometry,
[0248] o LIDAR-odometry,
[0249] o RADAR-odometry,
[0250] o Wheel encoders;
[0251] -(Self-) Localization, such as:
[0252] o Vision-based localization,
[0253] o GPS localization (or more generally satellite localization).
[0254] "Odometer" refers to the measurement of local relative motion, while "Localisation" refers to the measurement of global position on a map.
[0255] PSPM can be constructed in exactly the same way to model the perception output of such perception components using appropriate perception ground truth.
[0256] These allow realistic odometry and localization errors to be introduced into the simulated scenario in the same way as detection errors, classification errors, etc.
[0257] Ground Truth Pipeline
[0258] As described above, the generation of annotations in the ground truth pipeline 802 can be manual, automatic, or semi-automatic annotation.
[0259] Automatic or semi-automatic ground truth annotation can use high-quality sensor data that is not usually available (or at least not available all the time) during runtime in the AV. In fact, this can provide a way to test whether such components are needed.
[0260] Automatic or semi-automatic annotation can use offline processing to obtain a more accurate perception output, which can be used as the ground truth for PSPM construction. For example, to obtain the perception ground truth for localization or odometry components, offline processing such as bundle adjustment can be used to reconstruct the vehicle's path with high precision, which can then be used as the ground truth to measure and model the accuracy of the AV's online processing. Due to computational resource limitations or because the algorithms used are inherently non-real-time, this offline processing may not be feasible on the AV itself during runtime.
[0261] Examples of Confounding Factors
[0262] Figure 10 A high-level overview of the various factors that can lead to uncertainty in the perception output (i.e., the various sources of potential perception errors) is shown. This includes further examples of confounding factors c that can be incorporated as variables in the PSPM:
[0263] - Occlusion
[0264] - Lighting / Time of day
[0265] - Weather
[0266] - Season
[0267] -(Linear and / or angular) distance to the object
[0268] -(Linear and / or angular) velocity of the object
[0269] -Position in the sensor's field of view (e.g., angle from the center of the image)
[0270] -Other object characteristics, such as reflectivity, or other aspects of its response to different signals and / or frequencies (infrared, ultrasound, etc.)
[0271] Other examples of possible confounding factors include a map of the scene (indicating environmental structure) and inter-agent variables, such as "business" (a measure of the number or density of agents in the scene), the distance between agents, and the type of agent.
[0272] Each can be numerically or categorically characterized in one or more variable components (dimensions) of the confounder space C.
[0273] However, note that a confounding factor can be any variable that represents something about the physical world that might be related to the perception error. This doesn't necessarily have to be a directly measurable physical quantity such as speed, occlusion, etc. For example, another example of a confounding factor related to another actor might be "intent" (e.g., whether a cyclist intends to turn left or continue straight ahead at an upcoming turn at a particular moment, which can be determined from real-world data of the cyclist's actual actions at a given time by looking ahead). In a sense, variables such as intent are latent variables or unobserved variables, in that at a particular moment (in this case, before the cyclist takes a definitive action), the intent cannot be directly measured using the perception system 102 but can only be inferred through other measurable quantities; the point about confounding factors is that it is not necessary to know or measure those other measurable physical quantities in order to model the effect of intent on the confounding factor error. For example, the perception error associated with a cyclist with an "intend to turn left" intent might be statistically significantly higher compared to a cyclist with an "intend to continue straight" intent, which might be due to multiple, unknown, and potentially complex behavioral changes in the actions of the cyclist intending to turn left, meaning that in practice, the perception system perceives them worse. By introducing the "intent" variable as a confounding factor in the error model, there is no need to try to determine which observable physical manifestations of intent are related to the perception error - as long as the "intent" ground truth can be systematically assigned to the training data in a way consistent with the simulation (in this case, the intent of the cyclist is known in the simulation in order to simulate their behavior as the scene unfolds), then such data can be used to build appropriate behavioral models for different intents in order to simulate that behavior, as well as a perception error model that depends on intent, if any, without having to determine which physical manifestations of intent are actually related to the perception error. In other words, in order to model the effect of intent on the perception error, it is not necessary to understand why intent is related to the perception error, because intent itself can be modeled as a perception confounding factor (rather than trying to model the observable manifestations of intent as confounding factors).
[0274] Low Level Errors
[0275] Examples of low-level sensor errors include:
[0276] - Registration errors
[0277] - Calibration errors
[0278] - Sensor limitations
[0279] Such errors are not explicitly modeled in the simulation, but their effects are wrapped in the PSPM used to model the perception slices that interpret the applicable sensor data. That is, these effects will be encoded in the parameters θ that characterize the PSPM. For example, for a Gaussian-type PSPM, such errors result in larger covariance representing greater uncertainty.
[0280] High-level Perception Errors
[0281] Other errors may occur in the perception pipeline, such as:
[0282] - Tracking errors
[0283] - Classification errors
[0284] - Dynamic object detection failures
[0285] - Fixed scene detection failures
[0286] When it comes to detection, false positives and false negatives can cause the prediction stack 104 and / or the planner 106 to operate in an unexpected manner.
[0287] Construct specific PSPMs in a statistically robust manner to model such errors. These models can also take into account the effects of the variable confounding factor c.
[0288] Taking object detection as an example, detection probabilities can be measured and used to construct a detection distribution that depends on, for example, distance, angle, and occlusion level (the confounding factor c in this example). Then, when running the simulation, through ray tracing from the camera, it can be determined according to the model whether an object "might" be detectable. If so, the measured detection probability is checked, and the object is detected or not detected. This deliberately introduces the possibility of not detecting sensor-sensitive objects in the simulation in a way that reflects the behavior of the perception stack 102 in real life, because the detection failures have been modeled in a statistically robust manner.
[0289] This method can be extended in a Markov model to ensure proper modeling of conditional detection. For example, an object can be detected with an appropriate probability only if the object has been previously detected, otherwise the probability may be different. In this case, false negatives involve some temporal dependence on the simulated detections.
[0290] False positives can be randomly generated with a density similar to that of the density in space and time measured by PSPM. That is, in a statistically representative way.
[0291] 2. Problem Statement
[0292] As a further explanation, this section elaborates on the mathematical framework of PRISM and introduces the specific dynamic object detection problems to be addressed in the subsequent sections. Section 3 discusses the datasets used for training PRISM, the techniques for identifying relevant features, and the description of the evaluation methods. Section 4 describes the specific modeling decisions and how data science informs these decisions.
[0293] Note that in the following description, the symbols x g , y g , z g can be used to represent the coordinates of the position-aware ground truth t. Similarly, x s , y s , z s can be used to represent the coordinates of the position-aware stack output e. Thus, the distribution p(x s , y s , z s |x g , y g , z g ) is a form that the above-mentioned perception uncertainty distribution p(e|t) can take. Similarly, x can be used below to generically refer to the set of confounding factors, which is equivalent to the set of confounding factors c or c' described above.
[0294] The perception system has inputs that are difficult to simulate, such as camera images, lidar scans, and radar echoes. Since these inputs cannot be rendered with perfect realism, the perception performance in the simulation will not match that in the real world.
[0295] The goal is to construct a probabilistic surrogate model for the perception stack, called PRISM. PRISM uses a low-fidelity representation of the world state (perceptual ground truth) and produces perceptual outputs in the same format as the vehicle stack (or more precisely, the modeled perceptual slice 204). When the stack runs on real data, samples drawn from the surrogate model in the simulation should be similar to the outputs of the perception stack.
[0296] PRISM sampling should be fast enough to be used as part of a simulation system for the validation and development of downstream components such as planners.
[0297] 2.1 Intuition
[0298] The following sections state the most general case for the following considerations:
[0299] ● There is some stochastic function that maps from the true state of the world to the output of the perception stack.
[0300] ● This function can be modeled using training data. The function is modeled as a probability distribution.
[0301] ● Since the world state changes smoothly over time, the sampled perceptual outputs should also change smoothly over time. Since the world state is only partially observed, the appropriate way to achieve this is to make the probability distribution depend on the observed world state and the history of perceptual outputs.
[0302] · The simulator (Genie) is responsible for generating a representation of the world at runtime. The output of Genie is the 6D pose and range of dynamic objects, as well as some other information such as road geometry and weather conditions.
[0303] · For real-world training data, this world representation is obtained from annotations.
[0304] Mathematical statement
[0305] 2.2.1 Preliminaries
[0306] For any set S, let the set of histories of S be An element (t, h) ∈ histories(S) consists of t (the current time) and h (a function that returns elements of S at any time in the past). The notation denotes the simulation equivalent of x.
[0307] The sensing system is the stochastic function f: histories(World) → histories(Perception). In general, the form of f will be:
[0308]
[0309] sense: histories(World) → histories(SensorReading),
[0310] perceive: histories(SensorReading) → histories(Perception). (1)
[0311] The goal is to simulate some f. The world state can be decomposed into a set of properties ObservedWorld and a set of everything else UnobservedWorld (the exact pixel values of the camera image, the temperature at each point on each surface), such that there is a bijection between World and ObservedWorld × UnobservedWorld, and the set of properties ObservedWorld can be reliably measured (this may include the meshes and textures of each object in the scene, the positions of the light sources, the material density, etc.). In traditional photorealistic simulation methods, simulating f amounts to finding some stochastic function
[0312] : histories(ObservedWorld) → histories(SensorReading), which can be combined with perceive to form
[0313]
[0314] Let observe ∶ World → ObservedWorld be the function that maps world states to their observed counterparts. Note that this function is not one-to-one: there will be many world states that map to a single observed world state. An accurate and useful simulation of f For all histories (t, h) ∈ histories(World) will have
[0315]
[0316] where map ∶ ((S → T) × histories(S)) → histories(T) maps a function over histories.
[0317] One must then conclude that the best photorealistic simulation has such that
[0318]
[0319] Since the associativity, combinatorial equalities 1, 2, and 4 give equality 3 through , sense predicts the history of sensor readings, the joint distribution of history(SensorReading), and the correlations between different sensor readings enable the dependencies on unobserved properties of the world to be modeled more effectively. Therefore, similar correlations should be observed in the computation.
[0320] Because SensorReading has high dimensionality and sense is a stochastic function (since it depends strongly on unobserved properties of the world), it is very important to find such that equation 4 holds even approximately. Therefore, one can directly find
[0321] 2.2.2 Creating a surrogate model
[0322] The creation of the surrogate model can be characterized as a stochastic function estimation task. Let S+ be the set of finite sequences of elements of S. Let be a sequence of length N with elements s i . Obtain a data set of sequences of sensor readings
[0323]
[0324] where each I ij ∈ SensorReading is the sensor reading at time t ij in run i, and M i is the number of timestamps in a particular run. Use the function annotate∶SensorReading→ObservedWorld that recovers the observed scene parameters from the sensor readings to construct a new data set.
[0325]
[0326] Then, the task of PRISM is to estimate the distribution from samples in The implementation of can be obtained by sampling from this distribution.
[0327] For the previously sampled stack output Dependencies are included because the distribution of y meaningfully depends on the unobserved world, and the unobserved world changes smoothly over time. As discussed in Section 2.2.1, this dependence on the unobserved world means that y will change smoothly over time in a way that is difficult to model solely based on the dependencies. This time-related property of the stack output was explored for the perception system 102 in Section 4.2.3, where strong correlations over time were found.
[0328] Samples from the learned PRISM distribution give reasonable perception outputs conditioned on the low-fidelity scene representation and the history of previous samples. These factors are the independent variables in the generative model, and the dependent variable is the perceived scene. Independent variables that meaningfully affect the distribution of the dependent variable are called confounders in this paper. Part of the process of constructing the PRISM model is to identify the relevant confounders to include in the model and how these confounders should be combined. Section 3.2 explores a method for identifying relevant confounders.
[0329] 2.2.3 Dynamic Object Problem
[0330] A specific example of a perception system is presented—a system that uses RGBD images to detect dynamic objects in a scene. A "dynamic object" is a car, truck, cyclist, or other road user described by an oriented bounding box (6D pose and extent). The observed world is a collection of such dynamic objects. In this setting,
[0331]
[0332] SensorReading = Image = [0, 1] w×h×4 ,
[0333] where is the set of finite subsets of S, 1 Type represents the object type (Car, Van, Tram, Pedestrian), Spin(3) is the set of unit quaternions, and Info is an arbitrary set whose elements describe additional characteristics of the dynamic object, e.g., the degree to which the object is occluded by other (possibly static) objects in a scene closer to the camera. Dynamic objects are useful when characterizing the behavior of the perception system.
[0334] This example further simplifies the dynamic object problem by only choosing to model the positions of dynamic objects for a given ObservedWorld. This includes fitting a model for the possibility that an observable object is not perceived (false negative).
[0335] As shown in Section 4.2.8, false negatives are more common errors made by the perception system 102 than false positives (spurious dynamic object detections).
[0336] For simplicity, the following description only considers hazardous objects (poisons) in 3D space and omits the discussion of directions, ranges, object types, or other possible perception outputs. However, the principle can equally be applied to such other perception outputs.
[0337] 3 Method
[0338] 3.1 Data
[0339] A specific driving scenario is presented, and the data for this specific driving scenario have been recorded multiple times under similar conditions. The scenario mentioned herein by way of example is a roundabout located southeast of London on a test route. Figure 12 The background of the roundabout and the path of vehicles passing through can be seen therein, where it is observed from a camera as shown Figure 13 in the figure.
[0340] By restricting the PRISM training data to running on the same roundabout under similar climatic conditions, the influence of weather and sunlight as confounding factors of perception performance is minimized. The potential performance of PRISM tested on similar collected data is equally maximized. For example, by estimating how PRISM trained on roundabout data performs in a highway scenario, the performance of PRISM can be tested on out-of-domain data.
[0341] 3.1.1 Dataset Generation
[0342] PRISM training requires a dataset containing sufficient information to learn the distribution of perception errors. For simplicity, this section only considers the errors introduced by the perception system 102 when predicting the central position of a dynamic object in the camera frame of the observed dynamic object. To learn such errors, ground truth center and perceived center estimates are required.
[0343] The ground truth center positions are estimated from the human-annotated 3d bounding boxes present in each frame of the recorded video sequence of the roundabout and applied to all dynamic objects in the scene. These bounding boxes are fitted to the scene using a ground truth tooling suite. The ground truth tooling suite combines camera images, stereo depth pointclouds, and lidar pointclouds into a 3D representation of the scene to maximize annotation accuracy. It is assumed that the annotation accuracy is good enough to be used as ground truth.
[0344] Figure 9 The process of obtaining stack prediction objects from the recorded camera images is shown. It should be noted that the pipeline is stateless and each pair of camera frames is processed independently. This forces any temporal correlations found in the perception error data to be attributed to the behavior of the detector on closely related inputs rather than the internal state of the detection algorithm.
[0345] In general, the set of object predictions indexed by image timestamps combined with a similarly indexed set of ground truth data from the ground truth tooling suite is sufficient for PRISM training data. However, all models considered in this section are trained on data that has been passed through additional processing steps to generate associations between ground truth and predicted objects. This limits the space of alternative models but simplifies it by breaking the fitting task into the following parts: a model for fitting position error; a model for fitting false negatives; a model for fitting false positives. The association algorithm used runs independently on each frame. For each timestamp, the set of stack prediction objects and the set of ground truth objects are compared using intersection over union (IOU), where the prediction object with the highest confidence score (a measure indicating how good the prediction is generated by the perception stack 102) is considered first. For each predicted object, the ground truth object with the highest IOU is associated with it, forming pairs for learning the error distribution. Pairs with an IOU score less than 0.5 (an adjustable threshold) do not form any associations. After all predicted objects have been considered for association, the set of unassociated ground truth objects and the set of unassociated predicted objects are retained. The unassociated ground truth objects are stored as false negative examples, while the unassociated predicted objects are stored as false positive examples.
[0346] "Set-to-set" models that do not require such associations will be considered later.
[0347] 3.1.2 Content of Training Data
[0348] The previous section described how to generate PRISM training data and divide it into three sources: associations, false negatives, and false positives. Table 1 specifies the data present in each source and provides the following definitions:
[0349] centre_x, centre_y, centre_z: The x, y, and z coordinates of the center of the ground truth 3D box.
[0350] orientation_x, orientation_y, orientation_z: The x, y, and z components of the axis-angle representation of the rotation from the camera frame (right of the front stereo) to the ground truth 3D box frame.
[0351] height, width, length: The extent of the ground truth 3D box along the z, y, and x axes in the coordinate system of the 3D box.
[0352] manual_visibility: A label applied by a human annotator to indicate which of four visibility categories the ground truth object belongs to. The categories are: fully-occluded (100%), largely-occluded (80 - 99%), somewhat-occluded (1 - 79%), and fully-visible (0%).
[0353] occluded: The portion of the area where the ground truth 2D bounding box overlaps with the 2D bounding box of another ground truth object closer to the camera.
[0354] occluded_category: A combination of manual_visibility and occluded, which can be considered the maximum of the two. Combining manual_visibility and occluded in this way is useful for maximizing the number of correct occlusion labels. To see this, note that the occluded score of an object occluded by the static parts of the scene (bushes, trees, traffic lights) will be 0, but will have a manual_visibility correctly set by the human annotator. Objects occluded only by other ground truth objects do not have a manual_visibility field set by the human annotator, but will have a correct occluded field. These two cases can be handled by taking the maximum of the two values. Even with this logic, it is possible for the 2d bounding box of a ground truth object to completely obscure the 2d bounding box of an object behind it, even if some of the background objects are visible. This will generate some fully-occluded cases that can be detected by the perception system.
[0355] truncated: The portion of the eight vertices of the ground truth 3d box that lies outside the sensor frustum.
[0356] type: When attached to a ground truth object (false negative, the ground truth part of an association pair), this is the object type annotated by a human, such as Car or Tram. When attached to a predicted object (false positive, the predicted part of an association pair), this is the perception stack's best guess of the object type, limited to Pedestrian or Vehicle.
[0357] In addition to the above, this section will also mention the following derived quantities:
[0358] distance: The distance from the object center to the camera, calculated as the Euclidean norm of the object center position in the camera frame.
[0359] azimuth: The angle formed between the projection of the ray connecting the camera and the object center onto the camera's y = 0 plane and the camera's positive z-axis. The polarity is defined by the direction of rotation around the camera y-axis. Since objects behind the camera cannot be observed, the range is limited to [-π / 2, π / 2].
[0360] Table 1
[0361]
[0362] The composition of the dataset will be discussed in detail in the relevant section later. The following presents a high-level summary of the data.
[0363] · 15 traversals of the roundabout scenario, spanning a total footage of approximately 5 minutes.
[0364] · 8,600 unique frames containing 96k ground truth object instances visible to the camera.
[0365] · Among these 96k instances: 77% are cars; 14% are vans; 6% are pedestrians; 3% belong to smaller groups.
[0366] · Among these 96k instances: 29% are fully visible; 43% are somewhat occluded; 28% are mostly occluded.
[0367] In Table 1, specific data elements exist in each of the three generated PRISM data sources. X indicates that the column exists in the given data source. GT = ground truth, FN = false negative, FP = false positive. Each of these is actually three independent variables (e.g., center_x, center_y, center_z), but for readability, they are "compressed" here. *In the case with an asterisk, the type of content can be "Vehicle" or "Pedestrian", which are the only categories predicted by the five perception stacks. In the case without an asterisk, there are more classes (such as "Lorry" and "Van"), which are all the classes reported in the ground truth data.
[0368] 3.1.3 Training and Test Data
[0369] For all the modeling experiments described in this article, the roundabout dataset is split into roughly equal halves to form a training set and a test set. No hyperparameter optimisation is performed, so a validation set is not required.
[0370] 3.2 Identifying Relevant Confounding Factors
[0371] The PRISM model may consider many confounding factors. Instead of optimizing the model for every possible combination of confounding factors, it is preferred to perform such optimization on a limited set of known relevant confounding factors.
[0372] To identify relevant confounding factors, a Relief-based algorithm is used. The general outline of the Relief-based algorithm is given in Algorithm 1. The Relief algorithm produces an array of feature weights in the range [-1, 1], where weights greater than 0 indicate that the feature is relevant as changes in the feature tend to change the target variable. In practice, some features will accidentally have weights greater than 0, and only features with weights greater than some user-defined cutoff 0 < τ < 1 are selected.
[0373]
[0374] This algorithm has the following desirable properties:
[0375] · It is sensitive to non-linear relationships between features and the target variable. Other feature selection methods are not sensitive to these kinds of relationships, such as naive principal component analysis or comparison of Pearson correlations. Not all uncorrelated things are independent.
[0376] · It is sensitive to interactions between features.
[0377] · It is conservative. It will accidentally include uncorrelated or redundant confounding factors rather than accidentally excluding relevant confounding factors.
[0378] It is important to note the following considerations of this method:
[0379] ● It identifies correlations in the data but does not provide in-depth understanding of how or why the target variable is related to the confounding factors under investigation.
[0380] · The results are dependent on the parameterization of the confounding variables.
[0381] There are many extensions of the Relief algorithm. Here, an extension called MultiSURF is used. It is found that MultiSURF performs well in a wide range of problem types and is more sensitive to interactions of three or more features than other methods. The implementation is used from scikit-rebate, an open-source Python library that provides implementations of many Relief-based algorithms, which are extended to cover scalar features and target variables.
[0382] In the experiment, use where n is the size of the data set, and α = 0.2 is the desired false discovery rate. According to Chebyshev's inequality, we can say that the probability of accepting an irrelevant confounding factor as relevant is less than α.
[0383] Relief-based methods are useful tools for identifying possible confounding factors and their relative importance. However, not all features that affect the error characteristics of the perception system will be captured in the annotated training data. A manual process of examining the model's failure to assume new features as confounding factors is necessary.
[0384] 4 Model
[0385] 4.1 Heuristic Model
[0386] The camera coordinates represent the position of a point in the image in pixel space. In binocular vision, the camera coordinates of points in two images are available. This allows the reconstruction of the position of points in the 3D Cartesian world. The camera coordinates of a point p in 3D space are given by:
[0387]
[0388] where (u1, v1), (u2, v2) are the image pixel coordinates of p in the left and right cameras respectively, (x p , y p , z p ) are the 3D world coordinates of p relative to the left camera, b is the camera baseline, and f is the camera focal length. As Figure 14 shown. The disparity d is defined as:
[0389]
[0390] The 3D world coordinates of p can be written as:
[0391]
[0392]
[0393] The heuristic model is obtained by imposing a distribution in camera coordinates and propagating it to 3D coordinates using the above relationships. This distribution can equally be used for object center or object extent. When the image is discretized into pixels, the model allows for the consideration of the physical sensor uncertainty of the camera. The model is given by:
[0394] p(x s , y s , zs | x g , y g , z g ) = ∫∫∫ p(x s , y s , z s | u1, v, d) p(u1, v, d | x g , y g , z g ) du1 dv d d, (12)
[0395] where (x g , y g , z g ) are the coordinates of the ground truth points, and (x s , y s , z s ) are the coordinates of the stack prediction. Given the world coordinates, the probability distribution on the camera coordinates is: p(u1, v, d | x g , y g , z g ) = p(u1 | x g , y g , z g ) p(v | x g , y g , z g ) p(d | x g , y g , z g ) , (13)
[0396] where, assuming independence of the distributions in each camera coordinate:
[0397]
[0398] where σ is a constant, is the normal distribution, and Lognormal is the log - normal distribution (log - normal distribution). The Lognormal is chosen because it only supports positive real numbers. This defines the probability density of a normal distribution centered on the camera coordinates of points in 3D space. The normal distribution is chosen for mathematical simplicity. If only discretisation error is considered, a uniform distribution might be more appropriate. However, other errors are likely to lead to uncertainties in stereo vision, so the extended tails of the normal distribution are useful for modelling such phenomena in practice. For a front - facing stereo camera, α can be determined to be 0.7 by maximum likelihood estimation. p(x s , y s, z s |u1, v, d) is given by a Dirac distribution centered at the point values of x obtained from Equation 9 - Equation 11 s , y s and z s The point values of which are given by a Dirac distribution centered at the point values of x, y, and z
[0399] By forming a piecewise - constant diagonal multivariate normally distributed approximation of Equation 12, by solving the integral using Monte Carlo simulation, and by using the mean and variance of the sampled values for different x g , y g and z g values to estimate p(x s , y s , z s |x g , y g , z g ), a runtime model is obtained
[0400] The model can be improved by considering a more accurate approximation of the conditional distribution in Equation 12, or by modeling the uncertainty in the camera parameters f and b (set to their measured values in the model). How to extend the model to include time - dependence is an open question
[0401] 4.2 PRISM
[0402] An attempt to build a Probabilistic Reasoning about Implicit Statistical Models (PRISM) of the perception stack / sub - stack 204 guided by data analysis is described below. The model includes a non - zero probability of positional error with time - dependence and of objects not being detected, which are significant features of the data
[0403] 4.2.1 Positional errors
[0404] The central position of a dynamic object detected by the perception stack will be modeled using an additive error model given by:
[0405] y k = x k + e k
[0406] where y k is the observed position of the object, x k is the ground - truth position of the object, and e kis the error term, all at time t k The phrase "position error" will be used to refer to the additive noise component e of the model k .
[0407] Figure 15 shows the position error of a particular dynamic object detected by the perception stack relative to the human-labeled ground truth. A lag plot of the same data can be found in Figure 16 , indicating the strong temporal correlation of these errors. From these plots, it can be concluded that the generative model of the position error must condition each sample on the previous samples. An autoregressive model is proposed for the time-correlated position error, where each error sample linearly depends on the previous error samples and some noise. The proposed model can be written as:
[0408] e k = e k-1 + Δe k (18)
[0409] where e k is the position error sample at time step k, and Δe k is the random term, and Δe k can be a function of one or more confounding factors, often referred to as "error deltas". Figure 17 shows a diagram visualizing the model, including the dependencies on the hypothesized confounding factors C1 and C2.
[0410] The model is based on several assumptions. First, the subsequent error deltas are independent. This is explored in Section 4.2.3. Second, the empirical distribution of the error deltas can be reasonably captured by a parametric distribution. This is explored in Section 4.2.4. Third, the described model is stationary such that the mean error does not change over time. This is explored in Section 4.2.5.
[0411] 4.2.2 Piecewise Constant Model
[0412] It has been shown that modeling the position error requires subsequent errors to be conditioned on the previous errors, but how should the first error sample be selected? Now consider the task of fitting a time-independent position error distribution. If no temporal correlation is found in the data, the method adopted here can equally be applied to all samples of each dynamic object, not just to the first sample.
[0413] In general, such a model would be a complex joint probability distribution of all confounding factors. As discussed in Section 2, due to the incomplete scene representation in the perception stack (ObservedWorld≠World) and possible uncertainties, the distribution of possible perceptual outputs given the ground truth scene is expected. The expected variance is heteroskedastic; it varies based on the values of the confounding factors. As a simple example, it should not be surprising that the error in the position estimation of a dynamic object has a variance that increases with the distance of the object from the detector.
[0414] The conditional distribution modeled by PRISM is expected to have a complex functional form. This functional form can be approximated by discretizing each confounding factor. In this representation, categorical confounding factors (e.g., vehicle type) are mapped to bins. Continuous confounding factors (e.g., distance from the detector) are split into ranges, and each range is mapped to a bin. The combination of these discretizations is a multi-dimensional table for which an input set of confounding factors is mapped to bins. It is assumed that within each bin, the variance is homoskedastic and a distribution with constant parameters can be fit. The global heteroskedasticity is captured by different parameters in each bin. A model of a distribution with fixed parameters in each bin is called a Piecewise Constant Model (PCM) in this paper. Examples of general implementations of similar models can be found in the literature. Mathematically, this can be written as P(y|x) ∼ G(α[f(x)], β[f(x)],...), where y is the set of outputs, x is the set of confounding factors, f(·) is the function that maps the confounding factors to bins, and G is a probability distribution with parameters α[f(x)], β[f(x)],... fixed within each bin.
[0415] In the PCM for PRISM, it is assumed that the error is additive, i.e., the predicted position, pose, and range of the stack of dynamic objects equals the ground truth position, pose, and range plus some noise. The noise is characterized by a distribution in each bin. In this PCM, it is assumed that this noise is normally distributed. Mathematically this can be written as:
[0416]
[0417] where y is the stack observation, is the ground truth observation, and ∈ is the noise. The distribution in each bin is characterized by a mean μ and a covariance ∑. μ and ∑ can be regarded as functions of the confounding factor bin.
[0418] Figure 18 An example binning scheme is shown. A bin is composed of an azimuth angle and a distance to the center of the ground truth dynamic object.
[0419] Training the model requires the ground truth and stack predictions (actual perception outputs) collected as described in Section 3.1.1. (For example, using the maximum a posteriori method to incorporate the prior) Fit the mean and covariance of the normal distribution to the observations in that bin. For the mean of the normal distribution, use the prior of the normal distribution. For the scale of the normal distribution, use the Inverse Gamma prior.
[0420] To set the hyperparameters of the prior, physical knowledge can be combined with the intuition regarding how quickly the model should ignore the prior when data becomes available. This intuition can be represented by the concept of pseudo-observations, i.e., how much the weighted strength of the prior distribution is compared to the actual observations (encapsulated in the likelihood function) in the posterior distribution. Increasing the number of pseudo-observations results in a prior with lower variance. The hyperparameters for the normal distribution prior can be set to μ h = μ p and where μ p and σ p represent the prior point estimates of the mean and standard deviation of the bin under consideration, and n pseudo represents the number of pseudo-observations. The rate and scale hyperparameters of the Inverse Gamma prior can be set to and For this model, choose n pseudo = 1, and use the heuristic model described in Section 4.1 to provide prior point estimates for the parameters of each bin.
[0421] The advantage of the PCM method is that it takes into account global heteroscedasticity, provides a unified framework for capturing different types of confounding factors, and it utilizes simple probability distributions. In addition, the model is interpretable: the distribution in the bin can be examined, the training data can be directly inspected, and there are no hidden transformations. Furthermore, the parameters can be analytically fitted, which means that the uncertainties from lack of convergence in the optimization routine can be avoided.
[0422] Confounding factor selection
[0423] To select suitable confounding factors for the PCM, the methods described in Section 3.2 and the data described in Section 3.1.2 were used. The findings applied to the position, range, and orientation errors are shown in Table 2.
[0424] Table 2 - shows the confounding factors identified as important for the target variable under consideration
[0425]
[0426] As can be seen from Table 2, for d_centre_x and d_centre_z, the relevant confounding factors are some combination of the object's position relative to the camera and the degree to which the object is occluded. The perception system 102 assumes that the detected object exists on the ground plane, y = 0, which may be the reason why d_centre_y does not show a dependence on distance.
[0427] For the model of the position error of dynamic objects detected by the perception system 102, this analysis identified position and occlusion as good confounding factors to start with. The data did not show a strong preference for position confounding factors based on Cartesian grids (centre_x, center_y, center_z) over polar coordinates (distance, azimuth). Distance and azimuth are used in the PRISM prototype described herein, but a more in-depth assessment of the relative performance of each prototype can be performed. 4.2.3 Temporal correlation analysis of position error increments
[0428] The temporal correlation analysis performed on the position error can be repeated for the time series of error increments, giving Figure 19 the lag plots shown in. These plots show much less temporal correlation in the error increments than was found in the position errors. The Pearson correlation coefficients for the error increments are shown in Table 3. For each dimension, they are quite small in magnitude, with -0.35 being the furthest from zero. From this analysis, it can be concluded that a good model of the error increments can be formed from independent samples from the relevant distribution.
[0429] Table 3 - Pearson correlation coefficients for error increment samples and samples with a one-time step delay
[0430]
[0431] Distribution of position error increments
[0432] Typically, the x, y, z error increment dimensions are correlated. Here, they are considered independently, but note that future efforts could consider jointly modeling them. Figure 20A histogram of the error increment samples is presented, from which it can be clearly seen that the error increments are more likely to be close to zero, but with long tails of extreme values. Figure 21 The maximum likelihood best-fit of this data to some trial distributions is shown. Visual inspection of these plots suggests that the Student's t-distribution could be a good modeling choice for generating the error increments. Due to the presence of a large number of extreme error increments in the data, the normal distribution does not fit well.
[0433] Bounding the random walk
[0434] The autoregressive error increment model proposed in Section 4.2.1 is generally a non-bounded stochastic process. However, it is well known that the detected position of a dynamic object does not simply deviate but remains near the ground truth. This is an important property that must be captured in a time-dependent model. As a specific example of this point, consider modeling the position error as a Gaussian random walk, setting This results in a position error distribution at time t k where the variance increases with time without bound. Such a property should not exist in the PRISM model.
[0435] AR(1) is a first-order autoregressive process defined by:
[0436] y t = a1y t-1 + ∈ t (19)
[0437] where ∈ t is a sample from a zero-mean noise distribution, and y t is a sample of the variable of interest at time t. This process is known to be generalized stationary for |a1| < 1, otherwise the generated time series is non-stationary. Comparing Equation 18 and Equation 19, it can be seen that, given the known results of AR(1), if Δe k has zero mean, the error increment model proposed in Equation 18 will be non-stationary. Therefore, such a model is not sufficient to generate a reasonable stack output.
[0438] An extension to the model proposal in Equation 18 is presented, motivated by the nature of the collected error increment data. The extension conditions Δe on the previous errork Perform modeling such that the best fit to P(Δe k |e k -1) is found. A model of this form should learn to sample error increments that move the position error towards zero, with the probability being larger the further the position error is from zero. This has been found to be true.
[0439] Following the piecewise-constant modeling method described in Section 4.2.2, P(Δe k |e k -1) is approximated as follows:
[0440] Form M bins for the space of e k-1 values, where the boundaries are {b0, b1,..., b M}.
[0441] Characterize a separate distribution P m (Δe k ) for each bin, where 0 < m < M represents the bin index.
[0442] Given the previous time-step position error e k-1 , the next error increment is drawn from P m (Δe k ), where B m -1 < e k-1 < B m .
[0443] Figure 22 Shows the sample mean computed on the PRISM training data with M = 5. The revealed trend is consistent with expectations and has the following intuitive explanation. Consider a series of error increment samples of the same polarity that have accumulated an absolute position error away from the ground truth. To make the overall process appear stationary, subsequent error increment samples of the same polarity should be less likely to change than in the direction towards the true object position. This observation helps to explain the negative Pearson coefficients presented in Table 3, which indicate that subsequent error increments slightly tend to reverse polarity.
[0444] The binning scheme for P m (Δe k ) suffers from the typical PCM drawback of low sample counts in extreme bins. A simple prior can be used to reduce this risk, for example, setting the mean of the distribution in each bin to follow μ m = -ae m , where e m is the center value of the m-th bin and a > 0. It is worth noting that if P m (Δe k ) is chosen to be a Gaussian distribution such that The time-dependent model then becomes:
[0445]
[0446] The time-dependent model is a typical AR(1) process and is stationary when a < 2. In practice, a good prior will require a ∼ 0, so such a model is stationary by construction.
[0447] 4.2.6 Simple Verification
[0448] It is useful to see whether the samples from the proposed time-dependent position error model reproduce the features that motivated its construction. Figure 23 A plot of the position error of a single dynamic object trajectory sampled from the learned distribution is shown in. Figure 24 A lag plot of the same data is shown. In both cases, the true perception error data is provided for visual comparison. The similarity between the PRISM samples and the observed stack data is encouraging. Clearly, more quantitative evaluations (which will be the subject of Section 5) are needed to make any meaningful claims of reasonableness.
[0449] False Negatives and False Positives
[0450] The perception system has failure modes beyond noisy position estimates of dynamic objects. The detector may fail to identify an object in the scene, i.e., a false negative, or it may identify an object that does not exist, i.e., a false positive. A surrogate model like PRISM must simulate the observed false negative and false positive rates of the detector. This section discusses the importance of false negative modeling, investigates which confounding factors affect the false negative rate, and presents two simple Markov models. It has been shown that using more confounding factors can produce Markov models with better performance and highlights some issues with doing so using piecewise constant methods.
[0451] A survey was conducted to determine the frequencies of true positives (TP), false negatives (FN), and false positives (FP). The results are summarized in Table 4. False negative events are significantly more numerous than false positives. The counts in Table 4 apply to all object distances. The false negative count at such a distance from the detector seems unfair, i.e., humans would have a hard time identifying. Introducing a distance filter on the events reduces the factor by which false negatives are more prevalent than false positives, but the difference is still significant. When considering objects with a depth less than 50m, the number of TP / FN / FP events is 34046 / 13343 / 843. Reducing the distance threshold to a depth of 20m, the number of TP / FN / FP events is 12626 / 1236 / 201.
[0452] Table 4 - Shows the numbers describing false positive and false negative events in the dataset
[0453]
[0454] 4.2.9 False Negative Modelling
[0455] Following the method described in Section 3.2, the importance of different confounding factors for false negatives was explored using the relief algorithm. Milts 2 for 20% of the training data samples selected randomly. The results are as Figure 25 shown. A 20% random sample of the data allows the algorithm to run with manageable memory usage. The target variable is the association category produced by the detector: associated or false negative. The associated category is called the association state. The same list of confounding factors as in Section 3.2 is used, where distance and azimuth are substituted for center_x, center_y, and center_z. It has been found that the binning scheme based on distance and azimuth is as good as the binning of the center values but with lower dimensions. In addition, occluded_category is used, which is the most reliable occlusion variable. In addition, the association state of the object in the previous time step is included as a potential confounding factor. This is labeled as "from" in Figure 25 . Note that "from" has three possible values: associated, false negative, and empty. When an object is first visible to the detector, there will be no previous association state; in the previous time step, the detector did not detect the object. The association state for such time steps is considered empty, i.e., true negative. Similarly, for an object that disappears from the field of view, either by exiting the camera frustum or being fully occluded, the empty association state is used as the previous association state for the first frame in which the object reappears.
[0456] From Figure 25 it can be seen that the most important confounding factor is the "from" category. This means that the strongest predictor of the association state is the association state in the previous time step. This relationship is intuitive; if the detector fails to detect an object in one time step, it is expected to detect the object in multiple frames. The object may be inherently difficult to recognize by the detector, or some characteristics of the scene (e.g., lens flare of the camera) may affect its perception ability and persist over multiple frames. The next most important confounding factor is occluded_category. This is again intuitive - if an object is occluded, it is more difficult to detect and thus more likely to be a false negative. Distance is also important. Again, this is expected; the farther an object is, the less information there is about it (e.g., in a camera image, a farther car is represented by fewer pixels compared to a closer car).
[0457]
[0458] Under the guidance of this evaluation, a false negative model is constructed, where the only confounding variable is the association state at the previous time step. This is a Markov model because it assumes that the current state depends only on the previous state. This is modeled by determining the probability of transitioning from the state at time step t-1 to the state at time step t. Denoting the association state as X, this amounts to finding the conditional probability P(X t |X t-1 ). To determine these transition probabilities, the frequencies of these transitions in the training data are calculated. This is equivalent to Bayesian likelihood maximisation. Table 5 shows the transition probabilities and the number of instances for each type of transition in the data. Each bin in Table 5 has over 800 entries, indicating that the implied transition frequencies are reliable. The observed transition probability from false negative to false negative is 0.98, and from associated to associated is 0.96. As expected from the results of the Relief analysis, these values show a strong temporal correlation. Is there a reason attributable to the transition to the false negative state and the persistence of the false negative state? There is a 0.65 probability of transitioning from the empty (true negative) state to the false negative state, and a 0.35 probability of transitioning from the empty (true negative) state to the associated state. This means that when an object first becomes visible, it is more likely to be a false negative. Many objects enter the scene from a distance, which is likely to be an important factor in generating these initial false negatives. Some objects enter the scene from the side, particularly in the roundabout scenario considered in this example. Such objects are truncated in the first few frames, which may be a factor in the early false negatives. To explore these points in more detail, a model that depends on other factors is constructed.
[0459] Table 5 - Probabilities (left two columns) of transitioning from the association state in the previous time step (row) to the association state in the current time step (column), and counts of the number of transitions in the training dataset (right two columns)
[0460]
[0461] As a first step towards a more complex model, a Relief analysis is performed to identify confounding factors that are important for the transitions, regardless of the previous association state. MultiSURF is used on a randomly selected 20% of the training data samples. The results are as Figure 26 shown.
[0462] Figure 26The most important confounding factors that may affect the transition probability are: occluded_category, distance, and azimuth. In fact, using the criteria listed in Section 3.2, all confounding factors are good confounding factors. Based on this evidence, the next most complex Markov model is created; occluded_category is added as a confounding factor. Representing the association state X and the occluded category C, the conditional probability P(X t |X t-1 , C t ). As with the first Markov model, these transition probabilities are determined from the training data by counting the occurrence frequencies. Table 6 shows the transition probabilities and the number of instances for each transition in the data.
[0463] Table 6 - Probabilities (left two columns) of transitioning from the association state and occluded category in the previous time step (row) to the association state in the current time step (column), and counts of the number of transitions in the training dataset (right two columns).
[0464]
[0465] Table 6 shows that some of the frequencies are determined from very low counts. For example, there are only 27 transitions from a fully occluded false negative to the association state. However, this event is expected to be rare, and any of these transitions may indicate incorrect training data. These counts may come from an incorrect association of the annotated data with the detector's observations; if the object is truly fully occluded, then the detector would not be expected to observe it. Perhaps the least trustworthy transitions come from the association and full occlusion; there are only 111 observations in total for this category. The probability of transitioning from the association and full occlusion to the association is 0.61, which is highly likely; while the count of transitions from the false negative and full occlusion to the association is low, it actually has a zero probability (because the count of transitions from the false negative and full occlusion to the false negative is so high). Rows with low sums should be treated with caution.
[0466] Despite these limitations, there are still expected trends. When objects transition from the empty state (i.e., when they are first observed), if they are fully visible, there is a 0.60 chance of transitioning to the association, meaning that the object is more likely to be transitioned to the association than a false negative. However, if the object is mostly occluded, the probability of transitioning to the association is only 0.17.
[0467] Given the limitations of identification, it is possible to determine whether adding confounding factors improves the model. To compare these models, the approach of using the model with the smaller negative log predictive density (NLPD) to better explain the data is adopted. The corresponding NLPD is calculated on the held-out test set. The NLPD of the simple Markov model is 10,197 compared to 9,189 for the Markov model with confounding factors. Adding the occlusion_category confounding factor improves the model by this metric.
[0468] This comparison shows that including confounding factors can improve the model. To build a model that includes all relevant confounding factors, following the paradigm used in piecewise constant models, new confounding factors add additional bins (e.g., Table 6 has more rows than Table 5).
[0469] 5. Neural Network PRISM
[0470] This section describes how to implement PRISMS using a neural network or similar "black box" model.
[0471] As is well known in the art, a neural network is formed by a series of "layers", which in turn are formed by neurons (nodes). In a classical neural network, each node in the input layer receives a component of the network input (e.g., an image), which is typically multi-dimensional, and each node in each subsequent layer is connected to each node in the previous layer and computes a function of the weighted sum of the outputs of the nodes to which it is connected.
[0472] For example, Figure 27 illustrates node i in a neural network, where node i receives a set of inputs {u j} and computes a function of the weighted sum of those inputs as its output:
[0473]
[0474] Here, g is a possibly non-linear "activation function", and {w i,j} is the set of weights applied to node i. The weights of the entire network are adjusted during training.
[0475] Reference Figure 28, it is useful to conceptualize the input and output of the layers of a CNN as "volumes" in a discrete three-dimensional space (i.e., a three-dimensional array), each discrete three-dimensional space being formed by a stack of two-dimensional arrays referred to herein as "feature maps". More generally, a CNN takes "tensors" as input, which can typically have any number of dimensions. The following description may also refer to a layer of feature maps as a layer of tensors.
[0476] For example, Figure 28 shows a sequence of five such tensors 302, 304, 306, 308, and 310, which can be generated, for example, by a series of convolution operations, pooling operations, and non-linear transformations known in the art. For reference, the two feature maps within the first tensor 302 are labeled 302a and 302b, respectively, and the two feature maps within the fifth tensor 310 are labeled 310a and 310b, respectively. In this document, (x, y) coordinates refer to the applicable positions within a feature map or an image. The z-dimension corresponds to the "depth" of the feature map or the image, and can be referred to as the feature dimension. A color image has three depths corresponding to three color channels, i.e., the value at (x, y, z) is the value of color channel z at position (x, y). The tensors generated at a processing layer within a CNN have a depth corresponding to the number of filters applied at that layer, where each filter corresponds to a specific feature that the CNN learns to recognize.
[0477] A CNN differs from classical neural network architectures in that a CNN has processing layers that are not fully connected. Instead, processing layers are provided that are only partially connected to other processing layers. In particular, each node in a convolutional layer is only connected to a local 3D region of the processing layer, receives input from that local 3D region, and the nodes over that local 3D region perform a convolution with respect to a filter. The nodes to which a particular node is connected are said to be within the "receptive field" of that filter. A filter is defined by a set of filter weights, and the convolution at each node is a weighted sum (weighted according to the filter weights) of the outputs of the nodes within the filter's receptive field. The local partial connection from one layer to the (x, y) positions of the values within its corresponding tensor is such that the (x, y) position information is at least to some extent preserved in the CNN as the data passes through the network.
[0478] Each feature map is determined by convolving a given filter over an input tensor. Thus, the depth (range in the z - direction) of each convolutional layer is equal to the number of filters applied in that layer. The input tensor itself can be an image or a stack of feature maps, where the feature maps themselves are determined by convolution. When convolution is applied directly to an image, each filter acts as a low - level structure detector because “activation” (i.e., a relatively large output value) occurs when the pixels within the receptive field of the filter form some structure (i.e., match the structure of a particular filter). However, when convolution is applied to a tensor that is itself the result of an earlier convolution in the network, each convolution is performed on a set of feature maps for different features, and thus the network is further activated when a particular combination of lower - level features is present within the receptive field. Thus, for each successive convolution, the network detects the presence of increasingly higher - level structural features corresponding to particular combinations of features from the previous convolution. Thus, in the early layers, the network effectively performs lower - level structure detection, but gradually moves towards more high - level structural semantic understanding in the later layers. The filter weights are learned during training, and this is how the network learns what structures to look for. As is known in the art, convolution can be used in combination with other operations. For example, pooling (a form of dimensionality reduction) and non - linear transformations (such as ReLu, softmax, etc.) are typical operations used in combination with convolution in a CNN.
[0479] Figure 29 A highly schematic overview of the PSPM implemented as a neural network (net) or similar trainable function approximator is shown.
[0480] In this example, neural network A00 has an input layer A02 and an output layer A04. Although neural network A00 is schematically depicted as a simple feed - forward neural network, this is merely illustrative, and neural network A100 can take any form, including for example a Recurrent Neural Network (RNN) and / or a Convolutional Neural Network (CNN) architecture. The terms “input layer” and “output layer” do not imply any particular neural network architecture and include, for example, the input and output tensors in the case of a CNN.
[0481] At input layer A02, neural network A00 receives the sensed ground truth t as input. For example, the sensed ground truth t can be encoded as an input vector or tensor. In general, the sensed ground truth t can be related to any number of objects and any number of underlying sensor modalities.
[0482] Neural network A00 can be mathematically represented as a function as follows:
[0483] y = f(t; w)
[0484] where w is a set of adjustable weights (parameters) that processes the input t according to the set of adjustable weights (parameters). During training, the aim is to optimize the weights w with respect to some loss function defined on the output y.
[0485] In Figure 29 the example of, the output y is a set of distribution parameters that defines the predicted probability distribution p(e|t), where the predicted probability distribution p(e|t) is the probability of obtaining some predicted perceptual output e given the perceptual ground truth t at the input layer A02.
[0486] Taking the simple example of a Gaussian (normal) distribution, the output layer A04 can be configured to provide the predicted mean and variance for a given ground truth:
[0487] y = {μ(t; w), σ(t; w)}.
[0488] Note that either the mean or the variance can vary as a function of the input ground truth t as defined by the learned weights w, giving the neural network A00 the flexibility to learn such dependencies during training to the extent that they are reflected in the training data it is exposed to.
[0489] During training, the aim is to learn the weights w that match p(e|t) to the actual perceptual output A06 generated by the perceptual slice 204 to be modeled. This means that, for example, by optimizing a suitable loss function A08 via gradient descent or ascent, the distribution p(e|t) predicted at the output layer for a given ground truth t can be meaningfully compared to the actual perceptual output corresponding to the ground truth t. As described above, the ground truth input t for training is provided by the ground truth (annotation) pipeline 802, which has been defined by manual, automatic, or semi-automatic annotation of the sensor data to which the perceptual slice 204 has been applied. The set of sensor data to which the perceptual slice 204 is applied can be referred to in the following description as input samples or equivalently as frames and is denoted by the reference numeral A01. The actual perceptual output for each frame A01 is calculated by applying the perceptual slice 204 to the sensor data of that frame. However, according to the teachings above, the neural network A00 is not exposed to the underlying sensor data during training but receives the annotated ground truth t of the frame A01 as the input conveying the underlying scene.
[0490] Given a sufficient set of example {e, t} pairs, various existing neural network architectures can be trained to predict the conditional distribution of the form p(e|t). For simple Gaussian distributions (univariate or multivariate), a log-normal or (negative) log-PDF loss function A08 can be used. One way to extend this to non-Gaussian distributions is to use a Gaussian mixture model, where a neural network A00 and the mixture coefficients for combining these (learned according to the input t in the same way as the means and variances of each Gaussian component) are trained together to predict a multi-component Gaussian distribution. Theoretically, any distribution can be represented as mixed Gaussians, so the Gaussian mixture model is a useful way to approximate general distributions. References in this paper to "fitting Normal distributions", etc. include Gaussian mixture models. The relevant descriptions also apply more generally to other distribution parameterizations. As will be understood, there are various known techniques by which neural networks can be constructed and trained to predict conditional probability distributions given a sufficient representative set of input-output pairs. Therefore, further details are not described in this paper unless specifically relevant to the embodiments described.
[0491] At inference time, the trained network A00 is used as described above. The sensed ground truth t provided by the simulator 814 is provided to the neural network A00 of the input layer A02, which is processed by the neural network A00 to generate a predicted sensed output distribution of the form p(e|t) at the output layer A04, which can then be sampled by a sampling orchestration component (sampler) 816 in the manner described above.
[0492] It is important to note the terminology used in this paper. In this context, "ground truth" refers to the input of the neural network A00 from which its output is generated. In training, the ground truth input comes from annotations, and at inference time the ground truth is provided by the simulator 814.
[0493] Although the actual sensed output A06 can be considered as an example of the form of ground truth in the training context - since it is an example of the type of output that the neural network is trained to replicate - this term is generally avoided in this paper to avoid confusion with the input to the PSPM. The sensed output generated by applying the sensed slice 204 to the sensor data is referred to as the "actual" or "target" sensed output. The aim of training is to adjust the weights w via a suitable loss function that optimizes the deviation between the measured network output and the target sensed output, so that the distribution parameters of the output layer A04 match the actual sensed output A06.
[0494] Figure 29Not necessarily a complete representation of the input or output of neural network A00—it can take additional inputs on which the predictive distribution will depend and / or it can provide the output of other functions as its input.
[0495] 5.1 Confounding factors
[0496] Figure 30 Shows an extension of the neural network to incorporate one or more confounding factors c according to the above principle. Confounding factors are easily incorporated into the architecture as they can simply be provided as additional inputs to the input layer A02 (during both training and inference), and thus, during training, neural network A00 can learn the dependence of the output distribution on the confounding factors c. That is, network A00 can learn the distribution p(e|t, c) at the output layer A04, where any parameters of the distribution (e.g., mean, standard deviation, and mixing coefficients) can depend not only on the ground truth t but also on the confounding factors c, to the extent that those dependencies are captured in the training data.
[0497] 5.2 Temporal dependence
[0498] Figure 31 Shows another extension to incorporate explicit time dependency. In this case, the function (neural network) takes as input at the input layer A02:
[0499] ● The current ground truth t t
[0500] · The ground truth t at the previous time step t-1
[0501] · The previous detection output e t-1
[0502] where the subscript t (non-bold, italic) denotes the time instant. The output is the distribution of the current sensed output p(e t |t t , t t-1 , e t-1 ), where the current sampled sensed output e t is obtained by sampling from this distribution.
[0503] Here, e t-1 is also obtained by sampling from the distribution predicted at the previous time step, and thus, the distribution predicted at the current step will depend on the output of the sampler 816 at the previous step.
[0504] Figure 31 And Figure 32 's implementation can be combined to incorporate confounding factors and explicit temporal dependence.
[0505] One way to achieve the above is to model the characteristics of each detected object for the perceived ground truth t and the sampled perception output e respectively. For example, these characteristics may include Position, Extent, Orientation, and Type.
[0506] The output layer A04 of the neural network is used to predict a transformed real valued variable, and then parameterize the probability distribution of the variable of interest. Conceptually, this form of neural network models the perception slice 204 as a stochastic function.
[0507] Epistemic uncertainty motivates the stochastic modeling of the perception slice 204: even if the perception slice 204 is deterministic, it exhibits significant randomness, which stems from the lack of knowledge of many unknown variables that will affect its output in practice.
[0508] Typical scenarios may include multiple perceived objects. Note that e and t in this article are general symbols that can represent the sampled perception output / perceived ground truth of a single object or multiple objects.
[0509] Another challenge mentioned above is to model false positives (FPs, i.e., false positive detections of objects) and false negatives (FNs, i.e., failures to detect objects). The impact of FPs and / or FNs is that the number of ground truth objects (i.e., the number of objects providing the perceived ground truth) does not necessarily match the number of predicted objects (i.e., the number of objects for which true perception output samples are provided).
[0510] A distinction can be made between the "single object" method and the "set-to-set method". In the broadest sense, the single object PSPM relies on a clear one-to-one association between the ground truth object and the predicted object. In terms of the single object ground truth to which it is associated, the simplest way to implement the single object PSPM is to consider each object independently during training. At inference time, the PSPM receives the perceived ground truth for a single object and provides a single object perception output. False negatives can be directly accommodated by introducing some mechanisms that can model the failure detection of a single object.
[0511] 5.3 Single-object PSPMs
[0512] An example implementation of the single object PSPM using a neural network will now be described. A normal distribution is fitted to the position and extent variables (these can be multivariate normal if required).
[0513] For the modeling of orientation, the method of Section 3.2.2 of “Probabilistic Regression of Rotations using Quaternion Averaging and a Deep Multi-Headed Network” by Peretroukhin et al. [https: / / arxiv.org / pdf / 1904.03182.pdf] can be followed — the entire content of which is incorporated herein by reference. In this method, a quaternion representation of the orientation is used. Noise is injected into the tangent space around the quaternion and may be mixed with the surrounding quaternions where the noise is injected.
[0514] Model the fake “negativeness” using a Bernoulli random variable.
[0515] Given the final network layer A04, the individual variable distributions may be conditionally independent, but dependencies / correlations can be induced by feeding noise as an additional input into the neural network to form a stochastic likelihood function. This is actually a mixture distribution.
[0516] The neural network A00 is trained using stochastic gradient descent (maximum likelihood — using the negative log pdf of the random variable as the location and scale variables (which can be multivariate normal if appropriate)).
[0517] The single object method requires a clear association to be established between the ground truth object and the actual perception output A06. This is because the predicted distribution for a given object needs to match the appropriate single object perception output actually produced by the perception slice 204. Identifying and encoding these associations for the purpose of PSPM training can be implemented as an additional step within the ground truth pipeline 802.
[0518] 5.4 Set-to-set method
[0519] In the broadest sense, the set-to-set method is a method that does not rely on a clear association between the ground truth object and the predicted object, i.e., during training, the PSPM does not need to be told which ground truth object corresponds to which predicted object.
[0520] Figure 32Shows the set-to-set PSPM D00, which takes as input the perceived ground truth {t0, t1} for a set of ground truth objects of any size (two ground truth objects in this example, indexed 0 and 1), and provides a realistic perceived output or distribution {e0, e1} for the set of predicted perceived objects (also two in this example - but note the discussion of FPs and FNs below).
[0521] The set-to-set approach has various benefits.
[0522] The primary benefit is a reduced annotation burden - there is no need to determine the association between the ground truth and the actual perceived targets for training purposes.
[0523] Another benefit is that it can model the correlations between objects. In the Figure 32 example, the set-to-set neural network PRISM is shown. The set-to-set neural network PRISM takes as input the perceived ground truth for any number of input objects at its input layer and outputs the distribution for each predicted object. Notably, the architecture of the network is such that the predicted perceived output distribution p(e m |t0, t1) for any given predicted object m can generally depend on the perceived ground truth outputs for all ground truth objects (t0, t1 in this example). More precisely, the architecture is flexible enough to be able to learn these dependencies to the extent that they are reflected in the training data. The set-to-set approach can also learn the extent to which the perceived slice 204 provides overlapping bounding boxes, and any tendency for it to "swap" objects, which are further examples of learnable object correlations.
[0524] More generally, the result of the set-to-set approach is to consider the joint distribution of all detections at once, i.e., p(e1, e2,...|t0, t1) (which reduces to just the product of each p(e m |t0, t1) mentioned in the previous paragraph when the e m are independent of each other). The advantage of doing this is that it can model the correlations between detections. For example, e1 can have instance identifier 0, and e2 can too, but not simultaneously. While the previous paragraphs and Figure 32 assume that the e m are independent, this is not generally required - the output layer can alternatively be configured to represent the joint distribution p(e1, e2|t0, t1) more generally.
[0525] Another benefit is the ability to model false positives using certain set-to-set architectures. This is because the number of ground truth objects does not necessarily have to match the number of predicted perceived objects - the set-to-set architecture is viable where the latter is less than, equal to, or greater than the former, depending on the input to the network.
[0526] 5.5 CNN Set-to-Set Architecture
[0527] For example, the set-to-set CNN architecture will now be described with reference to Figures 33A - 33D The following is assumed: The actual perception output provided by the perception slice 204 includes 3D bounding boxes for any detected objects, with defined position, orientation, and extent (size / dimensions). The CNN uses an input tensor and produces an output tensor, constructed as described below.
[0528] The CNN PSPM D00 jointly models the output detections based on all ground truth detections in a particular frame. An RNN architecture can be used to induce temporal dependencies, which is a way to implement explicit temporal dependencies on previous frames.
[0529] The ground truth t and the output predictions are spatially encoded in the "PIXOR" format. Briefly, the PIXOR format allows for efficient encoding of 3D spatial data based on a top-down (bird's-eye) view. For details, see: Yang et al., "PIXOR: Real-time 3D Object Detection from Point Clouds" [https: / / arxiv.org / abs / 1902.06326], the entire content of which is incorporated herein by reference.
[0530] As Figure 33A shown, to represent the actual perception output A06 for training purposes, a low-resolution (e.g., 800px square) bird's-eye view (classification layer D12, or more generally an object map) of the actual perception 3D bounding boxes is generated. The output objects are drawn in only one "color" - namely, a classification detection image with a binary encoding of "detection-ness" (the "detect" pixels for a given object form the object region). This can be repeated for the ground truth objects, or generalized by coloring the input objects according to their occlusion status (one hot encoding).
[0531] To encode other characteristics of the object in space, more bird's-eye views are generated, which represent the position, extent, orientation, and any other important variables of the vehicles present in each pixel. These further images are referred to as regression layers or perception layers and are denoted by the reference numeral D14. This means that a single detection is represented multiple times in adjacent pixels and that some information is redundant, as shown. The images are stacked to produce a tensor of size (height x, width x, number of important variables (HEIGHT X WIDTH X NUMBER OF IMPORTANT VARIABLES)).
[0532] Note that it is the perception layer that encodes the 3D bounding box and there is redundancy. When "decoding" the output tensor, the actual numerical values of the regression layer D14 define the position, orientation, and extent of the bounding box. The purpose of the spatial encoding in the bird's-eye view is to provide the information encoded within the perception layer D14 of the input tensor in a form that is conducive to CNN interpretation.
[0533] One advantage of this model is that it can learn the correlations between different object detections - for example, it can learn if the stack does not predict overlapping objects. Additionally, the PSPM can learn if the stack is swapping object IDs between objects.
[0534] By feeding in additional input images, the CNN can be encouraged to predict false positives in physically meaningful places, as this provides the CNN with the information needed to determine the correlation between false positives and the map during training. The images can be, for example, a map of the scene (indicating the environmental structure).
[0535] The CNN can also have inputs to receive confounding factors in any suitable form. Object-specific confounding factors can be spatially encoded in the same way. Examples of such confounding factors include occlusion values, i.e., the degree to which an object is occluded by other objects and / or measurements of truncation (the degree to which the object is outside the sensor's field of view).
[0536] Figure 33B The training of the CNN is illustrated schematically.
[0537] The CNN D00 is trained to predict an output tensor D22 from an input tensor D20 using the classification (e.g., cross entropy) loss D32 of the classification layer and the regression (e.g., smooth L1) loss D34 of the regression layer of those tensors. The classification and regression layers of the input and output tensors D20, D22 are shown separately only for clarity. Typically, the information can be encoded in one or more tensors.
[0538] The regression layer of the input tensor D20 encodes the perceived ground truth t of the current frame for any number of ground truth objects.
[0539] The classification loss D32 is defined with respect to the target classification image D24A derived from the actual perceived output e of the current frame. The regression loss is defined with respect to the target perception layer D24B, which spatially encodes the actual perceived output of the current frame.
[0540] Each pixel of the classification layer of the output tensor D22 encodes the probability of detecting an object at that pixel (the probability of "detectability"). The corresponding pixel of the regression layer defines the corresponding object location, extent, and orientation.
[0541] The classification layer of the output tensor D22 is thresholded to produce a binary output classification image. During training, the binary output classification image D23 is used to mask the regression layer, i.e., the regression loss only considers regions of the regression layer where an object exists within the threshold image D23 and ignores regions of the regression layer outside of this.
[0542] Figure 33C Shows how the trained network is applied at test time or during inference.
[0543] During inference, the input tensor D20 now encodes the perceived ground truth t provided by the simulator 814 (for any number of objects).
[0544] The classification layer on the output image is thresholded and used to mask the regression layer of the output tensor D22. Unlabeled pixels within the object regions of the threshold image contain perceived values, which can then be treated as detections.
[0545] The predicted perceived 3D bounding boxes are decoded from the masked regression layers of the output tensor D22.
[0546] Recall that for any given pixel, the numerical value of that pixel in the regression layer defines the extent, location, and orientation of the bounding box, and thus, the predicted 3D bounding box for each unmasked pixel can be obtained directly. As shown, this typically results in a large number of overlapping boxes (proposed boxes) because each pixel within an object region is activated by the binary image (i.e., as a valid bounding box proposal).
[0547] Apply non-maximal suppression (NMS) to the decoded bounding boxes to ensure that objects are not detected multiple times. As is well known, NMS provides a systematic way to discard the proposed boxes based on the confidence of the boxes and their degree of overlap with other boxes. In this case, for the boxes corresponding to any given pixel of the output tensor D22, the detection probability of this pixel from the (non-thresholded) classification layer can be used as the confidence.
[0548] As an alternative, non-maximal suppression can be avoided by selecting an output classification image that only activates the center positions of the objects. Thus, only one detection will be obtained for each object, and NMS is not required. This can be combined with random probabilities (feeding noise as an additional input into the neural network) to mitigate the effect of activating the output classification image only at the center positions of the objects.
[0549] In addition to other losses, a GAN (generative adversarial network) can be used to obtain a more realistic network output.
[0550] The simple example described above does not provide a probability distribution for the output layer - that is, the probability distribution of the output layer is a one-to-one mapping between the perceived ground truth t and the set of predicted perceived outputs e directly encoded in the output tensor (the network is deterministic in this sense). This can be interpreted as the "average" response of the perceived slice 204 given the ground truth t.
[0551] However, as Figure 33D shown, the architecture can be extended to predict the distribution at the output tensor D22, thus applying exactly the same principle as described above with reference to Figure 29 In this case, the perceived values of the output tensor D22 encode the distribution parameters, and the L1 regression loss D34 is replaced with a log PDF loss or other loss suitable for learning conditional distributions.
[0552] Another option is to train an ensemble of deterministic neural networks in the same way but on different subsets of the training data. In the case of M neural networks trained in this way, combining those networks will directly provide sampled perceived outputs (a total of M samples for each ground truth t). Using a sufficient number of appropriately configured deterministic networks, the distribution of their output samples can capture the statistical characteristics of the perceived slice 204 being modeled in a manner similar to a learned parametric distribution.
[0553] 5.6 Modeling Online Error Estimation
[0554] Figure 34Illustrated is the modeling further extended to accommodate online error (e.g., covariance) estimation within the perception slice 816 to be modeled. The online error estimator 816U within the stack provides an error estimate (or a set of error estimates) associated with its perception output. The online error estimation is an estimate within the prediction system 816 regarding the error associated with its output. Note that this is an estimate by the prediction slice itself (potentially defective) of the uncertainty of its output, typically generated in real time using only the information available on the vehicle at runtime. This itself can be in error.
[0555] Such error estimation is important, for example, in the context of filtering or fusion, where multiple perception outputs (e.g., from different sensor modalities) can be fused in a way that respects their relative uncertainty levels. Incorrect online covariance can lead to fusion errors. The online error estimation can also be fed directly into prediction 104 and / or planning 106, for example, where the planning is based on probabilistic prediction. Thus, errors in the online error estimation can have a significant impact on the stack performance and, in the worst case, may lead to unsafe decisions (especially if the error level of a given perception output is underestimated).
[0556] The method of modeling online covariance (or other online error estimation) is different from position, range, and orientation because there is no available ground truth covariance, i.e., the ground truth input t does not include any ground truth covariance.
[0557] Therefore, the only change is to add additional distribution parameters in the output layer A04 to additionally model the distribution p(E|t), i.e., the probability that the online error estimator 816U provides an error estimate of E given the perception ground truth t.
[0558] Note that this also treats the online error estimator 816U as a random function. Without loss of generality, this can be referred to as the neural network A00 learning the "covariance of the covariance". Modeling the online error estimation component 816U in this way can accommodate the epistemic uncertainty regarding the online error estimator 816U in the same way as other such uncertainties regarding the perception system 204. This is particularly useful if the input to the online error estimator 816U is difficult to simulate or costly to simulate. For example, if the online error estimator 816U is directly applied to sensor data, this would be a way to model the online error estimator 816 without having to simulate those sensor data inputs.
[0559] The covariance is fitted by performing a Cholesky decomposition E02 on the covariance matrix to produce a triangular matrix. This results in positive diagonal elements, allowing the logarithm of the diagonal elements to be computed. Then, a normal distribution can be fitted to each component of the matrix (a multivariate normal distribution can be used if desired). At test time, the process is inverted to produce the desired covariance matrix (the lower triangular scale matrices are multiplied together). This allows the loss function to be formulated as a straightforward numerical regression loss on the unconstrained space of the Cholesky decomposition. To "decode" the neural network, an inverse transformation can be applied.
[0560] At inference time, p(E|t) can be sampled in the same way as p(e|t) to obtain a realistic sampled online error estimate.
[0561] Figures 29 - 34 All architectures are depicted in. For example, time and / or confounding factor dependencies can be incorporated into Figure 34 the model such that the covariance of the covariance depends on one or both.
[0562] More generally, the network can be configured to learn a joint distribution of the form P(e, E|t), which reduces to the above when e and E are independent of each other (but both depend on the ground truth t).
[0563] 6. PSPM Applications
[0564] PSPM has many useful applications, some of which will now be described.
[0565] 6.1 Planning under Uncertainty
[0566] The use cases listed above test planning under uncertainty. This means testing how the planner 106 performs in the presence of statistically representative perception errors. In that case, the benefit is being able to expose the planner 106 and the prediction stack 104 to realistic perception errors in a robust and efficient manner.
[0567] One benefit of the confounding factor approach is that when an instance of unsafe behavior occurs in a particular scenario, the contribution of any confounding factor to that behavior can be explored by running the same scenario but using a different confounding factor c (which may have the effect of changing the perception uncertainty p(e|t, c)).
[0568] As previously mentioned, when sampling from the PSPM, it is not necessary to sample in a uniform manner. Biasing the sampling towards outliers (i.e., PSPM samples with lower probabilities) deliberately can be beneficial.
[0569] The manner of incorporating the confounding factor c also facilitates testing more challenging scenarios. For example, if it is observed through simulation that the planner 106 makes relatively more errors in the presence of occlusions, this may be a trigger to test more scenarios where external objects are occluded.
[0570] 6.2 Separating Perception and Planning / Prediction Errors
[0571] Another somewhat related but still independent application is the ability to isolate the reasons for the unsafe decisions of the planner 106 in the runtime stack 100. In particular, it provides a convenient mechanism to infer whether the reason is a perception error rather than a prediction / planning error.
[0572] For example, consider a simulated scenario of an instance where an unsafe behavior occurs. Such an unsafe behavior may be caused by a perception error, but equally it may be caused by a prediction or planning error. To help isolate the reason, the same scenario can be run without the PSPM, i.e., directly on the perfect perception ground truth, to see how the planner 106 performs in the exact same scenario but with perfect perception output. If the unsafe behavior still occurs, this indicates that the unsafe behavior is at least partially attributed to errors outside the perception stack 102, which may indicate prediction and / or planning errors.
[0573] 6.3 Training
[0574] Simulation can also be used as a basis for training (e.g., reinforcement learning training). For example, simulation can be used as a training basis for components within the prediction stack 104, the planner 106, or the controller 108. In some cases, it may be beneficial to run training simulations based on the realistic perception output provided by the PSPM.
[0575] 6.4 Testing Different Sensor Arrangements
[0576] One possible advantage of the PSPM method is the ability to simulate sensor types / locations that have not been actually tested. This can be used to make reasonable inferences about the impact of a group-specific set of sensors on a moving AV or the use of different types of sensors.
[0577] For example, a relatively simple way to test the impact of reducing the pixel resolution of an in-vehicle camera is to reduce the pixel resolution of the annotated images in the annotated ground truth database 804, reconstruct the PSPM, and re-run the appropriate simulations. As another example, the simulations can be re-run with a particular sensor modality (e.g., LiDAR) completely removed to test the possible impact.
[0578] As a more complex example, the impact of changing a particular sensor on the perception uncertainty can be inferred. This is unlikely to be used as a basis for proving safety, but can be used as a useful tool when considering, for example, camera placement.
[0579] 6.6 PSPM for Simulated Sensor Data
[0580] While the above considered the PSPM generated by applying the perception slice 204 to real sensor data, the actual perception output used to train the PSPM can alternatively be derived by applying the perception slice 204 to simulated sensor data in order to model the performance of the perception slice 204 on the simulated sensor data. Note that the trained PSPM does not require simulated sensor data—it still applies to the perception ground truth without the need for simulated sensor input. The simulated sensor data is only used to generate the actual perception output for training. This can be used as a way to model the performance of the perception slice 204 on the simulated data.
[0581] 6.7 Online Application
[0582] Certain PSPMs can also be effectively deployed on an AV at runtime. That is, as part of the runtime stack 100 itself. This can in turn ultimately help the planner 106 take into account the knowledge of the perception uncertainty. The PSPM can be used in conjunction with existing online uncertainty models that form the basis of filtering / fusion.
[0583] Since the PSPM depends on confounding factors, in order to maximize the usefulness of the PSPM at runtime, the relevant confounding factors need to be measured in real time. This may not apply to all types of confounding factors, but when the appropriate confounding factors are measurable, the PSPM can still be effectively deployed.
[0584] For example, the uncertainty estimate made by the PSPM can be used as a prior in combination with an independent measurement of the uncertainty from one of the AV online uncertainty models at runtime. Collectively, these can provide a more reliable indication of the actual perception uncertainty.
[0585] Structural perception refers to a class of data - processing algorithms that can meaningfully interpret the structures captured in perceptual inputs (sensor outputs or perceptual outputs from lower - level perceptual components). This processing can be applied across different forms of perceptual inputs. Perceptual inputs generally refer to any structural representation, i.e., any data set that captures a structure. Structural perception can be applied in two - dimensional (2D) and three - dimensional (3D) spaces. The result of applying a structural - perception algorithm to a given structural input is encoded as a structural - perception output.
[0586] One form of perceptual input is a two - dimensional (2D) image; that is, an image with only color components (one or more color channels). The most basic form of structural perception is image classification, which simply classifies an image as a whole relative to a set of image - category classes. More complex forms of structural perception applied to 2D space include 2D object detection and / or localization (e.g., direction, pose, and / or distance estimation in 2D space), 2D instance segmentation, etc. Other forms of perceptual inputs include three - dimensional (3D) images, i.e., images with at least one depth component (depth channel); for example, 3D point clouds captured using radar or lidar or derived from 3D images; voxel - or mesh - based structural representations, or any other form of 3D structural representation. Examples of perceptual algorithms that can be applied in 3D space include 3D object detection and / or localization (e.g., distance, direction, or pose estimation in 3D space), etc. A single perceptual input can also be formed by multiple images. For example, stereo depth information can be captured in a stereo pair of 2D images, and this image pair can be used as the basis for 3D perception. 3D structural perception can also be applied to a single 2D image, an example being monocular depth extraction, which extracts depth information from a single 2D image (note that a 2D image can capture a certain degree of depth information within one or more of its color channels without any depth channel). This form of structural perception is an example of different "perception modalities" in the terms used herein. Structural perception applied to 2D or 3D images can be referred to as "computer vision".
[0587] Object detection refers to detecting any number of objects captured in a perceptual input and generally involves characterizing each such object as an instance of an object class. Such object detection may involve or be performed in combination with one or more forms of position estimation, e.g., 2D or 3D bounding - box detection (a form of object localization where the aim is to define a region or volume in 2D or 3D space that encloses the object), distance estimation, pose estimation, etc.
[0588] In the case of machine learning (ML), the structure-aware component can include one or more trained perception models. For example, machine vision processing often uses a convolutional neural network (CNN) to implement. Such a network requires a large number of training images, which have been annotated with the information that the neural network needs to learn (in the form of supervised learning). During training, the network is presented with thousands, or preferably hundreds of thousands, of such annotated images, and learns on its own how the features captured in the images themselves are related to the associated annotations. Each image is annotated in the sense of being associated with the annotation data. The image serves as the perception input, and the associated annotation data provides the "ground truth" for the image. A CNN and other forms of perception models can be constructed to receive and process other forms of perception inputs (e.g., point clouds, voxel tensors, etc.) and perceive structures in 2D and 3D spaces. In the case of general training, the perception input can be referred to as a "training example" or a "training input". In contrast, the training examples captured at runtime for processing by the trained structure-aware component can be referred to as "runtime inputs". The annotation data associated with the training input provides the ground truth for the training input, as the annotation data encodes the expected perception output for the training input. During the supervised training process, the parameters of the structure-aware component are systematically adjusted to minimize the overall measure of the difference between the perception output ("actual" perception output) produced by the structure-aware component when applied to the training examples in the training set and the corresponding ground truth (expected perception output) provided by the associated annotation data within a defined range. In this way, the perception input "learns" from the training examples and is furthermore able to "generalize" this learning, in the sense that it can be trained to provide a meaningful perception output for perception inputs that it did not encounter during training.
[0589] Such sensing components are the cornerstone of many mature and emerging technologies. For example, in the field of robotics, mobile robot systems capable of autonomously planning paths in complex environments are becoming increasingly common. An example of such a rapidly emerging technology is an autonomous vehicle (AV) that can navigate itself on urban roads. Such vehicles must not only perform complex maneuvers among people and other vehicles, but they must do so frequently while ensuring that the probability of adverse events occurring, such as collisions with these other entities in the environment, is strictly limited. For an AV to plan safely, it is crucial that it can observe its environment accurately and reliably. This includes the need to accurately and reliably detect real-world structures near the vehicle. An autonomous vehicle, also known as a driverless vehicle, is a vehicle that has a sensor system for monitoring its external environment and a control system that can automatically make and execute driving decisions using these sensors. This particularly includes the ability to automatically adapt the vehicle speed and driving direction based on perceptual inputs from the sensor system. A fully autonomous or "driverless" vehicle has sufficient decision-making capabilities to operate without any input from a human driver. However, the term "autonomous vehicle" as used in this article also applies to semi-autonomous vehicles, which have more limited autonomous decision-making capabilities and thus still require a certain degree of supervision from a human driver. Other mobile robots, for example, that transport goods within and outside industrial areas are being developed. Such mobile robots will have no one on board and belong to a class of mobile robots called UAVs (unmanned autonomous vehicles). Autonomous aerial mobile robots (drones) are also being developed.
[0590] Thus, in the more general context of autonomous driving and robotics, one or more sensing components may be needed to interpret perceptual inputs, i.e., one or more sensing components can determine the information about real-world structures captured in a given perceptual input.
[0591] Increasingly complex robotic systems (e.g., AVs) may need to implement multiple sensing modalities to accurately interpret multiple forms of perceptual input. For example, an AV can be equipped with one or more pairs of stereo optical sensors (cameras) from which relevant depth maps are extracted. In this case, the AV's data processing system can be configured to apply one or more forms of 2D structure sensing to the images themselves—e.g., 2D bounding box detection and / or other forms of 2D localization, instance segmentation, etc.—plus apply one or more forms of 3D structure sensing to the data of the relevant depth maps—e.g., 3D bounding box detection and / or other forms of 3D localization. Such depth maps can also come from LiDAR, RADAR, etc., or be derived by combining multiple sensor modalities.
[0592] The present technology can be used to simulate the behavior of various robotic systems for purposes such as testing / training. The runtime application can also be implemented in different robotic systems.
[0593] To train a perception component for a desired perception modality, the perception component is constructed such that it can receive perception inputs in a desired form and provide perception outputs in a desired form in response. Additionally, to train a perception component with an appropriate architecture based on supervised learning, annotations conforming to the required perception modality need to be provided. For example, to train a 2D bounding box detector, 2D bounding box annotations are required; similarly, to train a segmentation component that performs image segmentation (pixel-by-pixel classification of individual image pixels), the annotations need to encode suitable segmentation masks from which the model can learn; a 3D bounding box detector needs to be able to receive 3D structure data, as well as annotated 3D bounding boxes, etc.
[0594] A perception component can refer to any tangible embodiment (instance) of one or more underlying perception models of the perception component, which can be a software or hardware instance, or a combined software and hardware instance. Such an instance can be embodied using programmable hardware such as a general-purpose processor (e.g., a CPU, an accelerator such as a GPU, etc.) or a field programmable gate array (FPGA) or any other form of programmable computer. Thus, a computer program for programming a computer can take the form of program instructions executed on a general-purpose processor, circuit description code for programming an FPGA, etc. Instances of the perception component can also be implemented using non-programmable hardware, e.g., an application specific integrated circuit (ASIC), and such hardware can be referred to herein as a non-programmable computer. Generally, a perception component can be embodied in one or more computers, which can be programmable or not, and the one or more computers are programmed or otherwise configured to execute the perception component.
[0595] Referring to Figure 8 , the depicted pipeline component is a functional component of a computer system, which can be implemented in various ways at the hardware level: Although Figure 8Although not shown, the computer system includes one or more processors (computers), and the one or more processors perform the functions of the above components. The processor may be in the form of a general-purpose processor (e.g., a CPU (Central Processing Unit) or an accelerator (e.g., a GPU), etc.), or in the form of a more specialized hardware processor (e.g., an FPGA (Field Programmable Gate Array), or an ASIC (Application-Specific Integrated Circuit)). Although not shown separately, the UI generally includes at least one display and at least one user input device for receiving user input to allow the user to interact with the system, such as a mouse / touchpad, a touch screen, a keyboard, etc.
[0596] The various aspects of the present invention and their exemplary embodiments have been described above. Other aspects and exemplary embodiments of the present invention are described below.
[0597] On the other hand, a method for testing the performance of a robotic planner and a perception system is provided, the method comprising:
[0598] receiving at least one probability uncertainty distribution for modeling at least one perception component of the perception system, such as determining the at least one probability uncertainty distribution based on a statistical analysis of the actual perception output derived by applying the at least one perception component to inputs directly or indirectly obtained from one or more sensor components; and
[0599] running a simulation scenario in a simulator, wherein the simulated robot state changes according to autonomous decisions made by the robotic planner based on the realistic perception output calculated for each simulation scenario;
[0600] wherein the realistic perception output models the actual perception output to be provided by at least one perception component in the simulation scenario, but is calculated without applying the at least one perception component to the simulation scenario and without simulating one or more sensor components, but by:
[0601] (i) directly calculating the perception ground truth of at least one perception component based on the simulation scenario and the simulated robot state, and
[0602] (ii) modifying the perception ground truth according to the at least one probability uncertainty distribution to calculate the realistic perception output.
[0603] Note that the terms "perception pipeline", "perception stack", and "perception system" are used synonymously herein. The term "perception slice" is used to refer to all or part of a perception stack (including one or more perception components) modeled by a single PSPM. As described later, during simulation safety testing, the perception stack may be replaced in whole or in part by one or more PSPMs. The term slice may also be used to refer to the part of the prediction stack that is not modeled by a PSPM or is replaced by a PSPM, and the meaning will be clear from the context.
[0604] In a preferred embodiment of the present invention, the realistic perception output depends not only on the perception ground truth but also on one or more "confounders". That is, the influence of the confounders on the perception output modeled by the PSPM. Confounders represent real-world conditions that may affect the accuracy of the perception output (e.g., weather, lighting, speed of another vehicle, distance to another vehicle, etc.; examples of other types of confounders will be given later). It is said that the PSPM is mapped to a "confounder space" that represents all possible confounders or combinations of confounders that the PSPM can consider. This allows the PSPM to accurately model different real-world conditions represented by different points in the confounder space in an efficient manner, since the PSPM eliminates the need to simulate sensor data for those different conditions and does not require the application of the perception components themselves as part of the simulation.
[0605] The term "confounder" is sometimes used in statistics to refer to a variable that has a causal effect on both the dependent variable and the independent variable. However, in this document, the term is used in a more general sense to represent a variable of a perception error model (PSPM) that represents some physical condition.
[0606] In an embodiment, at least one probability uncertainty distribution can be used to model multiple collaborative perception components of the perception system.
[0607] In an embodiment, only a part of the perception system can be modeled, and at least a second perception component of the perception system can be applied to the realistic perception output to provide a second perception output for making the decision.
[0608] The second perception component can be a fusion component, such as a Bayesian filter or a non-Bayesian filter.
[0609] The modelled perception component can be a sensor data processing component that is highly sensitive to artefacts in the simulated data. In this case, the above method avoids the need to simulate high-quality sensor data for this component. For example, the perception component can be a Convolutional Neural Network (CNN) or other form of neural network.
[0610] Alternatively or additionally, the modelled perception component can be a sensor data processing component that processes sensor data that is inherently difficult to simulate. For example, a RADAR processing component.
[0611] The method can include the step of analyzing changes in the simulated robot state to detect instances of unsafe behavior of the simulated robot state and determining the cause of the unsafe behavior.
[0612] Instances of unsafe behavior can be detected based on a predefined set of acceptable behavior rules applied to the simulated scenario and the simulated robot state.
[0613] The rules of such acceptable behavior can take the form of a "Digital Highway Code" (DHC).
[0614] The PSPM combined with the DHC allows many realistic simulations to be efficiently run without knowing which will lead to unsafe / unacceptable behavior (as opposed to the movement variations of scenarios known to be unsafe from real-world test drives), where the predefined rules of the DHC are used to automatically detect instances of such behavior.
[0615] The perception component and / or the planner can be modified to mitigate the cause of the unsafe behavior.
[0616] The sensor output obtained from one or more sensors and the corresponding perceived ground truth associated with the sensor output can be used to determine a probability uncertainty distribution.
[0617] The probability uncertainty distribution can vary with one or more confounding factors, where the set of one or more confounding factors selected for the simulated scenario can be used to modify the perceived ground truth according to the probability uncertainty distribution, where each confounding factor represents a physical property.
[0618] The above one or more confounding factors can include one or more of the following:
[0619] - Occlusion level
[0620] - One or more lighting conditions
[0621] - Indication of the time of day
[0622] - One or more weather conditions
[0623] - Indication of season
[0624] - Physical characteristics of at least one external object
[0625] - Sensor conditions (e.g., object position in the field of view)
[0626] In a time-dependent model, another variable input on which the PSPM depends can be the previous ground truth calculated therefrom and / or at least one previous real-world perception output.
[0627] The simulated scenario can be derived from an observed real scenario.
[0628] The simulated scenario can be a fuzzed scenario determined by fuzzing an observed real-world scenario.
[0629] That is, as a result of the perception error, in addition to causing changes in the inputs that occur in the prediction and planning system, it can also be combined with a method of generating additional test scenarios by making changes (small or large) to the situation of the test scenario (e.g., other vehicles that accelerate or decelerate slightly in the scenario, e.g., slightly changing the initial position and orientation of the ego car and other vehicles, etc.). These two types of changes to the known real scenario will together have a higher chance of hitting dangerous situations, and the system needs to be able to handle them.
[0630] On the other hand, there is provided a computer-implemented method for training a perception statistical performance model (PSPM), wherein the PSPM models the uncertainty in the perception output calculated from perception slices, and the method includes:
[0631] Applying the perception slices to a plurality of training sensor outputs, thereby calculating a training perception output for each sensor output, wherein each training sensor output is associated with a perception ground truth;
[0632] Comparing each perception output with the associated perception ground truth, thereby calculating a set of perception errors Δ;
[0633] Using the set of perception errors Δ to train the PSPM, wherein the trained PSPM provides a probability perception uncertainty distribution of the form p(e|t), where p(e|t) represents the probability that the perception slice calculates a specific perception output e given the perception ground truth t.
[0634] On the other hand, there is provided a perception statistical performance model (PSPM) implemented in a computer system, which is used to model perception slices and is configured to:
[0635] Receive the calculated perception ground truth t;
[0636] Determine a probabilistic perceptual uncertainty distribution of the form p(e|t) from the sensed ground truth t based on a set of learning parameters θ, where p(e|t) represents the probability of the perceptual slice computing a particular perceptual output e given the computed sensed ground truth t, and the probabilistic perceptual uncertainty distribution is defined over a range of possible perceptual outputs, and the parameter θ is learned from a set of actual perceptual outputs generated using the perceptual slices to be modeled.
[0637] In a preferred embodiment, the PSPM can vary according to one or more confounding factors c, where each confounding factor characterizes a physical condition. In this case, the probabilistic perceptual uncertainty distribution takes the form p(e|t, c).
[0638] In an embodiment, the PSPM can take the form of a parametric distribution defined by a set of parameters θ learned from a set of perceptual errors Δ and varying as a function of the given sensed ground truth t.
[0639] To train the PSPM according to the confounding factors, each training perceptual output can also be associated with a set of one or more confounding factors that characterize one or more physical conditions in which the training sensor outputs are captured.
[0640] The ground truth for training the PSPM can be generated offline because more accurate and thus generally more computationally intensive algorithms can be used compared to the online case. These only need to be generated once.
[0641] Note that the term parameter includes hyperparameters, e.g., hyperparameters learned by variational inference.
[0642] In an embodiment, the sensed ground truth associated with the sensor output can be derived from the sensor output using offline processing (e.g., processing that cannot be performed in real time due to hardware limitations or because offline processing is inherently non-real-time).
[0643] Models suitable for the PSPM typically draw attention to confounding factors in the data used for modeling, which may or may not be initially obvious. The advantage of this is that only significant confounding factors need to be modeled separately, and their importance is determined by the degree to which the data deviates from the model.
[0644] The confounding factor c is a variable on which the trained PSPM depends. At runtime, by changing the value of the confounding factor c, realistic perceptual outputs (i.e., with realistic errors) can be obtained for different physical situations. The variable can be numerical (e.g., continuous / pseudo-continuous) or categorical (e.g., binary value or non-binary categorical value).
[0645] It is possible that the training of the PSPM reveals a statistically significant dependence on one or more physical characteristics that are not currently characterized by the existing confounding factor c. For example, when the trained PSPM is validated, its performance on certain types of data may be worse than expected, and the analysis can attribute it to the dependence on physical conditions that are not explicitly modeled in the PSPM.
[0646] Thus, in an embodiment, the method may include the steps of: analyzing the trained PSPM for the confounding factor c (e.g., validating the PSPM using a validation-aware error dataset), and in response thereto, retraining the PSPM for a new set of one or more confounding factors c', whereby the probability-aware uncertainty distribution of the retrained PSPM takes the form p(e|t, c').
[0647] For example, c' can be determined by adding or removing the confounding factor c. For example, if the confounding factor is considered to be statistically significant, it can be added, or if the analysis shows that the confounding factor is not actually statistically significant, it can be removed.
[0648] By modeling the PSPM in this way, it is possible to determine which confounding factors are statistically significant and need to be modeled, and which confounding factors are not statistically significant and do not need to be modeled.
[0649] For example, the above one or more confounding factors may include one or more of the following:
[0650] - The occlusion level of at least one external object (indicating the degree of occlusion of the object relative to the subject. The external object can be a moving actor or a static object)
[0651] - One or more lighting conditions
[0652] - An indication of the time of day
[0653] - One or more weather conditions
[0654] - An indication of the season
[0655] - The physical characteristics of at least one external object (e.g., position / distance from the subject, speed / rate / acceleration relative to the subject, etc.)
[0656] - The position of the external object in the subject's field of view (e.g., in the case of a camera, the angle from the image center)
[0657] Another aspect of the present disclosure provides a computer system for testing and / or training a runtime stack of a robotic system, the computer system comprising:
[0658] A simulator configured to run a simulation scenario, wherein a simulation subject interacts with one or more external objects;
[0659] A runtime stack including an input configured to receive a time series of perceptual outputs of each simulation scenario, a planner configured to make autonomous decisions based on the perceptual outputs, and a controller configured to generate a series of control signals to cause the simulation subject to execute the decisions as the simulation scenario progresses;
[0660] Wherein the computer system is configured to calculate each perceptual output of the time series by:
[0661] Calculating a perceptual ground truth based on the current state of the simulation scenario;
[0662] Applying the above PSPM to the perceptual ground truth to determine a probabilistic perceptual uncertainty distribution; and
[0663] Sampling a perceptual output from the probabilistic perceptual uncertainty distribution.
[0664] Preferably, the PSPM is applied to the perceptual ground truth and one or more sets of confounding factors associated with the simulation scenario.
[0665] Ray tracing can be used to calculate the perceptual ground truth for each external object.
[0666] Each external object can be a moving actor or a static object.
[0667] The same simulation scenario can be run multiple times.
[0668] The same simulation scenario can be run multiple times with different sets of confounding factors.
[0669] The runtime stack may include a prediction stack configured to predict the behavior of external actors based on the perceptual outputs, wherein the controller may be configured to make decisions based on the predicted behavior.
[0670] The computer system may be configured to record details of each simulation scenario in a test database, where the details include the decisions made by the planner, the perceptual outputs on which those decisions are based, and the behavior of the simulation subject when executing those decisions.
[0671] The computer system may include a scenario evaluation component configured to analyze the behavior of the simulation subject in each simulation scenario related to a set of predetermined behavior rules in order to classify the behavior of the subject.
[0672] The analysis results of the scenario evaluation component can be used to formulate simulation strategies. For example, the scenario can be "fuzzified" based on the analysis results (see below).
[0673] The behavior of the subject can be classified as safe or unsafe.
[0674] To model false negative detections, a probability-aware uncertainty distribution can provide the probability of successfully detecting a visible object, which is used to determine whether to provide an object detection output for that object. (In this case, a visible object means an object that is within the subject's sensor field of view in the simulation scenario but that the subject may still fail to detect.)
[0675] A time-dependent PSPM (e.g., a Hidden Markov Model) can be used in any of the above.
[0676] In the case of modeling false negatives, a time-dependent PSPM can be used such that the probability of detecting a visible object depends on at least one earlier determination regarding whether to provide an object detection output for the visible object.
[0677] To model false positive detections, a probability uncertainty distribution can provide the probability of false object detections, which is used to determine whether to provide a perception output for a non-existent object.
[0678] Once the "ground truth" is determined, if the scenario is run in a loop without a PSPM, potential errors in the planner can be explored. This can be extended to automatically classify data to indicate perception problems or planner problems.
[0679] In an embodiment, a simulation scenario in which the simulated subject exhibits unsafe behavior can be re-run without applying a PSPM but by directly providing the perception ground truth to the runtime stack.
[0680] Then, an analysis can be performed to determine whether the simulated subject still exhibits unsafe behavior in the re-run scenario.
[0681] Another aspect of the present invention provides a method for testing a robot planner that is used to make autonomous decisions using the perception output of at least one perception component, the method comprising:
[0682] Running a simulation scenario in a computer system, wherein the simulated robot state changes according to autonomous decisions made by the robot planner using realistic perception outputs calculated for each simulation scenario;
[0683] For each simulation scenario, determining an ontological representation of the simulation scenario and the simulated robot state; and
[0684] Apply a set of predefined acceptable behavior rules [e.g., DHC] to the ontology representation of each simulated scenario in order to record and flag violations of the predefined acceptable behavior rules in one or more simulated scenarios.
[0685] Another aspect of the invention provides a computer-implemented method that includes steps of implementing any of the above programs, systems, or PSPM functions.
[0686] A further aspect provides a computer system that includes one or more computers programmed or otherwise configured to perform any of the functions disclosed herein, and one or more computer programs for programming the computer system to perform the functions.
[0687] It should be understood that the various embodiments of the invention have been described by way of example only. The scope of the invention is not limited by the described embodiments, but is defined only by the appended claims.
Claims
1. A computer system for testing and / or training a runtime stack of a robotic system, the computer system comprising: A simulator configured to run a simulation scenario, wherein a simulated entity interacts with one or more external objects; A planner of the runtime stack, the planner of the runtime stack being configured to make autonomous decisions for each simulation scenario based on a time series of perception outputs calculated for the simulation scenario; and A controller of the runtime stack, the controller of the runtime stack being configured to generate a series of control signals to cause the simulated entity to execute the autonomous decisions as the simulation scenario progresses; wherein the computer system is configured to calculate each perception output by: Calculating a perception ground truth at the current moment based on the current state of the simulation scenario; Applying a perception statistical performance model PSPM to the perception ground truth to determine a probabilistic perception uncertainty distribution at the current moment; and Sampling the perception output from the probabilistic perception uncertainty distribution; wherein the PSPM is used to model a perception slice of the runtime stack and is configured to determine the probabilistic perception uncertainty distribution based on a set of parameters learned from a set of actual perception outputs generated using the perception slice to be modeled; wherein the PSPM includes a time-dependent model such that the perception output sampled at the current moment depends on at least one of: an earlier one of the perception outputs sampled at the previous moment, and an earlier one of the perception ground truths calculated for the previous moment.
2. The computer system according to claim 1, wherein, The time-dependent model is a hidden Markov model.
3. The computer system according to claim 1 or 2, wherein, The modeled perception slice includes at least one filtering component, and the time-dependent model is used to model the time-dependence of the filtering component.
4. The computer system according to claim 1 or 2, wherein, The PSPM is a time-dependent neural network architecture.
5. The computer system according to claim 4, wherein, The PSPM is configured to receive earlier perception outputs and / or earlier perception ground truths as inputs.
6. The computer system according to claim 4, wherein The PSPM has a recurrent neural network architecture.
7. The computer system according to any one of claims 1, 2, 5, and 6, wherein, The PSPM is applied to the perception ground truth and one or more confounding factors associated with the simulation scenario, wherein each confounding factor is a variable of the PSPM, the value of the variable characterizing the physical conditions applicable to the simulation scenario, and the probabilistic perception uncertainty distribution depends on the variable.
8. The computer system according to claim 7, wherein, The one or more confounding factors include one or more of the following confounding factors, the confounding factors at least partially determining the probabilistic uncertainty distribution from which the perception output is sampled: The occlusion level of at least one of the external objects; One or more lighting conditions; An indication of the time of day; One or more weather conditions; An indication of the season; The physical characteristics of at least one of the external objects; Sensor conditions, the position of at least one of the external objects in the sensor field of view of the entity; The number or density of the external objects; The distance between two of the external objects; The truncation level of at least one of the external objects; The type of at least one of the objects, and An indication as to whether at least one of the external objects corresponds to any external objects from an earlier time in the simulated scenario.
9. The computer system according to any one of claims 1, 2, 5, 6, and 8, comprising: A scenario evaluation component configured to evaluate the behavior of the simulated entity in each of the simulated scenarios by applying a set of predetermined rules.
10. The computer system according to claim 9, wherein, At least some of the predetermined rules are related to security, and the scenario evaluation component is configured to evaluate the security of the behavior of the simulated entity in each of the simulated scenarios.
11. The computer system according to claim 10, wherein, The scenario evaluation component is configured to automatically mark instances of unsafe behavior by the simulated entity for further analysis and testing.
12. The computer system according to claim 10 or 11, wherein, The computer system is configured to re-run the simulated scenario in which the simulated entity initially exhibited unsafe behavior, based on a time series of the perceived ground truth determined for re-running the scenario, without applying the PSPM to those perceived ground truths, and thus without perception error; and to evaluate whether the simulated entity still exhibits unsafe behavior in the re-run scenario.
13. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, and 11, wherein, Sampling from the probabilistic perception uncertainty distribution is non-uniform and biased towards lower probability perception outputs.
14. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, and 11, wherein, The planner is configured to make the autonomous decision based on a second time series of perception outputs, where the PSPM is a first PSPM, and the computer system is configured to compute the second time series of perception outputs using a second PSPM that models a second perception slice of the runtime stack, the first PSPM learning from data of a first sensor modality and a time series of the perception slice, and the second PSPM independently learning from data of a second sensor modality and the second time series of the second perception slice.
15. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, and 11, wherein, The planner is configured to make the autonomous decision based on a second time series of perception outputs, the time series and the second time series corresponding to a first sensor modality and a second sensor modality respectively, where the modeled perception slice is configured to process sensor inputs of two sensor modalities and to compute the two time series by applying the PSPM, the PSPM learning from data of the two sensor modalities to model any correlations between the inside of the perception slice.
16. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, and 11, wherein, The computer system is configured to apply at least one unmodeled perception component of the runtime stack to the perception output to compute a processed perception output, and the planner is configured to make the autonomous decision based on the processed perception output.
17. The computer system according to claim 16, wherein, The unmodeled perception component is a filtering component applied to the time series of the perception output, and the processed perception output is a filtered perception output.
18. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, 11, and 17, comprising a scenario fuzzification component configured to generate at least one fuzzified scenario for running in the simulator by fuzzifying at least one existing scenario.
19. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, 11 and 17, wherein, To model false negative detections, the probability-aware uncertainty distribution provides the probability of a visible object in a successfully detected object, which is used to determine whether to provide an object detection output for the object. The object is visible when it is within the sensor field of view of the simulated agent in the simulated scenario, and thus the detection of the visible object is not guaranteed.
20. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, and 11, wherein, To model false positive detections, the probability uncertainty distribution provides the probability of a false object detection, which is used to determine whether to provide a perception output for a non-existent object.
21. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, 11 and 17, wherein, Ray tracing is used to compute the perception ground truth for the one or more external objects.
22. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, 11, and 17, wherein, At least one of the external objects is a moving actor, and the computer system includes a prediction stack of the runtime stack, the prediction stack of the runtime stack being configured to predict the behavior of the external actor based on the perception output, and the planner being configured to make autonomous decisions based on the predicted behavior.
23. The computer system according to any one of claims 1, 2, 5, 6, 8, 10, 11, and 17, wherein the computer system is configured to record details of each simulation scenario in a test database, where The details include the decisions made by the planner, the perception outputs on which those decisions are based, and the behavior of the simulated agent when executing those decisions.
24. A computer-implemented method for performance testing a runtime stack of a robotic system, the method comprising: Running a simulated scenario in a simulator, wherein a simulated agent interacts with one or more external objects, wherein a planner of the runtime stack makes autonomous decisions for the simulated scenario based on a time series of perception outputs computed for the simulated scenario, and a controller of the runtime stack generates a series of control signals to cause the simulated agent to execute the autonomous decisions as the simulated scenario progresses; wherein each perception output is computed by: Computing a perception ground truth at the current moment based on the current state of the simulated scenario; Applying a perception statistical performance model PSPM to the perception ground truth to determine the probability-aware uncertainty distribution at the current moment; and Sampling the perception output from the probability-aware uncertainty distribution; wherein the PSPM is used to model the perception slice of the runtime stack and determines the probability-aware uncertainty distribution based on a set of parameters learned from a set of actual perception outputs generated using the perception slice to be modeled; wherein the PSPM includes a time-dependent model such that the perception output sampled at the current moment depends on at least one of the following: the earlier one of the perception outputs sampled at the previous moment, and the earlier one of the perception ground truths computed for the previous moment.
25. A computer system comprising a perception statistical performance model PSPM for modeling a perception slice of a runtime stack of a robotic system and configured to: Receive the current moment of the calculated perceived ground truth ; Determine a probabilistic perceptual uncertainty distribution of the form p(e |t t |t t ,t t-1 ) or p(e t |t t ,e t-1 ), from a set of learned parameters θ, where p(e t |t t ,t t-1 ) means that at a given current moment and the previous moment t The calculated perceptual ground truth t t ,t t-1 In the case of perceptual slices, a specific perceptual output e is calculated. t The probability of t |t t ,e t-1 ) represents the calculated perceptual ground truth at a given current time t t and the previous moment t The sampled perceptual output e of -1 t-1 In the case of perceptual slices, a specific perceptual output e is calculated. t wherein the probabilistic perceptual uncertainty distribution is defined over a range of possible perceptual outputs and the parameters θ are learned from a set of actual perceptual outputs generated using the perceptual slices to be modeled.
26. A computer program product for programming one or more computers to implement the computer system of any one of claims 1-23, or to implement the computer-implemented method of claim 24, or to implement the computer system of claim 25.
Citation Information
Patent Citations
Autonomous vehicle manoeuvres
GB201816852D0