System and procedure for evaluating the performance of an autonomous vehicle based on human logic

A cognitive system for autonomous vehicles simulates human driving behavior by training on datasets and adjusting performance metrics, addressing the mismatch between system-defined and human behaviors to improve passenger satisfaction.

DE102021114595B4Active Publication Date: 2025-12-31GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102021114595
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-03
Filing Date
2021-06-08
Publication Date
2025-12-31
Estimated Expiration
2041-06-08

AI Technical Summary

Technical Problem

Autonomous vehicles often exhibit behaviors that differ from those of human drivers, leading to passenger dissatisfaction and unfamiliarity with vehicle responses in various traffic situations.

Method used

A cognitive system is developed to simulate human driving behavior by training on a dataset, evaluating planned actions, and adjusting performance ratings to align with human-based ratings, using metrics such as safe following distance, lane change distance, collision state, and average traffic speed.

Benefits of technology

The cognitive system effectively mimics human driving behavior, enhancing passenger satisfaction by aligning vehicle responses with human-like performance in diverse traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for operating an autonomous vehicle (10), comprising: Operating a cognitive system (304) in response to a training data set (302) to generate a planned action for operating the autonomous vehicle (10); Evaluating the planned action to obtain a system performance rating; Updating the cognitive system (304) based on a comparison of the system performance rating with a human-based performance rating; and Operating the autonomous vehicle (10) using the cognitive system (304); wherein the human-based performance rating is obtained by evaluating a path driven by a human in relation to the training data set (302), wherein the evaluation is carried out by one or more humans; and / or the procedure further involves updating the cognitive system (304) by reducing a difference between the system performance rating and the human-based performance rating; and / or wherein an evaluation model generates the system performance rating, wherein the evaluation model includes at least one base metric weighted with a coefficient, and wherein the method further comprises adjusting the coefficient of the at least one base metric based on the comparison; and / or wherein the procedure further includes determining a complexity score which indicates a difficulty level of a driving scenario for the autonomous vehicle (10) and evaluating the planned action using the system performance rating, the human-based performance rating and the complexity score.
Need to check novelty before this filing date? Find Prior Art

Description

Introduction

[0001] The present invention relates to a system and a method for operating an autonomous vehicle and in particular to a system and a method for operating the autonomous vehicle in order to simulate the behavior of a human operator of the vehicle.

[0002] For background information, please refer to DE 10 2020 128 156 A1.

[0003] An autonomous vehicle operates by detecting objects in its environment or environmental conditions and taking action in response. Generally, an autonomous vehicle operates with a set of instructions that allow it to react to traffic conditions according to a system-defined behavior. However, this system-defined behavior does not always correspond to the behavior a real person driving the vehicle would exhibit. It is desirable, however, for a vehicle's passenger to be familiar with and satisfied with how the vehicle behaves in various traffic situations. Accordingly, it is desirable to provide a system and procedure for operating an autonomous vehicle that mimics or simulates the behavior of a human driver. Summary

[0004] According to the invention, a method for operating an autonomous vehicle is presented, characterized by the features of claim 1.

[0005] A cognitive system operates in response to a training dataset to generate a planned action for operating the autonomous vehicle. This planned action is evaluated to obtain a system performance rating. The cognitive system is updated based on a comparison of this system performance rating with a human-based performance rating. The autonomous vehicle is then operated using this cognitive system.

[0006] In addition to one or more of the features described here, the human-based performance rating is obtained by evaluating a path driven by humans in relation to the training dataset. The human-based performance rating is obtained by evaluating the planned action of one or more humans. The method includes updating the cognitive system by reducing any difference between the system performance rating and the human-based performance rating. In an embodiment in which a rating model generates the system performance rating, wherein the rating model includes at least one base metric weighted with a coefficient, the method further includes adjusting the coefficient of the at least one base metric based on the comparison.At least one basic metric relates to a deviation from a safe following distance, a deviation from a safe lane-change distance, a collision state, and / or a deviation from the average traffic speed. The procedure further includes determining a complexity score, which indicates the difficulty level of a driving scenario for the autonomous vehicle, and evaluating the planned action using the system performance rating, the human-based performance rating, and the complexity score.

[0007] Furthermore, according to the invention, a system for operating an autonomous vehicle is presented, which is characterized by the features of claim 2.

[0008] The system comprises a control system and a cognitive system. The control system executes a driving action on the autonomous vehicle. The cognitive system generates the driving action using an evaluation model. The evaluation model is generated by operating the cognitive system in response to a training dataset to generate a planned action for operating the autonomous vehicle. The planned action is evaluated to obtain a system performance rating, and the cognitive system is updated based on a comparison of the system performance rating with a human-based performance rating.

[0009] In addition to one or more of the features described here, the human-based performance rating is based on a path driven by a human, which relates to the training dataset. The human-based performance rating is based on an evaluation of the planned action by one or more humans. The system also includes a comparison module to update the cognitive system by reducing any difference between the system performance rating and the human-based performance rating. The comparison module evaluates the planned action using the system performance rating, the human-based performance rating, and the complexity score. In an embodiment where the rating model generates the system performance rating and includes at least one coefficient-weighted base metric, the system further includes a comparison model for setting or adjusting the system performance rating.Adjusting the coefficient of at least one basic metric based on the comparison. The at least one basic metric refers to a deviation from the safe following distance, a deviation from a safe lane change distance, a collision state, and / or a deviation from the average traffic speed.

[0010] Furthermore, an autonomous vehicle is described. The autonomous vehicle comprises a cognitive system for generating a driving behavior using an evaluation model. The evaluation model is generated by operating the cognitive system in response to a training dataset to generate a planned action for operating the autonomous vehicle using the cognitive system. The planned action is evaluated to obtain a system performance rating, and the cognitive system is updated based on a comparison of the system performance rating with a human-based performance rating.

[0011] In addition to one or more of the features described herein, the human-based performance rating is based on a path driven by a human with respect to the training dataset and / or an evaluation of the planned action by one or more humans. The vehicle further includes a comparison module to update the cognitive system by reducing any difference between the system performance rating and the human-based performance rating. The comparison module evaluates the planned action using the system performance rating, the human-based performance rating, and the complexity score. In an embodiment that includes at least one coefficient-weighted base metric, the vehicle further includes a comparison module to adjust the coefficient of the at least one base metric based on the comparison.At least one basic metric refers to a deviation from the safe following distance, a deviation from a safe lane change distance, a collision condition and / or a deviation from the average traffic speed.

[0012] The above-mentioned features and advantages, as well as other features and advantages of the invention, are readily apparent from the following detailed description when taken in conjunction with the accompanying drawings. Brief description of the drawings

[0013] Further features, advantages and details are listed in the following detailed description only as examples, with the detailed description referring to the drawings in which: Fig. 1 shows an autonomous vehicle with an associated trajectory planning system according to various embodiments; Fig.2 shows an illustrative control system that includes a cognitive processor integrated into an autonomous vehicle; Fig. 3 shows a schematic diagram illustrating a procedure for training the cognitive system for the operation of an autonomous vehicle to simulate a human driver; Fig. 4 shows a schematic diagram illustrating another method for training the cognitive system to simulate human driving behavior; Fig. Figure 5 shows a schematic diagram detailing the procedure of Fig. 4 illustrates how to train the cognitive system; Fig. 6 shows a first street scenario for evaluating the performance of the cognitive system; Fig.Figure 7 shows a second road scenario where a construction site creates an obstacle that requires the vehicle to move at least partially into an oncoming lane in order to pass the obstacle; Fig. 8 is an illustrative graphical representation showing performance ratings over time for a trial using the second road scenario; Fig. Figure 9 shows a third road scenario in which an obstacle creates a hidden area of ​​oncoming traffic at an intersection; Fig. 10 is an illustrative graphic representation showing performance ratings over time for a trial using the third road scenario; Fig. 11 graphical representations of the performance components and complexity for a cognitive system and a human driver responding to a training data set are shown; Fig.12 graphical representations of the performance components and complexity for a cognitive system and a human driver responding to a second set of historical data are shown; and Fig. Thirteen graphical representations illustrate different performance levels for the human driver or the cognitive system, which are based on the training dataset of Fig. 12 were received. Detailed description

[0014] The following description is by its very nature exemplary. It should be understood that in the drawings, corresponding reference numbers denote identical or equivalent parts and features. As used herein, the term module refers to a processing circuit that may contain an application-specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or grouped), memory that executes one or more software or firmware programs, a combinational logic circuit, and / or other suitable components that provide the described functionality.

[0015] According to an exemplary embodiment, Fig.Figure 1 shows an autonomous vehicle 10 with an associated trajectory planning system, illustrated at Figure 100, according to various embodiments. In general, the trajectory planning system 100 determines a trajectory plan for the automated driving of the autonomous vehicle 10. The autonomous vehicle 10 generally comprises a chassis 12, a body 14, front wheels 16, and rear wheels 18. The body 14 is arranged on the chassis 12 and essentially encloses components of the autonomous vehicle 10. The body 14 and the chassis 12 can together form a frame. The wheels 16 and 18 are rotatably coupled to the chassis 12 near their respective corners.

[0016] In various embodiments, the trajectory planning system 100 is integrated into the autonomous vehicle 10. The autonomous vehicle 10 is, for example, a vehicle that is automatically controlled to transport passengers from one place to another. In the illustrated embodiment, the autonomous vehicle 10 is depicted as a passenger car; however, it should be understood that any other vehicle, including motorcycles, trucks, sport utility vehicles (SUVs), recreational vehicles (RVs), etc., can also be used. At various levels, an autonomous vehicle can assist the driver through a range of methods, such as warning signals to indicate impending risky situations, displays or indicators that enhance the driver's situational awareness by predicting the movements of other actors and warning of potential collisions, etc.The autonomous vehicle has various levels of intervention or control, ranging from coupled assistive vehicle control to complete control of all vehicle functions. In an exemplary embodiment, the autonomous vehicle 10 is a so-called Level 4 or Level 5 automation system. A Level 4 system indicates a "high level of automation," which refers to the driving-mode-specific execution of all aspects of the dynamic driving task by means of an automated driving system, even if a human driver does not respond appropriately to a request for intervention. A Level 5 system indicates "full automation," which refers to the complete execution of all aspects of the dynamic driving task by means of an automated driving system under all road and environmental conditions that can be handled by a human driver.

[0017] As shown, the autonomous vehicle 10 generally comprises a drive system 20, a transmission system 22, a steering system 24, a braking system 26, a sensor system 28, an actuator system 30, a cognitive processor 32, and a controller 34. The drive system 20 may, in various embodiments, comprise an internal combustion engine, an electric machine such as a drive motor, and / or a fuel cell drive system. The transmission system 22 is configured to transmit power from the drive system 20 to the vehicle wheels 16 and 18 according to selectable speed ratios. According to various embodiments, the transmission system 22 may comprise a stepped automatic transmission, a continuously variable transmission, or another suitable transmission. The braking system 26 is configured to provide braking torque to the vehicle wheels 16 and 18.The braking system 26 can, in various embodiments, comprise friction brakes, brake-by-wire brakes, a regenerative braking system such as an electric motor, and / or other suitable braking systems. The steering system 24 influences the position of the vehicle wheels 16 and 18. Although a steering wheel is shown for illustrative purposes, the steering system 24 may, in some embodiments considered within the scope of the present invention, not include a steering wheel.

[0018] The sensor system 28 comprises one or more detection devices 40a-40n that detect observable conditions of the external environment and / or the internal environment of the autonomous vehicle 10. The detection devices 40a-40n may include radar systems, lidar systems, global positioning systems, optical cameras, thermal cameras, ultrasonic sensors, and / or other sensors. The detection devices 40a-40n receive measurements or data relating to various objects or actors 50 within the vehicle's environment. Such actors 50 may be other vehicles, pedestrians, bicycles, motorcycles, etc., as well as stationary objects. The detection devices 40a-40n may also acquire traffic data, such as information about traffic signals and signs, etc.

[0019] The actuator system 30 comprises one or more actuator devices 42a-42n that control one or more vehicle functions, such as the drive system 20, the transmission system 22, the steering system 24, and the braking system 26. In various embodiments, the vehicle features may also include internal and / or external vehicle features such as doors, a trunk, and cabin features such as ventilation, music, lighting, etc. (not numbered).

[0020] The controller 34 comprises a processor 44 and a computer-readable memory device or computer-readable storage medium 46. The processor 44 can be any custom or commercially available processor, central processing unit (CPU), graphics processing unit (GPU), an auxiliary processor among several processors assigned to the controller 34, a semiconductor-based microprocessor (in the form of a microchip or chipset), a macroprocessor, any combination thereof, or generally any device for executing instructions. The computer-readable memory device or computer-readable storage medium 46 can, for example, include volatile and non-volatile memory in read-only memory (ROM), random-access memory (RAM), and keep-alive memory (KAM).KAM is a persistent or non-volatile memory that can be used to store various operating variables while the processor 44 is switched off. The computer-readable memory device(s) or media 46 can be implemented using any number of known memory devices, such as PROMs (programmable read-only memory), EPROMs (electrical PROMs), EEPROMs (electrically erasable PROMs), flash memory, or any other electrical, magnetic, optical, or combined memory devices capable of storing data, some of which represent executable instructions used by the controller 34 in controlling the autonomous vehicle 10.

[0021] The instructions can comprise one or more separate programs, each containing an ordered list of executable instructions for implementing logical functions. When executed by the processor 44, the instructions receive and process signals from the sensor system 28, perform logic, calculations, procedures, and / or algorithms to automatically control the components of the autonomous vehicle 10, and generate control signals for the actuator system 30 to automatically control the components of the autonomous vehicle 10 based on the logic, calculations, procedures, and / or algorithms.

[0022] The controller 34 also communicates with the cognitive processor 32. The cognitive processor 32 receives various data from the controller 34 and from the sensor devices 40a-40n of the sensor system 28 and performs various calculations to provide the controller 34 with a trajectory so that the controller 34 can implement it on the autonomous vehicle 10 via one or more actuator devices 42a-42n. A detailed discussion of the cognitive processor 32 will follow in relation to Fig. 2 given.

[0023] Fig. Figure 2 represents an illustrative control system 200 that includes a cognitive processor 32 integrated into an autonomous vehicle 10. In various embodiments, the autonomous vehicle 10 can be a vehicle simulator that simulates different driving scenarios for the autonomous vehicle 10 and simulates different reactions of the autonomous vehicle 10 to the scenarios.

[0024] The autonomous vehicle 10 includes a data acquisition system 204 (e.g., the sensors 40a-40n of Fig.1) The data acquisition system 204 receives various data to determine the state of the autonomous vehicle 10 and various actors in the environment of the autonomous vehicle 10. Such data includes kinematic data, position or attitude data, etc., of the autonomous vehicle 10, as well as data about other actors, including distance, relative velocity (Doppler), altitude, angular position, etc. The autonomous vehicle 10 also includes a transmit module 206, which packetizes the acquired data and sends the packetized data to the communication interface module 208 of the cognitive processor 32, as discussed herein. The autonomous vehicle 10 also includes a receive module 202, which receives operating commands from the cognitive processor 32 and executes the commands on the autonomous vehicle 10 to navigate the autonomous vehicle 10.The cognitive processor 32 receives the data from the autonomous vehicle 10, calculates a trajectory for the autonomous vehicle 10 based on the provided state information and the methods disclosed herein, and makes the trajectory available to the autonomous vehicle 10 at the receiving module 202. The autonomous vehicle 10 then implements the trajectory provided by the cognitive processor 32.

[0025] The cognitive processor 32 contains various modules for communication with the autonomous vehicle 10, including the interface module 208 for receiving data from the autonomous vehicle 10 and a trajectory transmitter 222 for sending instructions, such as a trajectory, to the autonomous vehicle 10. The cognitive processor 32 also includes a working memory 210, which stores various data received from the autonomous vehicle 10 as well as various intermediate calculations of the cognitive processor 32. A hypothesis module 212 of the cognitive processor 32 is used to propose various hypothetical trajectories and movements of one or more actors in the environment of the autonomous vehicle 10 using a variety of possible prediction methods and state data stored in the working memory 210.A hypothesis resolver 214 of the cognitive processor 32 receives the multitude of hypothetical trajectories for each actor in the environment and determines from the multitude of hypothetical trajectories a most probable trajectory for each actor.

[0026] The cognitive processor 32 further comprises one or more decision-maker modules 216 and a decision resolver 218. The decision-maker module(s) 216 receive the most probable trajectory for each actor in the environment from the hypothesis resolver 214 and calculate a multitude of candidate trajectories and behaviors for the autonomous vehicle 10 based on the most probable actor trajectories. Each of the multitude of candidate trajectories and behaviors is provided to the decision resolver 218. The decision resolver 218 selects or determines an optimal or desired trajectory and a desired behavior for the autonomous vehicle 10 from the candidate trajectories and behaviors.

[0027] The cognitive processor 32 further includes a trajectory planner 220, which determines an autonomous vehicle trajectory that is provided to the autonomous vehicle 10. The trajectory planner 220 receives the vehicle behavior and trajectory from the decision resolver 218, an optimal hypothesis for each actor 50 from the hypothesis resolver 214, and the latest environmental information in the form of "state data" to adjust the trajectory plan. This additional step in the trajectory planner 220 ensures that any anomalous processing delays in the asynchronous calculation of the actor hypotheses are checked against the latest data acquired by the data acquisition system 204. This additional step updates the optimal hypothesis accordingly during the final trajectory calculation in the trajectory planner 220.

[0028] The specified vehicle trajectory is provided by the trajectory planner 220 to the trajectory transmitter 222, which provides a trajectory message to the autonomous vehicle 10 (e.g. at the controller 34) for implementation on the autonomous vehicle 10.

[0029] The cognitive processor 32 also includes a modulator 230, which controls various limits and thresholds for the hypothesis module(s) 212 and the decision module(s) 216. The modulator 230 can also modify parameters for the hypothesis resolver 214 to influence how it selects the optimal hypothesis object for a given actor 50, for the decision-maker, and for the decision resolver. The modulator 230 is a discriminator, which makes the architecture adaptive. The modulator 230 can modify both the computations being performed and the actual result of the deterministic computations by changing parameters in the algorithms themselves.

[0030] An evaluation module 232 of the cognitive processor 32 calculates context-related information, including error measurements, hypothesis confidence measurements, measurements of the complexity of the environment and the state of the autonomous vehicle 10, a performance evaluation of the autonomous vehicle 10 given environmental information, including actor hypotheses and a trajectory of the autonomous vehicle (either historical or future), and delivers this information to the cognitive processor. The modulator 230 receives information from the evaluator 232 to calculate changes to the processing parameters for the hypothesis modules 212, the hypothesis resolver 214, the decision-makers 216, and the thresholds of the decision resolution parameters for the decision resolver 218. A virtual controller 224 implements the trajectory message and determines a feedforward trajectory of various actors 50 in response to the trajectory.

[0031] The modulation occurs in response to the uncertainty measured by the evaluation module 232. In one embodiment, the modulator 230 receives confidence levels associated with hypothesis objects. These confidence levels can be collected from hypothesis objects at a single point in time or over a selected time window. The time window can be variable. The evaluation module 232 determines the entropy of the distribution of these confidence levels. Additionally, historical error measurements on hypothesis objects can also be collected and evaluated in the evaluation module 232.

[0032] These types of evaluations serve as internal context and a measure of uncertainty for the cognitive processor 32. These context-related signals from the evaluation module 232 are used for the hypothesis resolver 214, the decision resolver 218, and the modulator 230, which can modify parameters for hypothesis modules 212 based on the results of the calculations.

[0033] The various modules of the cognitive processor 32 operate independently of each other and are connected to (in Fig. 2. Individual update rates (e.g., LCM-Hz, h-Hz, d-Hz, e-Hz, m-Hz, t-Hz) have been updated.

[0034] During operation, the interface module 208 of the cognitive processor 32 receives the packetized data from the transmitting module 206 of the autonomous vehicle 10 at a data receiver 208a and parses the received data at a data splitter or parser 208b. The data parser 208b converts the data into a data format, referred to herein as a property bag, which is stored in the working memory 210 and can be used by the various hypothesis modules 212, decision modules 216, etc. of the cognitive processor 32.

[0035] Memory 210 extracts information from the collection of property bags during a configurable time window to create snapshots of the autonomous vehicle and various actors. These snapshots are published at a fixed frequency and distributed to participating modules. The data structure generated by Memory 210 from the property bags is a state data structure containing information ordered according to a timestamp. A sequence of generated snapshots therefore includes dynamic state information for a different vehicle or actor. Property bags within a selected state data structure contain information about objects such as other actors, the autonomous vehicle, route information, etc. The property bag for an object contains detailed information about the object, such as its location, speed, heading, etc.This state data structure flows through the rest of the cognitive processor 32 for computations. State data can refer to the states of the autonomous vehicle as well as to the states of actors, etc.

[0036] The hypothesis module(s) 212 pulls state data from memory 210 to compute possible outcomes of the actors in the local environment over a selected timeframe or time step. Alternatively, memory 210 can send state data to the hypothesis module(s) 212. The hypothesis module(s) 212 can comprise a plurality of hypothesis modules, each of which uses a different procedure or technique to determine the possible outcome of the actor(s). One hypothesis module can determine a possible outcome using a kinematic model that applies basic physics and mechanics to data in memory 210 to predict a subsequent state of each actor 50. Other hypothesis modules can predict a subsequent state of each actor 50 by, for example,Apply a kinematic regression tree to the data, apply a Gaussian Mixture Model / Markovian Mixture Model (GMM-HMM) to the data, apply a recursive neural network (RNN) to the data, apply other machine learning processes, apply logic-based inferences to the data, etc. The hypothesis modules 212 are modular components of the cognitive processor 32 and can be added to or removed from the cognitive processor 32 as desired.

[0037] Each hypothesis module 212 contains a hypothesis class for predicting actor behavior. The hypothesis class contains specifications for hypothesis objects and a set of algorithms. Upon invocation, a hypothesis object for an actor is created from the hypothesis class. The hypothesis object adheres to the specifications of the hypothesis class and uses the algorithms of the hypothesis class. A large number of hypothesis objects can be run in parallel. Each hypothesis module 212 creates its own prediction for each actor 50 based on the current working data and sends the prediction back to memory 210 for storage and future use. When new data is supplied to memory 210, each hypothesis module 212 updates its hypothesis and pushes the updated hypothesis back to memory 210.Each hypothesis module 212 can update its hypothesis with its own update rate (e.g., h-Hz). Each hypothesis module 212 can individually function as a subscription service, from which its updated hypothesis is forwarded to relevant modules.

[0038] Each hypothesis object generated by a hypothesis module 212 is a prediction in the form of a state data structure for a time vector, for defined entities such as location, velocity, course, etc. In one embodiment, the hypothesis module(s) 212 can contain a collision detection module that can modify the feedforward flow of information regarding predictions. Specifically, if a hypothesis module 212 predicts a collision between two actors 50, another hypothesis module can be called to generate adjustments to the hypothesis object to account for the expected collision or to send a warning flag to other modules to attempt to mitigate the dangerous scenario or to modify their behavior to avoid the dangerous scenario.

[0039] For each actor 50, the hypothesis resolver 214 receives the relevant hypothesis objects and selects a single hypothesis object from among them. In one embodiment, the hypothesis resolver 214 calls a simple selection process. Alternatively, the hypothesis resolver 214 can apply a fusion process to the different hypothesis objects to create a hybrid hypothesis object.

[0040] Since the architecture of the cognitive processor is asynchronous, the hypothesis resolver 214 and the downstream decision-making modules 216 receive the hypothesis object from this specific hypothesis module at the earliest possible time via a subscription push process if a computation method implemented as a hypothesis object takes longer to complete. Timestamps assigned to a hypothesis object inform the downstream modules about the relevant time frame for the hypothesis object, thus enabling synchronization with hypothesis objects and / or state data from other modules. The time period for which the prediction of the hypothesis object is valid is therefore temporally aligned across modules.

[0041] For example, when a decision-maker module 216 receives a hypothesis object, it compares the hypothesis object's timestamp with a timestamp for the most recent data (i.e., speed, position, direction, or course, etc.) from the autonomous vehicle 10. If the hypothesis object's timestamp is considered too old (e.g., before the autonomous vehicle's data by a selected time criterion), the hypothesis object can be ignored until an updated hypothesis object is received. Updates based on the latest information are also performed by the trajectory planner 220.

[0042] The decision-maker module(s) 216 contains modules that generate various decision proposals in the form of trajectories and behaviors for the autonomous vehicle 10. The decision-maker module(s) 216 receives a hypothesis for each actor 50 from the hypothesis resolver 214 and uses these hypotheses and a nominal target trajectory for the autonomous vehicle 10 as boundary conditions. The decision-maker module(s) 216 can comprise a plurality of decision-maker modules, each of which uses a different procedure or technique to determine a possible trajectory or behavior for the autonomous vehicle 10. Each decision-maker module can operate asynchronously and receives different input states from the working memory 210, such as the hypothesis generated by the hypothesis resolver 214.The decision-maker module(s) 216 are modular components and can be added to or removed from the cognitive processor 32 as needed. Each decision-maker module 216 can update its decisions at its own update rate (e.g., at d Hz).

[0043] Similar to a hypothesis module 212, a decision module 216 contains a decider class for predicting an autonomous vehicle trajectory and / or behavior. The decision class contains specifications for decision object objects and a set of algorithms. Upon invocation, a decision object for an actor 50 is created from the decision class. The decision object adheres to the specifications of the decision class and uses the algorithm of the decision class. A large number of decision objects can be executed in parallel.

[0044] The decision resolver 218 receives the various decisions generated by the one or more decision module(s) and creates a single trajectory and behavior object for the autonomous vehicle 10. The decision resolver can also receive various context information from evaluation module(s) 232, the context information being used to generate the trajectory and behavior object.

[0045] The trajectory planner 220 receives the trajectory and behavior objects from the decision resolver 218 along with the state of the autonomous vehicle 10. The trajectory planner 220 then generates a trajectory message, which is provided to the trajectory sender 222. The trajectory sender 222 provides the trajectory message to the autonomous vehicle 10 for implementation in a format suitable for communication with the autonomous vehicle 10.

[0046] The trajectory sender 222 also sends the trajectory message to a virtual controller 224. The virtual controller 224 provides data in a feedforward loop to the cognitive processor 32. The trajectory sent to the hypothesis module(s) 212 is refined in subsequent calculations by the virtual controller 224 to simulate a set of future states of the autonomous vehicle 10 resulting from the attempt to follow the trajectory. These future states are used by the hypothesis module(s) 212 to perform feedforward predictions.

[0047] Various aspects of the cognitive processor 32 provide feedback loops. A first feedback loop is provided by the virtual controller 224. The virtual controller 224 simulates the operation of the autonomous vehicle 10 based on the provided trajectory and determines or predicts future states that each actor 50 will assume in response to the trajectory taken by the autonomous vehicle 10. These future states of the actors can be made available to the hypothesis modules as part of the first feedback loop.

[0048] A second feedback loop occurs because various modules use historical information in their calculations to learn and update parameters. For example, the hypothesis module(s) 212 can implement their own buffers to store historical state data, regardless of whether the state data comes from an observation or a prediction (e.g., from the virtual controller 224). In a hypothesis module 212 that uses a kinematic regression tree, for instance, historical observation data for each actor is stored for several seconds and used in the calculation for state predictions.

[0049] Hypothesis Resolver 214 is also designed with feedback, as it too uses historical information for calculations. In this case, historical information about observations is used to calculate prediction errors over time and to adjust hypothesis resolution parameters using these errors. A sliding window can be used to select the historical information used for calculating prediction errors and for learning hypothesis resolution parameters. For short-term learning, the sliding window determines the update rate of Hypothesis Resolver 214's parameters. On longer timescales, the prediction errors during a selected episode (such as a left-turn episode) can be aggregated and used to update the parameters after the episode.

[0050] The Decision Resolver 218 also uses historical information for feedback calculations. Historical information about the performance of the autonomous vehicle's trajectories is used to calculate optimal decisions and adjust the decision resolution parameters accordingly. This learning can occur on multiple timescales within the Decision Resolver 218. On the shortest timescale, performance information is continuously calculated using evaluation modules 232 and fed back to the Decision Resolver 218. For example, an algorithm can be used to provide trajectory performance information, which is supplied by a decision module based on several metrics and other contextual information.This contextual information can be used as a reward signal in reinforcement learning processes for the operation of the decision resolver 218 across different time scales. The feedback can be provided asynchronously to the decision resolver 218, and the decision resolver 218 can adapt upon receiving the feedback.

[0051] In various embodiments, a cognitive system such as the cognitive processor 32 can be trained to operate the autonomous vehicle 10 in a manner that simulates or mimics the behavior of a human driver of the vehicle in various traffic situations. In other words, the cognitive system can be trained to suggest an action or trajectory that is the same as, or substantially the same as, an action or trajectory that a human driver would take behind the wheel of the vehicle. The cognitive system can be trained by evaluating its operation in a traffic scenario using one or more human-based evaluation techniques, as described below.

[0052] Fig.Figure 3 shows a schematic diagram 300 illustrating a procedure for training a cognitive system 304 to operate an autonomous vehicle by simulating a human driver. A training dataset 302 is provided to the cognitive system 304 and a human driver 306. The training dataset 302 can be a simulated dataset or a historical dataset. The simulated training dataset can be, for example, a ViRES dataset. The historical dataset can be, for example, a Next Generation Simulation (NGSIM) dataset. The historical data can include data on traffic traversing a selected road segment during a selected time interval.

[0053] In various embodiments, the training dataset 302 contains one or more vehicles belonging to actors. The training dataset 302 can be divided into time intervals of any duration, such as 2-second intervals. When the training dataset 302 is provided to either the cognitive system 304 or the human driver 306, one of the vehicles belonging to actors is selected and designated as the host vehicle (e.g., the autonomous vehicle), and the cognitive system 304 and the human driver 306 operate from the perspective of the designated host vehicle. The cognitive system 304 then plans a path for the autonomous vehicle based on the traffic conditions (i.e., the trajectories and speeds of the other vehicles belonging to actors).This process can be repeated by selecting a different vehicle from the actors as the host vehicle, or by performing the process with a different time interval, or any combination thereof. The planned path generated by cognitive system 304 is sent to a planned path evaluator 308, which generates a system performance evaluation based on the planned path. The planned path evaluator 308 subjects the planned path to various baseline metrics to determine a system performance rating.

[0054] Furthermore, the training dataset 302 is sent to a human driver 306 to evaluate a path driven by the human driver. In various embodiments, the same selected time intervals and host vehicle assignments can be sent to both the cognitive system 304 and the human driver 306. In another embodiment, the actions of the vehicle selected as the assigned host vehicle in the dataset can be used to represent the actions of a human driver. Thus, one of the human drivers 306 and the assigned host vehicle generates or provides a human-driven path from the training dataset. The human-driven path is sent to the planned path evaluator 308, which generates a human-based performance rating for the human driver.

[0055] The system performance rating and the human-based performance rating are sent to a comparison module 310. The comparison module 310 adjusts the rating model of the evaluator 308 for planned paths. In various embodiments, the adjustments reduce the difference between the system performance rating and the human-based performance rating. After the coefficients of the rating model have been adjusted, the rating model can be used in the autonomous vehicle 10 during real-world traffic situations.

[0056] Fig. Figure 4 shows a schematic diagram 400 illustrating another method for training the cognitive system 304 to simulate human driving behavior. The training dataset 302 is provided to the cognitive system 304. The cognitive system 304 plans a path for the autonomous vehicle, as shown in Fig. 3 was described.

[0057] The planned path is sent to a Planned Path Evaluator 308. The Planned Path Evaluator applies various baseline metrics to the planned path to determine a system performance rating. Additionally, the planned path is sent to a Human Evaluator 402. The Human Evaluator 402 assigns a human-based performance rating to the planned path. The system performance rating and the human-based performance rating are then sent to a Comparator 310. The Comparator 310 adjusts the Planned Path Evaluator 308's rating model.

[0058] Fig. Figure 5 shows a schematic diagram 500, which details the procedure of Fig.Figure 4 illustrates the training of the cognitive system. The procedure includes a system-based evaluation path 502 and a human-based evaluation path 504. The system-based evaluation path 502 contains a performance evaluation module 508, which is used in the evaluator 308 for planned paths. The performance evaluation module 508 receives input parameters 506 from the cognitive system in response to the training set. The input parameters 506 include a planned path of the host vehicle (i.e., a planned speed and course (orientation, direction) of the host vehicle), as well as speed and course for each of the multiple actor vehicles. The performance evaluation module 506 generates a system performance grade (Grade A) by subjecting the planned path and the actor input parameters to one or more of the base metrics described below.Alternatively, a complexity value of 516 or a complexity score can be determined based on the input parameters 506.

[0059] The performance evaluation module 506 generates a system performance rating based on a variety of base metrics. In other words, the planned path is evaluated against a multitude of criteria, with each criterion generating a sub-score or sub-rating. Once determined, these sub-ratings are multiplied by assigned coefficients and linearly combined to calculate the system performance rating.

[0060] The illustrative procedure described here has four sub-metrics or criteria: a deviation of the host vehicle from the safe following distance, a deviation of the host vehicle from a safe lane change distance, a collision state, and a deviation of the host vehicle from an average traffic speed.

[0061] The criterion for a deviation from the safe following distance is based on the distance between the host vehicle and the nearest vehicle of an actor that is ahead of the host vehicle and in the same lane. In one embodiment, the safe following distance is based on a two-second rule, which specifies a distance the host vehicle covers in two seconds. For the safe following distance criterion, the host vehicle is penalized as a function of the difference between the safe following distance and the actual following distance.

[0062] The criterion for a deviation from the safe lane-change distance is based on the distance between the host vehicle and the vehicle of an actor in a target lane (e.g., an adjacent lane). The actors located directly in front of and directly behind the host vehicle in the target lane are identified. Under this criterion, the host vehicle is penalized as a function of its distance to the actor behind the host vehicle and its distance to the actor in front of the host vehicle.

[0063] The criterion for a collision state is determined by ascertaining whether the distance between (a center point of the) host vehicle and the nearest actor lies within a collision threshold. This threshold can be calculated based on the shape of the convex hulls of both vehicles. If a collision state is detected, the maximum possible penalty is imposed.

[0064] The criterion for deviation from the average traffic speed is based on the difference between the speed of the host vehicle and the speeds of its surrounding vehicles (actors). The average speed of all actors within the sensor range of the host vehicle is calculated. For this criterion, the host vehicle is penalized as a function of the difference between its speed and the average speed of the other actors.

[0065] The human-based assessment path 504 comprises the clustering module 510 and the human assessment module 512. Clustering module 510 creates vehicle clusters containing host vehicles exhibiting similar behavior. These vehicle clusters are presented to one or more people in the human assessment module 512, who evaluate the behavior of the vehicles within the clusters and assign a rating (GradeH) to each cluster. This rating is then entered into the human assessment module 512.

[0066] The clustering module 510 groups or clusters the vehicles based on the input parameters 506 using a selected clustering method. In one embodiment, the clustering module 510 uses a k-means clustering method. Given a set of observations (x1, x2, ..., xn), where each observation is a d-dimensional real vector, k-means clustering aims to partition the n observations into k (≤n) sets S = {S1, S2, ..., Sk} in order to minimize the sum of squares within the cluster. Formally, the goal, as shown in Eq. (1), is: arg minS∑i=1k∑x∈Si‖x−μi‖2=arg minS∑i=1k|Si|VarSi to find, where µi is the mean of the point in Si This is equivalent to minimizing the pairwise squared deviations of points in the same cluster as in Eq. (2): arg minS∑i=1k12|Si|∑x,y∈Si‖x−y‖2 where x and y are observations.

[0067] In the representation of Fig.In section 5, the vehicles are divided into four categories using k-means clustering, as shown in Panel A. The algorithm attempts to divide datasets into clusters representing the most important elements of the data, maximizing the similarity of elements within the same group and minimizing the similarity of elements across different groups. k-means clustering divides groups of input data into a set num(k) of clusters through an iterative process that first assigns k points as means and then assigns each data point to a cluster based on the mean to which the point is closest. New means are then calculated to form the center of the cluster, which is the mean of the values ​​of the data points within the cluster. The data points are then reassigned to the clusters, and the process is repeated.This ultimately converges to a set of mean values ​​that no longer change and are therefore considered to be in the final grouping. Once the vehicles have been clustered, representative video clips can be shown to people for evaluation.

[0068] Voting folders are created. Each voting folder contains an equal number of vehicles from each cluster. For each folder, the performance evaluation algorithm is applied to generate a performance score based on the four subcomponents discussed here. Each folder is also provided to a human test subject who rates each vehicle on a scale of 1 to 4, where 1 is the worst and 4 is the best. The human assesses the vehicle based on its ability to maintain a safe distance, maintain speed relative to the traffic flow, change lanes, and avoid collisions. This allows for a comparison of the human assessment with the human-based performance score.

[0069] The results are analyzed in comparison module 514 using a generalized linear model (GLM) to determine which basic metrics of the evaluation model are relevant for the respective scenario. The numerical coefficients of the evaluation model are extracted, and weights are assigned to the basic metrics.

[0070] The comparison module 514 compares the human-based rating or classification with the system rating and determines adjustments to the rating model of the performance rating module 508 that align the system performance score with the human-based performance rating. These adjustments can then be applied to coefficients of the rating model of the performance rating module 508. Once the coefficients of the rating model have been adjusted, the rating model can be used in the autonomous vehicle 10 during real-world traffic situations.

[0071] Fig. 6, Fig. 7 and Fig.Figure 9 shows various street scenarios that can be used as a training dataset to familiarize the cognitive system with the information presented in Fig. to train according to the 3 disclosed methods.

[0072] Fig.Figure 6 shows a first road scenario 600 for evaluating the performance of the cognitive system. In the first road scenario 600, a host vehicle 602 is identified, and the host vehicle's performance is evaluated using the procedures disclosed here. In the first road scenario, the host vehicle 602 approaches an intersection to make a left turn. There is no stop sign at the intersection to stop the vehicles 604 of actors forming the cross traffic. A rating of the first scenario is based on the host vehicle 602 observing the actions of the vehicles 604 of actors in the cross traffic at the intersection and finding a safe time interval for the left turn. Table 1 shows some illustrative ratings obtained over four trials in response to the first road scenario by both a human driver and the cognitive system. Table 1 Attempt Person Cognitive 1 89 94 2 84 97 3 80 92 4 80 90 Average 83 93

[0073] The results from Table 1 show that the performance of the cognitive system is better than that of the human driver in each of the trials.

[0074] Fig. Figure 7 shows a second road scenario 700, in which a construction site forms an obstacle 702, forcing the host vehicle 602 to at least partially switch to an oncoming lane to pass the obstacle. The rating of the second scenario is based on the host vehicle 602 observing the actions of the other vehicles 604 in the oncoming lane and finding a safe time interval to move into the oncoming lane while passing the obstacle. Table 2 shows some illustrative ratings obtained over three trials in response to the second road scenario from both a human driver and the cognitive system. Table 2 Attempt Person Cognitive 1 89 82 2 90 92 3 88 93 Average 88 89

[0075] The results from Table 2 show that the performance of the cognitive system in each of the trials is similar to or better than that of the human driver.

[0076] Fig. Figure 8 is an illustrative graph showing the performance grades over time for a trial using the second road scenario. Time is represented on the x-axis and performance grade on the y-axis. Curve 802 shows the performance grade over time for the cognitive system, and curve 804 shows the performance grade over time for the human driver. By examining the performance grade curves of Fig.8. The cognitive system makes more decisions that comply with the rules and / or criteria for safe following distance and traffic flow. Furthermore, the cognitive system does not perform agile lane changes, which negatively impacts the system performance rating. Consequently, the average system performance rating (curve 802) is higher than the performance rating (curve 804) for the human driver, as shown in Table 2.

[0077] Fig.Figure 9 shows a third road scenario 900, in which an obstacle creates an obscured area 902 of oncoming traffic at an intersection. An evaluation of the third scenario is based on the host vehicle 602 approaching the intersection and watching for oncoming vehicles 604, in order to make an uncontrolled left turn. Table 3 shows some illustrative ratings obtained over five trials in response to the third road scenario from both a human driver and the cognitive system. Table 3 Attempt Person Cognitive 1 87 93 2 89 95 3 88 99 4 94 96 5 92 97 Average 90 96

[0078] The results from Table 3 show that the performance of the cognitive system is better than that of the human driver in each of the trials.

[0079] Similar to the second road scenario, the performance components within the observed period show that the cognitive system achieves higher performance values ​​or scores. This is due to the cognitive system employing a more conservative / cautious driving style than the human driver.

[0080] Fig.Figure 10 is an illustrative graph 1000, representing the performance ratings over time for a trial using the third road scenario. Time is plotted on the x-axis and performance rating on the y-axis. Curve 1002 shows the performance rating over time for the cognitive system, and curve 1004 shows the performance rating over time for the human driver. Similar to the second road scenario, the performance components within the trial show that the cognitive system achieves higher performance scores. This is because the cognitive system employs a more conservative / cautious driving pattern than the human driver.

[0081] Fig.Figure 11 shows graphical representations 1100 of the performance components 1102 and the complexity 1104 for a cognitive system 1108 and a human driver 1106 in response to a training dataset. The training dataset is a historical dataset recorded from above a section of highway during a selected time interval of approximately 175 seconds to approximately 205 seconds. This time interval can be divided into a first sub-interval, lasting from approximately 175 seconds to approximately 190 seconds, and a second sub-interval, lasting from approximately 190 seconds to approximately 205 seconds, based on the nature of the complexity during these intervals. The complexity in the first sub-interval is, on average, higher than the complexity in the second sub-interval. Furthermore, the complexity in the first sub-interval shows greater temporal variation than the complexity in the second sub-interval.During the first sub-interval, the human driver (1106) performs better on average than the cognitive system (1108). During the second sub-interval, the cognitive system (1108) performs better on average than the human driver (1106). For the entire time interval, the average performance of the cognitive system (1108) is 72, and the average performance compared to the human driver (1106) is 80.

[0082] Fig.Figure 12 shows graphical representations of the performance components 1202 and the complexity 1204 for a cognitive system 1108 and a human driver 1106 in response to a second set of historical data recorded over a time interval from 230 seconds to 265 seconds. The complexity (C) of the time interval remains fairly constant over this interval, except for a dramatic increase at approximately 249 seconds. For the entire time interval, the average performance (%) of the cognitive system is 77, and the average performance relative to the human driver is 71.

[0083] Fig. Figure 13 shows graphical representations 1300 and 1310 illustrating different performance sub-grades for the human driver and the cognitive system, respectively, which are derived from the training data set of Fig.12 were obtained. Graph 1300 shows the performance sub-grades or partial scores obtained over time for the human driver. Curve 1302 is a partial classification for the safety distance criterion, curve 1304 is the partial classification for the collision criterion, curve 1306 is the partial classification for the speed deviation criterion, and curve 1308 is the partial classification for the lane change criterion. Curve 1310 shows the performance sub-classifications obtained over time for the cognitive system. Curve 1312 is the partial classification for the safe distance criterion, curve 1314 is the partial classification for the collision criterion, curve 1316 is the partial classification for the speed deviation criterion, and curve 1318 is the partial classification for the lane change criterion. If one considers the graphic representations from 1300 and 1310 as well as the graphic representation from 1200 by Fig. 12, thus one comes to the conclusion that the cognitive system achieves the higher average performance rating primarily by reducing the deviation from the average traffic speed.

Claims

[1] Method for operating an autonomous vehicle (10), comprising: Operating a cognitive system (304) in response to a training data set (302) to generate a planned action for operating the autonomous vehicle (10); Evaluating the planned action to obtain a system performance rating; Updating the cognitive system (304) based on a comparison of the system performance rating with a human-based performance rating; and Operating the autonomous vehicle (10) using the cognitive system (304); wherein the human-based performance rating is obtained by evaluating a path driven by a human in relation to the training data set (302), wherein the evaluation is carried out by one or more humans; and / or the procedure further involves updating the cognitive system (304) by reducing a difference between the system performance rating and the human-based performance rating; and / or wherein an evaluation model generates the system performance rating, wherein the evaluation model includes at least one base metric weighted with a coefficient, and wherein the method further comprises adjusting the coefficient of the at least one base metric based on the comparison; and / or wherein the procedure further includes determining a complexity score which indicates a difficulty level of a driving scenario for the autonomous vehicle (10) and evaluating the planned action using the system performance rating, the human-based performance rating and the complexity score. [2] System for operating an autonomous vehicle (10), comprising: a control system (200) for performing a driving action in the autonomous vehicle; (10) and a cognitive system (304) for generating the driving action using an evaluation model, wherein the evaluation model is generated by: Operating the cognitive system (304) in response to a training data set (302) to generate a planned action for operating the autonomous vehicle (10) by the cognitive system; Evaluating the planned action to obtain a system performance rating; and Updating the cognitive system (304) based on a comparison of the system performance rating with a human-based performance rating; wherein the human-based performance rating is based on a path driven by a human in relation to the training data set (302), wherein the assessment is carried out by one or more humans; and / or wherein the system further comprises a comparison module (310) for updating the cognitive system (304) by reducing a difference between the system performance rating and the human-based performance rating; and / or wherein the rating model generates the system performance rating and includes at least one base metric weighted with a coefficient, wherein the system further includes a comparison model for adjusting the coefficient of the at least one base metric based on the comparison. [3] System according to claim 2, wherein the comparison module (310) evaluates the planned action using the system performance rating, the human-based performance rating and a complexity score.

Citation Information

Patent Citations

  • EVALUATING AUTONOMOUS VEHICLE TRAJECTORIES USING ADEQUATE QUANTITY DATA

    DE102020128156A1