Automated driving systems with expected levels of driving aggression

By selecting reinforcement learning agents based on sensed driving environment and user preferences, the problem of poor driving style adjustment in autonomous driving systems has been solved, thereby improving driving experience and safety.

CN116300853BActive Publication Date: 2026-04-03GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing autonomous driving systems struggle to adjust driving styles in real time based on driving environment and user preferences, resulting in a poor driving experience.

Method used

The system employs a sensor-based approach to select reinforcement learning agents based on the driving environment and user preferences. By calculating challenge scores and desired driving styles, it selects the appropriate agent from multiple reinforcement learning agents to generate driving actions.

Benefits of technology

It enables dynamic adjustment of driving style based on driving environment and user preferences, improving driving experience and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116300853B_ABST
    Figure CN116300853B_ABST
Patent Text Reader

Abstract

This invention relates to an automated driving system with a desired level of driving aggression. One system includes a computer comprising a processor and a memory. The memory includes instructions that program the processor to: receive sensor data representing a perceived driving environment; select a reinforcement learning agent from a plurality of reinforcement learning agents based on a challenge score calculated using the sensor data and a desired driving style; and generate driving actions via the selected reinforcement learning agent based on the sensor data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to selecting reinforcement learning agents to operate a vehicle based on sensed driving environment and user preferences. Background Technology

[0002] Reinforcement learning systems include agents that interact with the environment by performing actions selected by the reinforcement learning system in response to observations received that represent the current state of the environment. Summary of the Invention

[0003] A system includes a computer comprising a processor and a memory. The memory includes instructions that cause the processor to: receive sensor data representing a perceived driving environment; select a reinforcement learning agent from a plurality of reinforcement learning agents based on a challenge score calculated using the sensor data and a desired driving style; and generate driving actions via the selected reinforcement learning agent based on the sensor data.

[0004] Among other features, each of the plurality of reinforcement learning agents corresponds to a different challenge score and a desired driving style.

[0005] Among the other characteristics, the desired driving style corresponds to the desired level of driving aggression.

[0006] Among other characteristics, the expected level of driving aggression corresponds to completing driving actions within a specific time period.

[0007] Among other features, the plurality of reinforcement learning agents includes M×N reinforcement learning agents, where M is an integer representing M driving preference levels and N is an integer representing N numbers of driving environments.

[0008] Among other features, the processor is further programmed to automatically select another reinforcement learning agent from the plurality of reinforcement learning agents based on sensor data representing different perceptions of the driving environment.

[0009] Among the other features, the desired driving style is received from the user.

[0010] Among other features, the desired driving style is received from the human-machine interface (HMI).

[0011] Among other features, the vehicle is operated based on driving actions.

[0012] Among the other features, the vehicle includes at least one of a land vehicle, an air vehicle, or a water vehicle.

[0013] One method includes: receiving sensor data representing a perceived driving environment; selecting a reinforcement learning agent from a plurality of reinforcement learning agents based on a challenge score calculated using the sensor data and a desired driving style; and generating driving actions via the selected reinforcement learning agent based on the sensor data.

[0014] Among other features, each of the plurality of reinforcement learning agents corresponds to a different challenge score and a desired driving style.

[0015] Among the other characteristics, the desired driving style corresponds to the desired level of driving aggression.

[0016] Among other characteristics, the expected level of driving aggression corresponds to completing driving actions within a specific time period.

[0017] Among other features, the plurality of reinforcement learning agents includes M×N reinforcement learning agents, where M is an integer representing M driving preference levels and N is an integer representing N numbers of driving environments.

[0018] Among other features, the method includes: automatically selecting another reinforcement learning agent from the plurality of reinforcement learning agents based on sensor data representing different perceptions of the driving environment.

[0019] Among the other features, the desired driving style is received from the user.

[0020] Among other features, the desired driving style is received from the human-machine interface (HMI).

[0021] Among other features, the method includes: operating based on driving actions.

[0022] Among the other features, the vehicle includes at least one of a land vehicle, an air vehicle, or a water vehicle.

[0023] 1. A system comprising a computer, the computer including a processor and a memory, the memory including instructions that program the processor to:

[0024] Receive sensor data representing the perceived driving environment;

[0025] Based on the challenge score calculated using the sensor data and the desired driving style, a reinforcement learning agent is selected from multiple reinforcement learning agents; and

[0026] Driving actions are generated based on the sensor data via a selected reinforcement learning agent.

[0027] 2. The system according to Scheme 1, wherein each of the plurality of reinforcement learning agents corresponds to a different challenge score and a desired driving style.

[0028] 3. The system according to Scheme 1, wherein the desired driving style corresponds to the desired level of driving aggression.

[0029] 4. The system according to Scheme 1, wherein the desired level of driving aggression corresponds to completing the driving action within a specific time period.

[0030] 5. The system according to Scheme 1, wherein the plurality of reinforcement learning agents includes M×N reinforcement learning agents, where M is an integer representing M driving preference levels and N is an integer representing N numbers of driving environments.

[0031] 6. The system according to claim 5, wherein the processor is further programmed to automatically select another reinforcement learning agent from the plurality of reinforcement learning agents based on the sensor data representing different perceptions of the driving environment.

[0032] 7. The system according to Scheme 1, wherein the desired driving style is received from the user.

[0033] 8. The system according to claim 7, wherein the desired driving style is received from a human-machine interface (HMI).

[0034] 9. The system according to claim 1, wherein the vehicle is operated according to the driving action.

[0035] 10. The system according to claim 9, wherein the vehicle includes at least one of a land vehicle, an air vehicle, or a water vehicle.

[0036] 11. A method comprising:

[0037] Receive sensor data representing the perceived driving environment;

[0038] Based on the challenge score calculated using the sensor data and the desired driving style, a reinforcement learning agent is selected from multiple reinforcement learning agents; and

[0039] Driving actions are generated based on the sensor data via a selected reinforcement learning agent.

[0040] 12. The method according to Scheme 11, wherein each of the plurality of reinforcement learning agents corresponds to a different challenge score and a desired driving style.

[0041] 13. The method according to Scheme 11, wherein the desired driving style corresponds to the desired level of driving aggression.

[0042] 14. The method according to Scheme 11, wherein the expected level of driving aggression corresponds to completing the driving action within a specific time period.

[0043] 15. The method according to Scheme 11, wherein the plurality of reinforcement learning agents includes M×N reinforcement learning agents, where M is an integer representing M driving preference levels and N is an integer representing N numbers of driving environments.

[0044] 16. The method according to claim 11, further comprising: automatically selecting another reinforcement learning agent from the plurality of reinforcement learning agents based on the sensor data representing different perceptions of the driving environment.

[0045] 17. The method according to Scheme 11, wherein the desired driving style is received from the user.

[0046] 18. The method according to claim 17, wherein the desired driving style is received from a human-machine interface (HMI).

[0047] 19. The method according to claim 11, further comprising: operating according to the driving action.

[0048] 20. The method according to claim 19, wherein the vehicle includes at least one of a land vehicle, an air vehicle, or a water vehicle.

[0049] Other application areas will become apparent from the description provided herein. It should be understood that the description and specific examples are for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description

[0050] The accompanying drawings described herein are for illustrative purposes only and are not intended to limit the scope of this disclosure in any way.

[0051] Figure 1 This is a block diagram of an example system including a vehicle;

[0052] Figure 2 This is a block diagram of the sample server within the system;

[0053] Figure 3 This is a block diagram of an example computing device;

[0054] Figure 4 This is a diagram of an example neural network;

[0055] Figure 5This is a diagram illustrating an example process for training multiple reinforcement learning agents;

[0056] Figure 6 This is a block diagram illustrating a reinforcement learning system used to select a reinforcement learning agent from multiple reinforcement learning agents to operate a vehicle.

[0057] Figure 7 This is a floor plan of an example driving environment; and

[0058] Figure 8 This is a flowchart illustrating an example process for selecting a reinforcement learning agent from multiple reinforcement learning agents to operate a vehicle. Detailed Implementation

[0059] The following description is exemplary in nature and is not intended to limit this disclosure, application, or use.

[0060] Reinforcement learning (RL) is a form of goal-oriented machine learning. For example, an agent can learn from direct interactions with its environment without relying on explicit supervision and / or a complete model of the environment. Reinforcement learning is a framework that models the interaction between a learning agent and its environment based on state, action, and reward. At each time step, the agent receives a state, selects an action based on a policy, receives a scalar reward, and transitions to the next state. This state can be based on one or more sensor inputs indicative of environmental data. The agent's goal is to maximize the expected cumulative reward. The agent can receive positive scalar rewards for positive actions and negative scalar rewards for negative actions. Thus, the agent "learns" by attempting to maximize the expected cumulative reward. Although the agent is described in the context of a vehicle in this paper, it should be understood that the agent can include any suitable reinforcement learning agent.

[0061] As discussed in more detail herein, a vehicle may include multiple reinforcement learning agents. Each agent is trained to generate outputs representing driving actions based on a challenge score and a user selection representing a desired level of driving aggression, the challenge score corresponding to the perceived difficulty from the sensed driving environment. The desired level can correspond to the user's preferred driving style, such as a relatively conservative or relatively aggressive driving style.

[0062] Figure 1This is a block diagram of an example vehicle system 100. System 100 includes a vehicle 105, which can include land vehicles (such as cars, trucks, etc.), air vehicles, and / or water vehicles. Vehicle 105 includes a computer 110, vehicle sensors 115, actuators 120 for actuating various vehicle components 125, and a vehicle communication module 130. The communication module 130 allows the computer 110 to communicate with a server 145 via a network 135.

[0063] Computer 110 can operate vehicle 105 in autonomous, semi-autonomous, or non-autonomous (manual) modes. For the purposes of this disclosure, autonomous mode is defined as a mode in which each of the propulsion, braking, and steering of vehicle 105 is controlled by computer 110; in semi-autonomous mode, computer 110 controls one or both of the propulsion, braking, and steering of vehicle 105; and in non-autonomous mode, a human operator controls each of the propulsion, braking, and steering of vehicle 105.

[0064] Computer 110 may include programming to operate vehicle 105 braking, propulsion (e.g., controlling vehicle acceleration by controlling one or more of an internal combustion engine, electric motor, hybrid engine, etc.), steering, climate control, interior lights and / or exterior lights, etc., and to determine whether and when computer 110 (not a human operator) controls such operations. Additionally, computer 110 may be programmed to determine whether and when a human operator controls such operations.

[0065] Computer 110 may include more than one processor, or be communicatively coupled to said more than one processor, for example via vehicle 105 communication module 130 as further described below. The more than one processor may be included, for example, in an electronic control unit (ECU) or the like included in vehicle 105 to monitor and / or control various vehicle components 125, such as powertrain controllers, brake controllers, steering controllers, etc. Further, computer 110 may communicate with a navigation system using a Global Positioning System (GPS) via vehicle 105 communication module 130. As an example, computer 110 may request and receive location data of vehicle 105. The location data may be in a known form, such as geographic coordinates (latitude and longitude coordinates).

[0066] Computer 110 is typically arranged for communication on vehicle 105 communication module 130 and also utilizes wired and / or wireless networks within vehicle 105 (e.g., buses in vehicle 105, such as controller area network (CAN) and / or other wired and / or wireless mechanisms) for communication.

[0067] Via the vehicle 105 communication network, the computer 110 can transmit messages to and / or receive messages from various devices within the vehicle 105, such as vehicle sensors 115, actuators 120, vehicle components 125, human-machine interfaces (HMIs), etc. Alternatively or additionally, where the computer 110 actually comprises multiple devices, the vehicle 105 communication network can be used for communication between devices represented herein as computer 110. Further, as mentioned below, various controllers and / or vehicle sensors 115 can provide data to the computer 110. The vehicle 105 communication network can include one or more gateway modules that provide interoperability between various networks and devices within the vehicle 105, such as protocol converters, impedance matching devices, rate converters, etc.

[0068] Vehicle sensor 115 may include various devices, such as those known to provide data to computer 110. For example, vehicle sensor 115 may include multiple light detection and ranging (LiDAR) sensors 115 positioned on the top of vehicle 105, behind the windshield of vehicle 105, or surrounding vehicle 105, providing information on the relative position, size, and shape of objects around vehicle 105 and / or the conditions around vehicle 105. As another example, one or more radar sensors 115 fixed to the bumper of vehicle 105 may provide data to provide and range position relative to vehicle 105, speed of objects (potentially including a second vehicle 106), etc. Vehicle sensor 115 may further include multiple camera sensors 115 (e.g., forward-looking, side-looking, rear-looking, etc.) to provide images from the interior and / or exterior of vehicle 105.

[0069] The actuator 120 of vehicle 105 is implemented via circuits, chips, motors, or other electronic and / or mechanical components that can actuate various vehicle subsystems according to appropriate control signals, as is known. The actuator 120 can be used to control components 125, including braking, acceleration, and steering of vehicle 105.

[0070] In the context of this disclosure, vehicle component 125 is one or more hardware components adapted to perform mechanical or electromechanical functions or operations, such as moving vehicle 105, decelerating or stopping vehicle 105, steering vehicle 105, etc. Non-limiting examples of component 125 include propulsion components (which include, for example, internal combustion engines and / or electric motors), transmission components, steering components (e.g., which may include one or more of a steering wheel, steering rack, etc.), braking components (as described below), parking assist components, adaptive cruise control components, adaptive steering components, movable seats, etc.

[0071] Additionally, computer 110 may be configured to communicate with devices outside vehicle 105 via vehicle-to-vehicle communication module or interface 130, for example, with another vehicle or remote server 145 (typically via network 135) via vehicle-to-vehicle (V2V) or vehicle-to-infrastructure (V2X) wireless communication. Module 130 may include one or more mechanisms through which computer 110 can communicate, including any desired combination of wireless (e.g., cellular, wireless, satellite, microwave, and radio frequency) communication mechanisms and any desired network topology (or topology when multiple communication mechanisms are utilized). Exemplary communications provided via module 130 include cellular, Bluetooth®, IEEE 802.11, Private Short Range Communication (DSRC), and / or Wide Area Network (WAN), including the Internet, thereby providing data communication services.

[0072] Network 135 can be one or more of various wired or wireless communication mechanisms, including any desired combination of wired (e.g., cable and fiber optic) and / or wireless (e.g., cellular, wireless, satellite, microwave, and radio frequency) communication mechanisms, and any desired network topology (or topology when multiple communication mechanisms are utilized). Exemplary communication networks include wireless communication networks (e.g., using Bluetooth, Bluetooth Low Energy (BLE), IEEE 802.11, vehicle-to-vehicle (V2V) (such as Dedicated Short Range Communication (DSRC), etc.), local area networks (LANs), and / or wide area networks (WANs), including the Internet, to provide data communication services.

[0073] Computer 110 is capable of receiving and analyzing data from sensor 115 substantially continuously, periodically, and / or when instructed by server 145. Furthermore, based on data from lidar sensor 115, camera sensor 115, etc., object classification or recognition techniques can be used in computer 110 to identify the type of object (e.g., vehicle, person, rock, pothole, bicycle, motorcycle, etc.) and the physical characteristics of the object.

[0074] As described in more detail herein, computer 110 is configured to implement a neural network-based reinforcement learning program. Computer 110 generates a set of state-action (Q-values) as output in response to observed input states. Computer 110 is capable of selecting the action corresponding to the maximum state-action value (e.g., the highest state-action value). Computer 110 obtains sensor data corresponding to the observed input states from sensor 115.

[0075] Figure 2 The illustration shows an example server 145 including a reinforcement learning (RL) system 205. As shown, the RL system 205 may include a reinforcement learning (RL) agent module 210, one or more RL agents 215, and a storage module 220.

[0076] Specifically, the RL agent module 210 is capable of managing, maintaining, training, implementing, utilizing, or communicating with one or more RL agents 215. For example, the RL agent module 210 can communicate with the storage module 220 to access one or more RL agents 215. The RL agent module 210 is also capable of accessing data that specifically describes the policies of different numbers of learners, which will be discussed below. Figure 5 Let me describe it in more detail.

[0077] Figure 3 An example computing device 300 is illustrated, namely, a computer 110 and / or / more than one server 145 that may be configured to perform one or more of the processes described herein. As shown, the computing device may include a processor 305, a memory 310, a storage device 315, an I / O interface 320, and a communication interface 325. Furthermore, the computing device 300 may include input devices such as a touchscreen, a mouse, a keyboard, etc. In some embodiments, the computing device 300 may include a larger... Figure 3 The components shown are fewer or more components.

[0078] In a particular implementation, processor(s) 305 includes hardware for executing instructions (such as instructions constituting a computer program). By way of example and not limitation, in order to execute instructions, processor(s) 305 may retrieve (or obtain) instructions from internal registers, internal caches, memory 310, or storage device 315 and decode and execute them.

[0079] The computing device 300 includes a memory 310 coupled to processor(s) 305. The memory 310 can be used to store data, metadata, and programs to be executed by the processor(s). The memory 310 may include one or more of volatile and non-volatile memories, such as random access memory ("RAM"), read-only memory ("ROM"), solid-state drive ("SSD"), flash memory, phase-change memory ("PCM"), or other types of data storage devices. The memory 310 may be internal or distributed memory.

[0080] Computing device 300 includes storage device 315, which includes memory for storing data or instructions. By way of example and not limitation, storage device 315 may include the non-transitory storage media described above. Storage device 315 may include hard disk drive (HDD), flash memory, universal serial bus (USB) drive, or combinations of these or other storage devices.

[0081] The computing device 300 also includes one or more input or output (“I / O” devices / interfaces 320, which are provided to allow a user to provide input (such as user strokes) to the computing device 300, receive output from the computing device 300, and otherwise transfer data to and from the computing device 300. These I / O devices / interfaces 320 may include a mouse, keypad or keyboard, touchscreen, camera, optical scanner, network interface, modem, other known I / O devices, or combinations of such I / O devices / interfaces 320. The touchscreen can be activated using a writing device or a finger.

[0082] I / O device / interface 320 may include one or more means for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In some embodiments, device / interface 320 is configured to provide graphics data to a display for presentation to a user. The graphics data may represent one or more graphical user interfaces and / or any other graphical content that may be available for a particular embodiment.

[0083] The computing device 300 may further include a communication interface 325. The communication interface 325 may include hardware, software, or both. The communication interface 325 provides one or more interfaces for communication (such as, for example, packet-based communication) between the computing device and one or more other computing devices 300 or one or more networks. By way of example and not by way of limitation, the communication interface 325 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks (such as Wi-Fi). The computing device 300 may further include a bus 330. The bus 330 may include hardware, software, or both, which interconnects components of the computing device 300.

[0084] Figure 4 This is a graph of an example deep neural network (DNN) 400 that may be used in this paper. Within this context, the DNN 400 may include a single RL agent 215. The DNN 400 includes multiple nodes 405, and the nodes 405 are arranged such that the DNN 400 includes an input layer 410, one or more hidden layers 415, and an output layer 420. Each layer of the DNN 400 can include multiple nodes 405. Although... Figure 4 The diagram illustrates three (3) hidden layers 415, but it should be understood that the DNN 400 can include additional or fewer hidden layers. The input layer 410 and the output layer 420 may also include more than one (1) node 405.

[0085] Nodes 405 are sometimes referred to as artificial neurons because they are designed to mimic biological (e.g., human) neurons. A set of inputs to each node 405 (indicated by arrows) is each multiplied by its respective weight. The weighted inputs are then summed in an input function to provide a net input, which may be adjusted for bias. This net input is then provided to an activation function, which in turn provides the output to the connected nodes 405. The activation function can be a variety of suitable functions, typically chosen based on empirical analysis. For example, by... Figure 4 As illustrated by the arrow in the diagram, it is then possible to provide the output of node 405 so that it can be included in a set of inputs to one or more neurons 305 in the next layer.

[0086] The DNN 400 can be trained to accept sensor data as input and generate an output-action, such as a reward value, based on that input. The DNN 400 can be trained with training data (e.g., a known set of sensor inputs) to train an agent for the purpose of determining an optimal policy. In one or more embodiments, the DNN 400 is trained via server 145, and the trained DNN 400 can be transmitted to vehicle 105 via network 135. For example, weights can be initialized using a Gaussian distribution, and the bias of each neuron 405 can be set to zero. Training the DNN 400 can include updating the weights and biases via appropriate techniques, such as optimized backpropagation.

[0087] During operation, computer 110 acquires sensor data from sensor 115 and provides this data as input to DNN 400, such as (multiple) RL agents 215. Once trained, RL agent 215 is able to accept sensor input and provide one or more state-action values ​​(Q-values) as output based on the sensed input. During the execution of RL agent 215, state-action values ​​can be generated for each action available to the agent within the environment. In an example implementation, RL agent 215 is trained according to a baseline policy. The baseline policy can include one or more state-action values ​​corresponding to a set of sensor input data (corresponding to a baseline driving environment).

[0088] In other words, once the RL agent 215 has been trained, it generates output data reflecting its decisions to take specific actions in response to specific input data. The input data includes, for example, the values ​​of multiple state variables related to the environment the RL agent 215 is exploring or the task it is performing. In some cases, one or more state variables may be one-dimensional. In others, one or more state variables may be multi-dimensional. State variables may also be referred to as features. The mapping from input data to output data may be called a policy and governs the decisions of the RL agent 215. For example, a policy may include a probability distribution of a specific action given a specific value of a given state variable at a given time step.

[0089] Figure 5 The illustration shows an example procedure 500 for initializing each RL agent 215 at server 145. Blocks of procedure 500 can be executed by server 145. At block 505, the RL agent 215 is trained according to the baseline policy.

[0090] At block 510, it is determined whether the average reward value has converged. If the average reward value has not yet converged, one or more hyperparameters are modified at block 515 and provided to the RL agent 215. Hyperparameters are values ​​that are initialized before training begins and modified during the training of the RL agent 215, and that affect the training process of the RL agent 215. For example, hyperparameters can include reward hyperparameters, which include the feedback or observations used by the RL agent 215 to determine appropriate state-actions based on sensor data. Otherwise, at block 520, the RL agent 215 is replicated based on the baseline policy and / or the modified hyperparameters provided in the converged average reward value. It should be understood that the driving environment may include M levels of driving preference, such as the expected level of driving aggression.

[0091] At block 525, a prescriptive analysis framework is used to generate N numbers of driving environments, where M and N are integers greater than or equal to one (1). At block 530, a descriptive challenge score is calculated, which corresponds to one of the driving environments generated at block 525. In the example implementation, the descriptive analysis framework can be used to generate the descriptive challenge score. The descriptive challenge score can include a quantified level of difficulty corresponding to the driving environment. It should be understood that the challenge score can be calculated based on perceived traffic conditions, weather conditions (i.e., severe weather), snow-covered roads, traffic congestion, planned maneuvers, etc.

[0092] At block 535, N driving environments are sorted and provided to RL agent 215. For example, the N driving environments are sorted according to descriptive challenge scores. Figure 5As shown, it is possible to generate M×N RL agents 215, for example, RL agents 215-1 to 215-P, according to the steps described above, where P is the value of M×N.

[0093] At block 540, the reward value of RL agent 215 is set and / or updated. Initially, the reward value can be set to the converged average reward value described above in block 510. At block 545, one or more weights are generated based on the initial benchmark score and provided to each RL agent 215. For example, one or more weights are calculated based on a desired level of aggression. For example, the desired level of aggression may represent completing a driving action within a specific time frame given a sensed environment. In another example, the desired level of aggression may include avoiding stop-and-run actions given a sensed environment.

[0094] At block 550, each RL agent 215 can be trained for a predetermined number of AI training epochs. For example, each RL agent 215 is provided with data representing a specific driving environment as input and generates an output representing driving actions within that specific driving environment. The resulting output can represent the updated baseline score. At block 555, the updated baseline score can be compared with the previous score to determine whether one or more reward values ​​need to be updated. If the comparison result is greater than a predetermined difference threshold, the reward value is updated, and process 500 returns to block 550. It should be understood that process 500 can be executed offline.

[0095] Figure 6 This is an example environment 600 used to operate vehicle 105. Therefore, computer 110 receives sensing data 605 representing the observed environment. Sensing data 605 is generated by vehicle sensors 115. Computer 110 calculates a challenge score 610 based on the sensing data 605. The challenge score 610 represents the driving difficulty corresponding to the sensed driving environment, such as the number of other vehicles approaching vehicle 105, traffic congestion, etc.

[0096] Computer 110 also receives user selections representing a desired driving style 615 (i.e., a desired driving aggressiveness preference). For example, a first desired driving style 615 could correspond to performing driving actions that favor caution. In another example, a second desired driving style 615 could correspond to performing driving actions that favor saving driving time. User selections can be received via an HMI (such as vehicle component 125, a telecomputing device, etc.). In some implementations, computer 110 can suggest modifying the desired driving style 615 based on the driving environment. For example, computer 110 can calculate a challenge score 610 based on the perceived driving environment. Computer 110 can access a lookup table that associates the challenge score 610 with an acceptable driving aggressiveness preference 615 determined during the training of the RL agent 215.

[0097] Then, computer 110 can use selector module 620 to select one of P stored RL agents 215. For example, selector module 620 selects RL agent 215 based on challenge score 610 and desired driving style 615. The selected RL agent 215 is capable of generating output representing driving actions provided to vehicle actuator 120 to operate vehicle 105 accordingly. In some embodiments, computer 110 may automatically select RL agent 215 based on the calculated challenge score 610. For example, computer 110 may determine based on the perceived driving environment that the current driving aggression preference 615 does not correspond to the calculated challenge score 610. In this example, computer 110 suggests to (multiple) passengers via HMI that the desired driving style 615 be modified to an appropriate level.

[0098] Figure 7 An example driving environment 700 is illustrated. As shown, environment 700 includes vehicle 105, such as an autonomous vehicle, and vehicles 705, 710, and 715. As described herein, one or more vehicle sensors 115 generate sensor data representing the sensed driving environment. Based on the sensed driving environment, computer 110 is able to calculate a challenge score representing the difficulty of driving within the sensed driving environment. Based on the challenge score and user preferences, computer 110 selects and determines an RL agent 215 to perform one or more driving actions within environment 700.

[0099] Figure 8 This is a flowchart of an example process 800 for operating vehicle 105 according to the techniques described herein. The blocks of process 800 can be executed by computer 110. Process 800 begins at block 805, where sensed data representing the driving environment is received. At block 810, a challenge score is calculated based on the received data. At block 815, a user selection representing the desired driving style is received.

[0100] At block 820, an RL agent 215 is selected from P RL agents 215. At block 825, the RL agent 215 generates driving actions based on the sensed data. At block 830, a determination is made as to whether an updated user selection has been received. If an updated user selection has been received, process 800 returns to block 820. Otherwise, process 800 returns to block 825.

[0101] The description in this disclosure is exemplary in nature only, and variations thereof without departing from the spirit and scope of this disclosure are intended to be made within the scope of this disclosure. Such variations shall not be considered as departing from the spirit and scope of this disclosure.

[0102] Generally, the described computing system and / or device may employ any of a number of computer operating systems, including, but not limited to, the following versions and / or variations thereof: Microsoft Automotive® operating system, Microsoft Windows® operating system, Unix operating system (e.g., Solaris® operating system released by Oracle Corporation of Redwood Coast, California), AIX UNIX operating system released by International Business Machines Corporation of Armonk, New York, Linux operating system, Mac OSX and iOS operating systems released by Apple Inc. of Cupertino, California, BlackBerry OS released by BlackBerry Ltd. of Waterloo, Canada, and Android operating system developed by Google and the Open Handset Alliance, or the QNX® CAR entertainment system platform provided by QNX Software Systems. Examples of computing devices include, but are not limited to, in-vehicle computers, computer workstations, servers, desktop computers, laptops, notebooks, or handheld computers, or some other computing system and / or device.

[0103] Computers and computing devices typically include computer-executable instructions, which may be executable by one or more computing devices (such as those listed above). Computer-executable instructions can be compiled or interpreted from computer programs created using various programming languages ​​and / or technologies, including, but not limited to, Java™, C, C++, Matlab, Simulink, Stateflow, Visual Basic, JavaScript, Perl, HTML, etc. Some of these applications can be compiled and executed on virtual machines, such as the Java Virtual Machine, Dalvik Virtual Machine, etc. Generally, a processor (e.g., a microprocessor) receives and executes instructions, such as from memory, computer-readable media, etc., thereby performing one or more processes, including one or more of those described herein. Various computer-readable media can be used to store and transfer such instructions and other data. Files in a computing device are typically collections of data stored on computer-readable media, such as storage media, random access memory, etc.

[0104] Memory may include computer-readable media (also known as processor-readable media), which includes any non-transitory (e.g., tangible) medium involved in providing data (e.g., instructions) that can be read by a computer (e.g., by the computer's processor). Such media can take many forms, including but not limited to non-volatile and volatile media. Non-volatile media may include, for example, optical discs or magnetic disks, and other persistent storage. Volatile media may include, for example, dynamic random access memory (DRAM), which typically constitutes main memory. Such instructions may be transmitted via one or more transmission media, including coaxial cables, copper wires, and optical fibers, which include wires containing a system bus connected to a processor of an ECU. Common forms of computer-readable media include, for example, floppy disks, floppy disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs, any other optical media, punched cards, paper tape, any other physical media with a perforated pattern, RAM, PROM, EPROM, FLASH EEPROM, any other memory chip or cartridge, or any other medium that a computer can read from.

[0105] The databases, data repositories, or other data stores described herein can include a variety of mechanisms for storing, accessing, and retrieving various types of data, including hierarchical databases, a set of files in a file system, application databases in proprietary formats, relational database management systems (RDBMS), and so on. Each such data store is typically contained within a computing device employing a computer operating system (such as one of the computer operating systems mentioned above) and is accessed via a network in one or more of various ways. The file system can be accessible from the computer operating system and can include files stored in various formats. In addition to the languages ​​used to create, store, edit, and execute the stored procedures, RDBMS typically employs a structured query language (SQL), such as the PL / SQL language mentioned above.

[0106] In some examples, system elements may be implemented as computer-readable instructions (e.g., software) stored on one or more computing devices (e.g., servers, personal computers, etc.) on an associated computer-readable medium (e.g., disks, storage, etc.). Computer program products may include such instructions stored on computer-readable media to perform the functions described herein.

[0107] In this application (including the definitions below), the term "module" or "controller" may be replaced by the term "circuit". The term "module" may refer to, be part of, or include the following: application-specific integrated circuit (ASIC); digital, analog, or mixed-signal analog / digital discrete circuit; digital, analog, or mixed-signal analog / digital integrated circuit; combinational logic circuit; field-programmable gate array (FPGA); processor circuitry (shared, dedicated, or grouped) that executes code; memory circuitry (shared, dedicated, or grouped) that stores code executed by the processor circuitry; other suitable hardware components that provide the described functionality; or combinations of some or all of the above, such as in a system-on-a-chip.

[0108] A module may include one or more interface circuits. In some examples, the interface circuits may include wired or wireless interfaces that connect to a local area network (LAN), the Internet, a wide area network (WAN), or a combination thereof. The functionality of any given module of this disclosure may be distributed among multiple modules connected via the interface circuits. For example, multiple modules may allow for load balancing. In further examples, a server (also referred to as a remote or cloud) module may perform a function on behalf of a client module.

[0109] Regarding the media, processes, systems, methods, trial-and-error methods, etc., described herein, it should be understood that although the steps of such processes, etc., are described as occurring according to an ordered sequence, such processes can be practiced where the described steps are performed in an order different from that described herein. Furthermore, it should be understood that some steps may be performed simultaneously, other steps may be added, or some steps described herein may be omitted. In other words, the description of processes herein is provided for illustrative purposes and should in no way be construed as limiting the claims.

[0110] Therefore, it will be understood that the above description is intended to be illustrative rather than restrictive. Many embodiments and applications, other than the examples provided, will be apparent to those skilled in the art upon reading the above description. The scope of the invention should not be determined by reference to the above description, but rather by reference to the full scope of the appended claims together with their equivalents. It is contemplated and intended that future developments will occur in the techniques discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In conclusion, it should be understood that the invention is capable of modifications and variations and is limited only by the appended claims.

[0111] All terms used in the claims are intended to be given their simple and common meaning as understood by those skilled in the art, unless explicitly indicated otherwise herein. In particular, the use of singular articles (such as “a,” “the,” “said,” etc.) should be interpreted as referring to one or more of the indicated elements, unless the claims state an explicit limitation to the contrary.

Claims

1. An automated driving system for a vehicle, the automated driving system comprising a computer, the computer including a processor and a memory, the memory including instructions that program the processor to: Receive sensor data representing the perceived driving environment around the vehicle; A challenge score is calculated based on the sensor data. Receive the desired driving style from the user; Based on the calculated challenge score and the desired driving style, one reinforcement learning agent is selected from M×N reinforcement learning agents, where, M is an integer representing M driving preference levels, and N is an integer representing N numbers of driving environments; as well as Driving actions are generated based on the sensor data via a selected reinforcement learning agent. The M×N reinforcement learning agents are generated through the following steps, such that each of the M×N reinforcement learning agents corresponds to a different challenge score and a desired driving style: Train the reinforcement learning agent based on the baseline policy; Determine whether the average reward value has converged; If the average reward value has not converged, modify one or more hyperparameters and provide the modified one or more hyperparameters to the reinforcement learning agent, retrain the reinforcement learning agent and re-determine whether the average reward value has converged. If the average reward value converges, the reinforcement learning agent is replicated based on the baseline policy and / or the modified hyperparameters provided in the converged average reward value. A prescriptive analysis framework is used to generate N driving environments, each of which includes M driving preference levels; Calculate a descriptive challenge score based on each of the N driving environments; and The N driving environments are sorted according to the descriptive challenge scores, and the sorted N driving environments are provided to the reinforcement learning agents to generate the M×N reinforcement learning agents.

2. The automated driving system according to claim 1, wherein, The desired driving style corresponds to the desired level of driving aggression.

3. The automated driving system according to claim 2, wherein, The desired level of driving aggression corresponds to completing the driving action within a specific time period.

4. The automated driving system according to claim 1, wherein, The processor is further programmed to automatically select another reinforcement learning agent from the M×N reinforcement learning agents based on the sensor data representing different perceptions of the driving environment.

5. The automated driving system according to claim 1, wherein, The desired driving style is received from the human-machine interface.

6. The automated driving system according to claim 1, wherein, The vehicle is operated according to the driving actions described.

7. The automated driving system according to claim 6, wherein, The vehicle includes at least one of land vehicles, air vehicles, or water vehicles.

8. An automated driving method for a vehicle, the automated driving method comprising: Receive sensor data representing the perceived driving environment around the vehicle; A challenge score is calculated based on the sensor data. Receive the desired driving style from the user; Based on the calculated challenge score and the desired driving style, a reinforcement learning agent is selected from M×N reinforcement learning agents, where M is an integer representing M driving preference levels and N is an integer representing N numbers of driving environments; and Driving actions are generated based on the sensor data via a selected reinforcement learning agent. The M×N reinforcement learning agents are generated through the following steps, such that each of the M×N reinforcement learning agents corresponds to a different challenge score and a desired driving style: Train the reinforcement learning agent based on the baseline policy; Determine whether the average reward value has converged; If the average reward value has not converged, modify one or more hyperparameters and provide the modified one or more hyperparameters to the reinforcement learning agent, retrain the reinforcement learning agent and re-determine whether the average reward value has converged. If the average reward value converges, the reinforcement learning agent is replicated based on the baseline policy and / or the modified hyperparameters provided in the converged average reward value. A prescriptive analysis framework is used to generate N driving environments, each of which includes M driving preference levels; Calculate a descriptive challenge score based on each of the N driving environments; and The N driving environments are sorted according to the descriptive challenge scores, and the sorted N driving environments are provided to the reinforcement learning agents to generate the M×N reinforcement learning agents.

9. The automated driving method according to claim 8, wherein, The desired driving style corresponds to the desired level of driving aggression.

10. The automated driving method according to claim 9, wherein, The desired level of driving aggression corresponds to completing the driving action within a specific time period.

11. The automated driving method according to claim 8, further comprising: Based on the sensor data representing different perceptions of the driving environment, another reinforcement learning agent is automatically selected from the M×N reinforcement learning agents.

12. The automated driving method according to claim 8, wherein, The desired driving style is received from the human-machine interface.

13. The automated driving method according to claim 8, further comprising: Operate according to the driving actions described.

14. The automated driving method according to claim 13, wherein, The vehicle includes at least one of land vehicles, air vehicles, or water vehicles.

Citation Information

Patent Citations

  • Automatic driver agent and policy server for providing policies to driver agents

    CN110850854A

  • Controlling an autonomous vehicle using smart control architecture selection

    US20190113919A1

  • Cognitive state vehicle navigation based on image processing and modes

    US20210339759A1