Operation decision method and device applied to unmanned vehicle
Through reinforcement learning training decision model based on state space and action space, the problem of insufficient decision accuracy and safety of autonomous vehicles in complex scenarios is solved, and the operation safety and intelligence level of unmanned vehicles are improved.
Patent Information
- Application Number
- CN202210388019.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-13
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-04-13
AI Technical Summary
The accuracy and safety of existing autonomous decision-making methods of autonomous driving vehicles in complex scenarios are difficult to ensure, especially the difficulty in setting the critical value of regular decision-making methods, which affects driving performance.
By determining the state space and action space based on the state information and environmental information of the unmanned vehicle, using reinforcement learning to train the decision model, and optimizing the decision model with the preset return function to improve the decision accuracy.
It improves the operation safety and decision-making accuracy of unmanned vehicles in complex scenarios, and enhances the intelligence level of autonomous driving.
Smart Images

Figure CN114735027B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, in particular to the field of autonomous driving, and specifically to an operation decision-making method and device for an unmanned vehicle. Background Art
[0002] Autonomous vehicles, which can effectively improve driving performance, have seen rapid growth in recent years. Autonomous decision-making is one of the core technologies that enable these vehicles to achieve intelligent driving. Rule-based autonomous decision-making methods struggle to define critical values for behavior selection criteria, making them difficult to handle in complex scenarios. Furthermore, existing decision-making algorithms struggle to select critical values, severely impacting the accuracy and safety of rule-based decision-making methods. Summary of the Invention
[0003] The embodiments of the present application propose an operation decision-making method and device for an unmanned vehicle, as well as a decision-making method and device, a computer-readable medium, and an electronic device for an unmanned vehicle.
[0004] In the first aspect, an embodiment of the present application provides an operation decision-making method applied to an unmanned vehicle, including: determining the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located; determining the action space of the unmanned vehicle based on the parameters to be controlled during the operation of the unmanned vehicle; training an initial decision model for controlling the operation of the unmanned vehicle according to the state space, action space and preset reward function to obtain a trained decision model.
[0005] In some embodiments, the above-mentioned training of the initial decision model for controlling the operation of the unmanned vehicle according to the state space, action space and preset reward function to obtain the trained decision model includes: looping through the following training operations to obtain the decision model according to the obtained cumulative reward value: starting from the initialization state of the preset training task corresponding to the training operation, updating the initial decision model after training in the previous training operation according to the state space, action space and preset reward function until the preset end condition is reached, ending the current training operation, and determining the task reward value corresponding to the current training operation; determining the cumulative reward value up to the present according to the task reward values corresponding to each training operation that has been executed.
[0006] In some embodiments, the above starts from the initialization state of the preset training task, and updates the initial decision model after the previous training operation according to the state space, action space and preset reward function, including: looping the following update operations: determining whether the preset end condition is met according to the state space at the current moment; in response to determining that the preset end condition is not met, inputting the state space at the current moment into the preset reward function to determine the single-step reward value; updating the initial decision model according to the single-step reward value; outputting the action space at the next moment according to the updated initial decision model to control the simulation operation of the unmanned vehicle and obtain the state space corresponding to the next moment.
[0007] In some embodiments, the preset end conditions include the unmanned vehicle completing a preset training task and the unmanned vehicle colliding; and the above-mentioned determination of the task reward value corresponding to the current training operation includes: in response to determining that the unmanned vehicle completes the preset training task according to the state space corresponding to the current moment, determining a first task reward value, wherein the first task reward value is a positive value; in response to determining that the unmanned vehicle collides according to the state space corresponding to the current moment, determining a second task reward value, wherein the second task reward value is a negative value.
[0008] In some embodiments, in response to determining that the preset end condition has not been met, the state space at the current moment is input into the preset reward function to determine the single-step reward value, including: in response to determining that the preset end condition has not been met, the state space at the current moment is input into the preset reward function, and the single-step reward value is determined based on the information in the state space that represents the driving safety, task completion, driving efficiency and driving comfort up to the current moment.
[0009] In some embodiments, the above-mentioned determination of the single-step reward value based on the information in the state space representing the driving safety, task completion, driving efficiency and driving comfort up to the current time includes: determining a first single-step reward sub-value corresponding to driving safety based on the information in the state space representing the distance of the unmanned vehicle relative to the centerline of the lane in which it is located and the closest distance between the unmanned vehicle and surrounding vehicles; determining a second single-step reward sub-value corresponding to task completion based on the information in the state space representing the distance between the unmanned vehicle and the target lane, the angle of the unmanned vehicle relative to the centerline of the lane in which it is located, and the position of the unmanned vehicle; determining a third single-step reward sub-value corresponding to driving efficiency based on the information in the state space representing the speed of the unmanned vehicle; determining a fourth single-step reward sub-value corresponding to driving comfort based on the information in the state space representing the yaw angular velocity, steering wheel angle and longitudinal acceleration of the unmanned vehicle; and determining the single-step reward value based on the first single-step reward sub-value, the second single-step reward sub-value, the third single-step reward sub-value and the fourth single-step reward sub-value.
[0010] In some embodiments, the above-mentioned determination of the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located includes: determining the state space based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the surrounding vehicles of the unmanned vehicle, and the relative relationship between the unmanned vehicle and the surrounding vehicles.
[0011] In some embodiments, the above-mentioned determination of the state space based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the unmanned vehicle's surrounding vehicles, and the relative relationship between the unmanned vehicle and the surrounding vehicles includes: determining the initial state space of the unmanned vehicle based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the unmanned vehicle's surrounding vehicles, and the relative relationship between the unmanned vehicle and the surrounding vehicles; normalizing each parameter in the initial state space to obtain the state space.
[0012] In a second aspect, an embodiment of the present application provides a decision-making method applied to an unmanned vehicle, comprising: obtaining a pre-trained decision model, wherein the decision model is obtained by the method described in any implementation of the first aspect.
[0013] On the third aspect, an embodiment of the present application provides an operation decision-making device for an unmanned vehicle, including: a first determination unit, configured to determine the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located; a second determination unit, configured to determine the action space of the unmanned vehicle based on the parameters to be controlled during the operation of the unmanned vehicle; a training unit, configured to train an initial decision model for controlling the operation of the unmanned vehicle according to the state space, action space and preset reward function, to obtain a trained decision model.
[0014] In some embodiments, the above-mentioned training unit is further configured to: loop through the following training operations to obtain a decision model based on the cumulative reward value obtained: starting from the initialization state of the preset training task corresponding to the training operation, update the initial decision model after the previous training operation according to the state space, action space and preset reward function until the preset end condition is reached, end the current training operation, and determine the task reward value corresponding to the current training operation; determine the cumulative reward value up to the current time based on the task reward values corresponding to each training operation that has been executed.
[0015] In some embodiments, the above-mentioned training unit is further configured to: loop and perform the following update operations: determine whether the preset end condition is reached according to the state space at the current moment; in response to determining that the preset end condition is not reached, input the state space at the current moment into the preset reward function to determine the single-step reward value; update the initial decision model according to the single-step reward value; output the action space at the next moment according to the updated initial decision model to control the simulation operation of the unmanned vehicle and obtain the state space corresponding to the next moment.
[0016] In some embodiments, the preset end conditions include the unmanned vehicle completing a preset training task and the unmanned vehicle colliding; and the above-mentioned training unit is further configured to: in response to determining that the unmanned vehicle completes the preset training task according to the state space corresponding to the current moment, determine a first task reward value, wherein the first task reward value is a positive value; in response to determining that the unmanned vehicle collides according to the state space corresponding to the current moment, determine a second task reward value, wherein the second task reward value is a negative value.
[0017] In some embodiments, the above-mentioned training unit is further configured to: in response to determining that the preset end condition has not been met, input the state space at the current moment into a preset reward function, and determine the single-step reward value based on the information in the state space that represents the driving safety, task completion, driving efficiency and driving comfort up to the current moment.
[0018] In some embodiments, the above-mentioned training unit is further configured to: determine a first single-step reward sub-value corresponding to driving safety based on information in the state space representing the distance between the unmanned vehicle and the center line of the lane in which it is located, and the closest distance between the unmanned vehicle and surrounding vehicles; determine a second single-step reward sub-value corresponding to task completion based on information in the state space representing the distance between the unmanned vehicle and the target lane, the angle between the unmanned vehicle and the center line of the lane in which it is located, and the position of the unmanned vehicle; determine a third single-step reward sub-value corresponding to driving efficiency based on information in the state space representing the speed of the unmanned vehicle; determine a fourth single-step reward sub-value corresponding to driving comfort based on information in the state space representing the yaw angular velocity, steering wheel angle and longitudinal acceleration of the unmanned vehicle; determine a single-step reward value based on the first single-step reward sub-value, the second single-step reward sub-value, the third single-step reward sub-value and the fourth single-step reward sub-value.
[0019] In some embodiments, the above-mentioned first determination unit is further configured to: determine the state space based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the unmanned vehicle's surrounding vehicles, and the relative relationship between the unmanned vehicle and the surrounding vehicles.
[0020] In some embodiments, the above-mentioned first determination unit is further configured to: determine the initial state space of the unmanned vehicle based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the unmanned vehicle's surrounding vehicles, and the relative relationship between the unmanned vehicle and the surrounding vehicles; normalize each parameter in the initial state space to obtain the state space.
[0021] In a fourth aspect, an embodiment of the present application provides a decision-making device applied to an unmanned vehicle, comprising: an acquisition unit configured to acquire a pre-trained decision model, wherein the decision model is obtained by the device described in any implementation method of the third aspect.
[0022] In a fifth aspect, an embodiment of the present application provides a computer-readable medium on which a computer program is stored, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect or the second aspect is implemented.
[0023] In the sixth aspect, an embodiment of the present application provides an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation method of the first aspect or the second aspect.
[0024] The operation decision-making method and device for unmanned vehicles provided in the embodiments of the present application determine the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located; determine the action space of the unmanned vehicle based on the parameters to be controlled during the operation of the unmanned vehicle; train an initial decision model for controlling the operation of the unmanned vehicle based on the state space, action space and preset reward function to obtain a trained decision model, thereby providing an operation decision-making method for unmanned vehicles based on reinforcement learning, and obtaining a trained decision model based on the established state space, action space and reward function of the unmanned vehicle, thereby improving the accuracy of the decision model and thus improving the operation safety of the unmanned vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0026] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present application may be applied;
[0027] Figure 2 is a flowchart of an embodiment of an operation decision-making method applied to an unmanned vehicle according to the present application;
[0028] FIG3 is a schematic diagram of a ramp merging scenario according to this embodiment;
[0029] Figure 4 is a specific schematic diagram of the training task according to this embodiment;
[0030] Figure 5 is a schematic diagram illustrating various parameters in the state space according to this embodiment;
[0031] Figure 6 is a schematic diagram of surrounding vehicles of the unmanned vehicle in a ramp merging scenario according to this embodiment;
[0032] Figure 7 is a schematic diagram of the relative relationship between the unmanned vehicle and surrounding vehicles according to this embodiment;
[0033] Figure 8 is a schematic diagram of an application scenario of the operation decision-making method applied to an unmanned vehicle according to this embodiment;
[0034] Figure 9 is a flowchart of another embodiment of the operation decision-making method applied to an unmanned vehicle according to the present application;
[0035] Figure 10 is a flowchart of an embodiment of a decision-making method applied to an unmanned vehicle according to the present application;
[0036] Figure 11 This is a structural diagram of an embodiment of an operation decision-making device applied to an unmanned vehicle according to the present application;
[0037] Figure 12 is a structural diagram of an embodiment of a decision-making device applied to an unmanned vehicle according to the present application;
[0038] Figure 13 It is a structural diagram of a computer system suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION
[0039] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.
[0040] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0041] Figure 1 An exemplary architecture 100 is shown to which the operation decision-making method and apparatus for an unmanned vehicle and the decision-making method and apparatus for an unmanned vehicle of the present application can be applied.
[0042] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The communication connections between terminal devices 101, 102, and 103 constitute a topological network, and network 104 is used to provide a medium for communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0043] The unmanned vehicle can use terminal devices 101, 102, and 103 to interact with the server 105 via the network 104 to receive or send messages, etc. The terminal devices 101, 102, and 103 can be hardware devices or software that support network connection for data interaction and data processing. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connection, information acquisition, interaction, display, processing, and other functions, including but not limited to vehicle-mounted control devices, tablet computers, e-book readers, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules for providing distributed services, for example, or as a single software or software module. No specific limitations are given here.
[0044] Server 105 can be a server that provides various services. For example, based on training instructions for the decision model sent by a terminal device, it can train an initial decision model for controlling the operation of the unmanned vehicle based on the state space, action space, and preset reward function designed for the unmanned vehicle, thereby obtaining a trained decision model. After obtaining the decision model, the operation of the unmanned vehicle can be controlled by the decision model. As an example, server 105 can be a cloud server.
[0045] It should be noted that the server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (e.g., software or software modules for providing distributed services), or as a single software or software module. No specific limitations are given here.
[0046] It should also be noted that the operation decision-making method and decision-making method for unmanned vehicles provided in the embodiments of the present application can be executed by a server, a terminal device, or a server and a terminal device in cooperation with each other. Accordingly, the various parts (e.g., various units) included in the operation decision-making device and decision-making device for unmanned vehicles can be all set in the server, all set in the terminal device, or separately set in the server and the terminal device.
[0047] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the system is merely illustrative. Any number of terminal devices, networks, and servers may be provided as required. When the operation decision-making method applied to the unmanned vehicle and the electronic device on which the decision-making method applied to the unmanned vehicle is executed do not need to transmit data with other electronic devices, the system architecture may only include the operation decision-making method applied to the unmanned vehicle and the electronic device (e.g., server or terminal device) on which the decision-making method applied to the unmanned vehicle is executed.
[0048] Continue to refer Figure 2 , shows a process 200 of an embodiment of an operation decision method applied to an unmanned vehicle, including the following steps:
[0049] Step 201: Determine the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located.
[0050] In this embodiment, the execution subject (e.g. Figure 1 The terminal device or server in the system determines the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located.
[0051] Unmanned vehicles can be driverless vehicles, assisted-driving vehicles, and other intelligent vehicles that utilize computer technology to achieve unmanned or assisted operation. In this embodiment, the ultimately trained decision model used to control the operation of the target vehicle can be applied to any vehicle operation scenario, including but not limited to merging onto a highway ramp, merging onto a non-highway auxiliary road, and operating on a main road.
[0052] This embodiment aims to represent the real-time state of the vehicle during operation through the determined state space. The state space includes the state information of the unmanned vehicle and the environmental information of the unmanned vehicle's environment. It should be noted that a unified state space can be designed for each operating environment for the unmanned vehicle in different operating scenarios; different state spaces can also be designed based on the actual operating conditions of different operating scenarios.
[0053] In view of the complexity of ramp merging scenarios, taking ramp merging scenarios as an example, there are three types of existing ramp merging scenarios: the first type is a single-lane ramp with a known merging point, such as Figure 3A The second type is a single lane ramp with unknown merging point, such as Figure 3B As shown; the third type is a composite multi-lane ramp, which is a combination of the first two types, such as Figure 3C As shown in Figure 2, this type of scenario is the most common. Here, we continue to design training tasks for the decision model using the second type of unknown merging ramp as a typical scenario to clarify the training tasks corresponding to this typical scenario.
[0054] like Figure 4 As shown in the figure, a specific schematic diagram of the training task is shown. The main road is a three-lane highway with a length of 350 meters and a single lane width of 3.75 meters. The ramp merges into the highway at an angle of 10° at 100 meters. There is a 130-meter-long acceleration lane and a 90-meter-long gradient lane. The initial position of the unmanned vehicle is in the ramp. When the lateral position is greater than 100 meters, the unmanned vehicle merges into the main line and maintains an angle of 3° with the lane line without colliding with surrounding vehicles and road boundaries. ° If the vehicle remains within the specified range for more than 3 seconds, it is considered that the vehicle has successfully merged. In other words, the vehicle can have multiple target locations that indicate the vehicle has successfully merged. It should be noted that the specific scenario information shown in the figure is only an example and does not limit the specific scenario.
[0055] The status information of the unmanned vehicle may include, for example, vehicle speed, position, and other information, and the environmental information may include, for example, the relative relationship between the unmanned vehicle and target objects in the surrounding environment.
[0056] In some optional implementations of this embodiment, the execution entity may perform step 201 as follows:
[0057] The state space is determined based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the vehicles surrounding the unmanned vehicle, and the relative relationship between the unmanned vehicle and the surrounding vehicles.
[0058] As an example, in the ramp merging scenario, the state information of the autonomous vehicle s ego The expression is:
[0059]
[0060] Among them, such as Figure 5 The state space diagram 500 of the unmanned vehicle shown in FIG. e represents the speed of the unmanned vehicle, w e 、l e Represent the width and length of the unmanned vehicle respectively. is the angle between the driverless car and the center line of the lane it is in, d c is the lateral distance of the autonomous vehicle relative to the centerline of the lane it is in, d l d r are the lateral distances of the autonomous vehicle relative to the left and right boundaries of the lane, d m It is the lateral distance between the unmanned vehicle and the center line of the outer lane of the highway.
[0061] like Figure 6 As shown, a schematic diagram 600 of the surrounding vehicles of the unmanned vehicle in the ramp merging scene is shown, including the schematic situation of the unmanned vehicle and the surrounding vehicles when it is on the ramp, and the schematic situation of the unmanned vehicle and the surrounding vehicles when it is on the main line (such as the main road of the highway). Since the unmanned vehicle is generally only affected by the adjacent surrounding vehicles during operation, the status information of the surrounding vehicles and the relative relationship between the surrounding vehicles and the unmanned vehicle s veh Only consider the area within the sensor detection range (before the unmanned vehicle d f m, then d b Specifically, for the lane where the autonomous vehicle is located, the two closest vehicles, veh1 and veh2, are considered; for the left and right lanes adjacent to the autonomous vehicle, the closest vehicle behind it, veh3 and veh4, as well as the two closest or parallel vehicles in front of it, veh5, veh6, veh7, and veh8, are considered, for a total of up to eight surrounding vehicles. Specifically, s veh Expressed as:
[0062] s veh =[veh1,…,veh j ,…,veh8]
[0063] The veh j include:
[0064]
[0065] d xj =x j -x e
[0066] d yj =y j -y e
[0067] Among them, such as Figure 7 The relative relationship diagram 700 of the unmanned vehicle and surrounding vehicles is shown. j Indicates veh j speed, It's a veh j The heading angle, w j 、lj Respectively represent veh j Width and length, d xj d yj Indicates veh j The horizontal and vertical distance from the unmanned vehicle, x j 、y j Indicates veh j Location information, x e 、y e Indicates the location information of the unmanned vehicle.
[0068] It is worth noting that in actual road scenarios, when an unmanned vehicle is driving on a ramp, it may not be able to observe vehicles in the outer lanes of the main line due to obstructions by trees and buildings. However, in this embodiment, in order to make better decisions, it is assumed that they are observable within the sensor detection range.
[0069] In some optional implementations of this embodiment, the above-mentioned execution entity can obtain the state space specifically in the following manner: first, based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the unmanned vehicle's surrounding vehicles, and the relative relationship between the unmanned vehicle and the surrounding vehicles, determine the initial state space of the unmanned vehicle; then, normalize the parameters in the initial state space to obtain the state space.
[0070] In this embodiment, each parameter in the initial state space is normalized to between 0 and 1 to obtain a normalized state space, so as to improve the information processing efficiency and training effect during the entire training process.
[0071] Step 202: Determine the action space of the unmanned vehicle based on the parameters to be controlled during the operation of the unmanned vehicle.
[0072] In this embodiment, the execution entity can determine the autonomous vehicle's action space based on the parameters to be controlled during the autonomous vehicle's operation. The action space is used to represent the control information for the autonomous vehicle. Each updated action space is used to control the autonomous vehicle's next operational action to be adjusted.
[0073] As an example, the parameters to be controlled when the unmanned vehicle is running include acceleration and front wheel angle, so the action space a can be expressed as
[0074] a=[a x ,δ]
[0075] Among them, a x represents acceleration, and δ represents the front wheel angle.
[0076] Step 203 : Based on the state space, the action space, and the preset reward function, an initial decision model for controlling the operation of the unmanned vehicle is trained to obtain a trained decision model.
[0077] In this embodiment, the above-mentioned execution entity can train an initial decision model for controlling the operation of the unmanned vehicle based on the state space, action space and preset reward function to obtain a trained decision model.
[0078] The preset reward function is an extremely important part of reinforcement learning. By making the training task objectives specific and numerical, it reflects the goal of machine learning. Reinforcement learning is a machine that achieves its own goals by maximizing the cumulative rewards of the reward function.
[0079] As an example, the execution entity can be based on a simulation operation system for an unmanned vehicle. Starting from the initialization state of the training task, the current state space of the unmanned vehicle is input into a preset reward function to obtain a reward value. The initialization model updates parameters based on the reward value and outputs the action space to be executed by the unmanned vehicle. This allows the simulation operation system to simulate the unmanned vehicle according to the action space, obtain a changed state space, and further determine the reward value based on the changed state space. During the process of looping and executing the above operations, the initial decision model is updated with the goal of maximizing the cumulative reward value of each reward value obtained, resulting in a trained decision model.
[0080] In some optional implementations of this embodiment, the execution entity may perform step 203 as follows:
[0081] The following training operations are performed cyclically to obtain a decision model based on the accumulated reward values:
[0082] First, starting from the initialization state of the preset training task corresponding to the training operation, the initial decision model after the previous training operation is updated according to the state space, action space and preset reward function until the preset end condition is reached, the current training operation is ended, and the task reward value corresponding to the current training operation is determined.
[0083] If the current training operation is the first training operation and no training operation has been performed before this training operation, the model to be updated by this training operation is the initial decision model for training.
[0084] A pre-set training task can be a task designed to enable the decision model to learn a desired function. For example, if the decision model is designed to enable the autonomous vehicle to merge onto a ramp, the pre-set training task could be a task that demonstrates the autonomous vehicle's ability to reach a target destination on the main road from a target starting point on the ramp in a ramp-merge scenario.
[0085] It should be noted that the preset training tasks corresponding to each training operation can be the same or different. Continuing with the example of a model designed to achieve ramp-merge functionality, the preset training task for each training operation could be a training task that demonstrates the autonomous vehicle's ability to reach its target destination on the main road from its target starting point on the ramp in a ramp-merge scenario. When the preset training tasks for each training operation are different, decision-making models with a variety of desired capabilities can be obtained.
[0086] Each training operation begins with the initialization state of the preset training task. For example, the initialization state represents the unmanned vehicle in the simulation operation system at the target starting point on the ramp. When the preset end condition is obtained, the training operation ends. The preset end condition can be set according to the specific preset training task. Continuing with the ramp merging scenario as an example, the preset end condition includes the successful completion of the preset training task, that is, according to the instructions of the action space given by the decision model, the unmanned vehicle successfully reaches the target end point on the main road from the target starting point on the ramp; the preset end condition also includes: unsuccessful completion of the preset training task.
[0087] In response to reaching a preset end condition, the execution entity may determine a specific task reward value using a preset reward function based on which preset end condition is reached. The task reward value represents the reward value corresponding to each preset training task after the task is completed.
[0088] Second, the cumulative reward value up to the current time is determined based on the task reward value corresponding to each training operation that has been executed.
[0089] As an example, the execution entity may accumulate the task reward values corresponding to each executed training operation to determine the current cumulative reward value.
[0090] In multiple training operations, the initial decision model is updated with the goal of maximizing the cumulative reward value to obtain the trained decision model.
[0091] In this embodiment, by cyclically executing the preset training tasks, the initial decision model can more fully learn how to execute the preset training tasks, thereby improving the accuracy of the obtained decision model.
[0092] In some optional implementations of this embodiment, the execution entity may update the initial decision model trained in the previous training operation according to the state space, the action space, and the preset reward function by performing the following operations:
[0093] The following update operations are performed in a loop:
[0094] First, based on the current state space, a determination is made as to whether a preset termination condition has been met. Then, if the preset termination condition has not been met, the current state space is input into a preset reward function to determine a single-step reward value. The initial decision model is then updated based on the single-step reward value. Finally, based on the updated initial decision model, the action space for the next moment is output to control the unmanned vehicle simulation and obtain the corresponding state space for the next moment.
[0095] Based on the parameters in the state space, it can be determined whether a preset end condition has been met. For example, based on the information representing the position of the unmanned vehicle in the state space, it can be determined whether the unmanned vehicle has reached the target destination and whether the preset end condition of completing the training task has been met.
[0096] In each update operation of the same training operation, the current initial decision model can be continuously updated with the goal of maximizing the accumulated single-step reward value up to the current point until the preset end condition is reached.
[0097] In this implementation, in the training operation corresponding to the same preset training task, the initial decision model is further adjusted through the change information at each moment when the unmanned vehicle performs the preset training task, thereby further improving the accuracy of the trained decision model.
[0098] In some optional implementations of this embodiment, the preset termination conditions include the unmanned vehicle completing a preset training task or the unmanned vehicle experiencing a collision. Completion of the preset training task indicates that the unmanned vehicle has successfully reached the target destination corresponding to the task. A collision may include, for example, a collision with a road edge, roadside structure, or a surrounding vehicle.
[0099] In this implementation, the execution entity can determine the task reward value corresponding to the current training operation by executing the following method:
[0100] In response to determining that the unmanned vehicle completes a preset training task according to a state space corresponding to the current moment, a first task reward value is determined.
[0101] The first task reward value is a positive value. The value of the first task reward value can be set according to actual conditions and is not limited here.
[0102] In response to determining that the unmanned vehicle has collided according to the state space corresponding to the current moment, a second task reward value is determined.
[0103] The reward value of the second task is negative. The value of the reward value of the first task can be set according to the actual situation and is not limited here.
[0104] In some optional implementations of this embodiment, the above-mentioned execution entity can determine the single-step reward value in the following manner: in response to determining that the preset end condition has not been met, the state space at the current moment is input into the preset reward function, and the single-step reward value is determined based on the information in the state space that represents the driving safety, task completion, driving efficiency and driving comfort up to the current moment.
[0105] Specifically, the above-mentioned execution entity can pre-determine the parameters corresponding to various indicators such as driving safety, task completion, driving efficiency, and driving comfort in the state space, and determine the single-step reward value based on the corresponding parameters through a preset reward function.
[0106] In some optional implementations of this embodiment, the execution entity determines the single-step reward value in the following manner:
[0107] First, the first single-step reward sub-value corresponding to driving safety is determined based on the information in the state space that represents the distance of the unmanned vehicle relative to the center line of the lane it is in and the closest distance between the unmanned vehicle and surrounding vehicles.
[0108] As an example, the execution entity may determine the first single-step return sub-value in the following manner:
[0109]
[0110]
[0111] Among them, d v is the shortest distance between the unmanned vehicle and surrounding vehicles, d safe Indicates the preset safety distance, such as 10 meters; d c Indicates the distance of the autonomous vehicle relative to the center line of the lane it is in. k s1 、k s2 are the corresponding weights, which are all negative numbers, indicating that if the autonomous vehicle maintains a safe distance from surrounding vehicles, the autonomous vehicle will receive a positive reward, otherwise, the autonomous vehicle will receive a penalty; the more the autonomous vehicle deviates from the lane line, the more unsafe it is, and the greater the penalty it receives. xj d yj Indicates veh j The lateral and longitudinal distances from the unmanned vehicle, m represents the minimum value when j is 1 to 8.
[0112] Second, the second single-step reward sub-value corresponding to the task completion degree is determined based on the information in the state space that represents the distance between the unmanned vehicle and the target lane, the angle between the unmanned vehicle and the center line of the lane, and the position of the unmanned vehicle.
[0113] As an example, the execution entity may determine the second single-step return sub-value in the following manner:
[0114]
[0115] Among them, d m represents the distance between the unmanned vehicle and the target lane, represents the angle between the autonomous vehicle and the center line of the lane where it is located, x e It represents the distance between the current position of the unmanned vehicle and the target destination, which can be determined based on the real-time position of the unmanned vehicle and the location of the target destination. t1 ,k t2 ,k t3 are the corresponding weights, where k t1 ,k t2 are all negative, k t3 It is a positive number.
[0116] In the ramp-in scenario, the target lane is the lane of the main line after merging. The greater the lateral distance from the target lane and the greater the relative heading angle, the greater the penalty for the task item. When the autonomous vehicle successfully merges into the highway, the penalty is close to 0. In order to prevent dense traffic from flowing, the autonomous vehicle waits in place to merge into the main line, through k t2 Encourage driverless cars to move forward.
[0117] Third, based on the information representing the speed of the unmanned vehicle in the state space, the third single-step reward sub-value corresponding to the driving efficiency is determined.
[0118] As an example, the execution entity may determine the third single-step return sub-value in the following manner:
[0119] r efficiency =k e (v max -v e )
[0120] Among them, v max Indicates the maximum speed allowed during the operation of the unmanned vehicle, v e k represents the speed of the unmanned vehicle. e Indicates the corresponding weight coefficient, which is a positive number, indicating that the unmanned vehicle will be penalized if its speed is less than the maximum speed. e It is also the weight of efficiency. The smaller the absolute value, the smaller the weight.
[0121] Fourth, based on the information representing the yaw angular velocity, steering wheel angle, and longitudinal acceleration of the unmanned vehicle in the state space, the fourth single-step report sub-value corresponding to the driving comfort is determined.
[0122] As an example, the execution entity may determine the fourth single-step return sub-value in the following manner:
[0123]
[0124] in, δ sw 、a x They represent the yaw rate, steering wheel angle and longitudinal acceleration of the autonomous vehicle, respectively, and k c1 、k c2 、k c3 are the corresponding weights, both are negative.
[0125] Fifth, determine the single-step return value based on the first single-step return sub-value, the second single-step return sub-value, the third single-step return sub-value, and the fourth single-step return sub-value.
[0126] As an example, the execution entity may add the first single-step return sub-value, the second single-step return sub-value, the third single-step return sub-value, and the fourth single-step return sub-value to obtain a single-step return value.
[0127] Based on the above, the preset reward function can be expressed as a whole by the following formula:
[0128]
[0129] The weights of each item in the preset reward function require repeated experimentation to determine the weights that result in good decision-making performance. Understandably, task completion and driving safety are weighted more heavily, while driving efficiency and comfort are weighted less heavily.
[0130] As an example, the weights of each item in the preset reward function are shown in the following table:
[0131] <![CDATA[r succeed ]]> <![CDATA[r f ]]> <![CDATA[k s1 ,k s2 ]]> <![CDATA[k t1 ,k t2 ,k t3 ]]> <![CDATA[k e ]]> <![CDATA[k c1 ,k c2 ,k c3 ]]> -200 1000 -3,-2 -1,-20,20 0.5 -15,-15,-15
[0132] Continue to see Figure 8 , Figure 8 FIG8 is a schematic diagram 800 of an application scenario of the operation decision method for an unmanned vehicle according to this embodiment. Figure 8 In an application scenario, in a simulated vehicle operation environment, unmanned vehicle 801 merges onto highway 803 from ramp 802. Using this simulated operation environment, a decision model is trained to ensure that the unmanned vehicle successfully merges onto the main line. First, the state space 804 of the unmanned vehicle is determined based on its state information and environmental information. Then, the action space 805 of the unmanned vehicle is determined based on the parameters to be controlled during operation. Finally, based on the state space 804, action space 805, and preset reward function 806, an initial decision model for controlling the operation of unmanned vehicle 801 is trained, resulting in a trained decision model 807.
[0133] The method provided by the above-mentioned embodiments of the present application determines the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located; determines the action space of the unmanned vehicle based on the parameters to be controlled during the operation of the unmanned vehicle; and trains the initial decision model for controlling the operation of the unmanned vehicle according to the state space, action space and preset reward function to obtain a trained decision model, thereby providing an operation decision method for unmanned vehicles based on reinforcement learning. Based on the established state space, action space and reward function of the unmanned vehicle, a trained decision model is obtained, which improves the accuracy of the decision model and thus improves the operation safety.
[0134] Continue to refer Figure 9 , shows a schematic process 900 of an embodiment of an operation decision method applied to an unmanned vehicle according to the present application, including the following steps:
[0135] Step 901: Determine the state space of the unmanned vehicle based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the vehicles surrounding the unmanned vehicle, and the relative relationship between the unmanned vehicle and the surrounding vehicles.
[0136] Step 902: Determine the action space of the unmanned vehicle based on the parameters to be controlled during the operation of the unmanned vehicle.
[0137] Step 903: cyclically perform the following training operations to obtain a decision model based on the obtained cumulative reward values:
[0138] Step 9031: Starting from the initialization state of the preset training task corresponding to the training operation, the following update operations are executed cyclically:
[0139] Step 90311, based on the current state space, determine whether the preset end condition is met.
[0140] Step 90312, in response to determining that the preset end condition has not been met, the state space at the current moment is input into the preset reward function, and the single-step reward value is determined based on the information in the state space that represents the driving safety, task completion, driving efficiency and driving comfort up to the current moment.
[0141] Step 90313: Update the initial decision model based on the single-step reward value.
[0142] Step 90314: Output the action space at the next moment based on the updated initial decision model to control the simulation operation of the unmanned vehicle and obtain the state space corresponding to the next moment.
[0143] Step 90315, in response to determining that the preset end condition is that the unmanned vehicle completes the preset training task based on the state space corresponding to the current moment, determine the first task reward value.
[0144] Among them, the reward value of the first task is positive.
[0145] Step 90316, in response to determining that the preset end condition is a collision of the unmanned vehicle according to the state space corresponding to the current moment, determine the second task reward value.
[0146] Among them, the reward value of the second task is negative.
[0147] Step 9032: Determine the task reward value corresponding to the current training operation; and determine the current cumulative reward value based on the task reward values corresponding to each executed training operation.
[0148] It can be seen from this embodiment that Figure 2 Compared with the corresponding embodiments, process 900 of the operation decision method applied to the unmanned vehicle in this embodiment specifically illustrates the state space determination process and the decision model training process, further improving the accuracy of the obtained decision model and the operation safety of the unmanned vehicle.
[0149] Continue to refer Figure 10 , shows a process 1000 of an embodiment of a decision-making method applied to an unmanned vehicle, including the following steps:
[0150] Step 1001: Obtain a pre-trained decision model.
[0151] In this embodiment, the execution subject of the decision-making method applied to the unmanned vehicle (for example, Figure 1 The server or terminal device in the network can obtain the pre-trained decision model.
[0152] The decision model, obtained through any of the implementations described in Example 200, represents the correspondence between the autonomous vehicle's current state space and its next action space. The state space includes the autonomous vehicle's state information, the relative relationship between the autonomous vehicle and the road, the state information of the autonomous vehicle's surrounding vehicles, and the relative relationship between the autonomous vehicle and its surrounding vehicles. The action space is used to represent the autonomous vehicle's expected next operating parameters, such as acceleration and front wheel steering angle.
[0153] Step 1002: Control the operation of the unmanned vehicle through the decision model.
[0154] In this embodiment, the above-mentioned execution entity can control the operation of the unmanned vehicle through a decision-making model.
[0155] In this embodiment, the operation of the unmanned vehicle is controlled by a decision model, thereby improving the operational safety of the unmanned vehicle.
[0156] Continue to refer Figure 11As an implementation of the methods shown in the above figures, the present application provides an embodiment of an operation decision-making device for an unmanned vehicle. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0157] like Figure 11 As shown, the operation decision-making device applied to the unmanned vehicle includes: a first determination unit 1101, which is configured to determine the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located; a second determination unit 1102, which is configured to determine the action space of the unmanned vehicle based on the parameters to be controlled when the unmanned vehicle is running; a training unit 1103, which is configured to train the initial decision model for controlling the operation of the unmanned vehicle according to the state space, the action space and the preset reward function to obtain a trained decision model.
[0158] In some optional implementations of this embodiment, the above-mentioned training unit 1103 is further configured to: cyclically execute the following training operations to obtain a decision model based on the cumulative reward value obtained: starting from the initialization state of the preset training task corresponding to the training operation, according to the state space, action space and preset reward function, update the initial decision model after training of the previous training operation until the preset end condition is reached, end the current training operation, and determine the task reward value corresponding to the current training operation; determine the cumulative reward value up to the present based on the task reward values corresponding to each training operation that has been executed.
[0159] In some optional implementations of this embodiment, the above-mentioned training unit 1103 is further configured to: cyclically perform the following update operations: determine whether the preset end condition is met based on the state space at the current moment; in response to determining that the preset end condition is not met, input the state space at the current moment into the preset reward function to determine the single-step reward value; update the initial decision model based on the single-step reward value; output the action space at the next moment based on the updated initial decision model to control the simulation operation of the unmanned vehicle and obtain the state space corresponding to the next moment.
[0160] In some optional implementations of this embodiment, the preset end conditions include the unmanned vehicle completing the preset training task and the unmanned vehicle colliding; and the above-mentioned training unit 1103 is further configured to: in response to determining that the unmanned vehicle completes the preset training task according to the state space corresponding to the current moment, determine a first task reward value, wherein the first task reward value is a positive value; in response to determining that the unmanned vehicle collides according to the state space corresponding to the current moment, determine a second task reward value, wherein the second task reward value is a negative value.
[0161] In some optional implementations of this embodiment, the above-mentioned training unit 1103 is further configured to: in response to determining that the preset end condition has not been met, input the state space at the current moment into the preset reward function, and determine the single-step reward value based on the information in the state space that represents the driving safety, task completion, driving efficiency and driving comfort up to the current moment.
[0162] In some optional implementations of this embodiment, the above-mentioned training unit 1103 is further configured to: determine a first single-step reward sub-value corresponding to driving safety based on information representing the distance of the unmanned vehicle relative to the center line of the lane in which it is located and the closest distance between the unmanned vehicle and surrounding vehicles in the state space; determine a second single-step reward sub-value corresponding to task completion based on information representing the distance between the unmanned vehicle and the target lane, the angle of the unmanned vehicle relative to the center line of the lane in which it is located, and the position of the unmanned vehicle in the state space; determine a third single-step reward sub-value corresponding to driving efficiency based on information representing the speed of the unmanned vehicle in the state space; determine a fourth single-step reward sub-value corresponding to driving comfort based on information representing the yaw angular velocity, steering wheel angle and longitudinal acceleration of the unmanned vehicle in the state space; and determine a single-step reward value based on the first single-step reward sub-value, the second single-step reward sub-value, the third single-step reward sub-value and the fourth single-step reward sub-value.
[0163] In some optional implementations of this embodiment, the above-mentioned first determination unit 1101 is further configured to: determine the state space based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the unmanned vehicle's surrounding vehicles, and the relative relationship between the unmanned vehicle and the surrounding vehicles.
[0164] In some optional implementations of this embodiment, the above-mentioned first determination unit 1101 is further configured to: determine the initial state space of the unmanned vehicle based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of the unmanned vehicle's surrounding vehicles, and the relative relationship between the unmanned vehicle and the surrounding vehicles; normalize each parameter in the initial state space to obtain the state space.
[0165] In this embodiment, the first determination unit in the operation decision device applied to the unmanned vehicle determines the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located; the second determination unit determines the action space of the unmanned vehicle based on the parameters to be controlled during the operation of the unmanned vehicle; the training unit trains the initial decision model for controlling the operation of the unmanned vehicle according to the state space, action space and preset reward function, and obtains the trained decision model, thereby providing an operation decision method based on reinforcement learning applied to the unmanned vehicle. Based on the established state space, action space and reward function of the unmanned vehicle, the trained decision model is obtained, which improves the accuracy of the decision model and thus improves the operation safety of the unmanned vehicle.
[0166] Continue to refer Figure 12 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a decision-making device for an unmanned vehicle. Figure 10 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0167] like Figure 12 As shown, the decision-making device applied to the unmanned vehicle includes: an acquisition unit 1201, configured to obtain a pre-trained decision model, wherein the decision model represents the correspondence between the state space of the unmanned vehicle at the current moment and the action space at the next moment; a control unit 1202, configured to control the operation of the unmanned vehicle through the decision model.
[0168] In this embodiment, the operation of the unmanned vehicle is controlled by a decision model, thereby improving the operational safety of the unmanned vehicle.
[0169] Reference below Figure 13 , which shows a device suitable for implementing the embodiments of the present application (eg Figure 1 Schematic diagram of the structure of the computer system 1300 of the devices 101, 102, 103, 105 shown. Figure 13 The device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.
[0170] like Figure 13 As shown, the computer system 1300 includes a processor (e.g., a CPU, a central processing unit) 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage unit 1308 into a random access memory (RAM) 1303. Various programs and data required for the operation of the system 1300 are also stored in the RAM 1303. The processor 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.
[0171] The following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, a mouse, and the like; an output section 1307 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 1308 including a hard disk; and a communication section 1309 including a network interface card such as a LAN card or a modem. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as needed. Removable media 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1310 as needed, so that computer programs read therefrom can be installed into the storage section 1308 as needed.
[0172] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1309 and / or installed from the removable medium 1311. When the computer program is executed by the processor 1301, the above-mentioned functions defined in the method of the present application are performed.
[0173] It should be noted that the computer-readable medium of the present application may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0174] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the client computer, partially on the client computer, as a stand-alone software package, partially on the client computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the client computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0175] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0176] The units involved in the embodiments described in this application can be implemented by software or by hardware. The units described can also be set in a processor. For example, they can be described as: a processor including a first determination unit, a second determination unit and a training unit; they can also be described as: a processor including an acquisition unit and a control unit. The names of these units do not constitute a limitation on the units themselves in some cases. For example, the training unit can also be described as "a unit that trains an initial decision model for controlling the operation of an unmanned vehicle based on a state space, an action space and a preset reward function to obtain a trained decision model."
[0177] As another aspect, the present application also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the computer device: determines the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located; determines the action space of the unmanned vehicle based on the parameters to be controlled during the operation of the unmanned vehicle; trains the initial decision model for controlling the operation of the unmanned vehicle according to the state space, action space and preset reward function to obtain a trained decision model. The computer device also: obtains a pre-trained decision model, wherein the decision model represents the correspondence between the state space of the unmanned vehicle at the current moment and the action space at the next moment; and controls the operation of the unmanned vehicle through the decision model.
[0178] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. An operation decision-making method for an unmanned vehicle, comprising: Determining a state space of the unmanned vehicle based on state information of the unmanned vehicle and environmental information of an environment in which the unmanned vehicle is located; Determining an action space of the unmanned vehicle based on parameters to be controlled during operation of the unmanned vehicle; Training an initial decision model for controlling the operation of the unmanned vehicle according to the state space, the action space, and a preset reward function to obtain a trained decision model includes: The following training operations are performed cyclically to obtain the decision model based on the obtained cumulative reward value: Starting from the initialization state of the preset training task corresponding to the training operation, updating the initial decision model after the previous training operation according to the state space, the action space and the preset reward function until the preset end condition is reached, ending the current training operation, and determining the task reward value corresponding to the current training operation; Determine the current cumulative reward value based on the task reward value corresponding to each training operation that has been executed; The step of starting from the initialization state of the preset training task corresponding to the training operation and updating the initial decision model trained in the previous training operation according to the state space, the action space, and the preset reward function includes: The following update operations are performed in a loop: Determining whether the preset end condition is met according to the current state space; In response to determining that the preset end condition has not been met, inputting the current state space into the preset reward function to determine a single-step reward value; updating the initial decision model according to the single-step reward value; The action space at the next moment is output according to the updated initial decision model to control the simulation operation of the unmanned vehicle and obtain the state space corresponding to the next moment.
2. The method according to claim 1, wherein The preset end condition includes the unmanned vehicle completing the preset training task or the unmanned vehicle colliding; as well as Determining the task reward value corresponding to the current training operation includes: In response to determining that the unmanned vehicle completes the preset training task according to the state space corresponding to the current moment, determining a first task reward value, wherein the first task reward value is a positive value; In response to determining that the unmanned vehicle has collided according to the state space corresponding to the current moment, a second task reward value is determined, wherein the second task reward value is a negative value.
3. The method according to claim 1, wherein In response to determining that the preset end condition is not met, inputting the current state space into the preset reward function to determine a single-step reward value includes: In response to determining that the preset end condition has not been met, the state space at the current moment is input into the preset reward function, and the single-step reward value is determined based on the information in the state space that represents the driving safety, task completion, driving efficiency and driving comfort up to the current moment.
4. The method according to claim 3, wherein: The determining of the single-step reward value based on the information in the state space representing the driving safety, task completion, driving efficiency, and driving comfort to date includes: Determining a first single-step reward sub-value corresponding to driving safety based on information in the state space representing the distance of the unmanned vehicle relative to the centerline of the lane in which it is located and the closest distance between the unmanned vehicle and surrounding vehicles; Determining a second single-step reward sub-value corresponding to the task completion degree based on information representing the distance between the unmanned vehicle and the target lane, the angle of the unmanned vehicle relative to the centerline of the lane, and the position of the unmanned vehicle in the state space; determining a third single-step reward sub-value corresponding to driving efficiency based on information representing the speed of the unmanned vehicle in the state space; determining a fourth single-step feedback sub-value corresponding to driving comfort based on information representing the yaw angular velocity, steering wheel angle, and longitudinal acceleration of the unmanned vehicle in the state space; The single-step report value is determined according to the first single-step report sub-value, the second single-step report sub-value, the third single-step report sub-value, and the fourth single-step report sub-value.
5. The method according to claim 1, wherein The determining of the state space of the unmanned vehicle based on the state information of the unmanned vehicle and the environmental information of the environment in which the unmanned vehicle is located includes: The state space is determined based on state information of the unmanned vehicle, a relative relationship between the unmanned vehicle and a road, state information of vehicles surrounding the unmanned vehicle, and a relative relationship between the unmanned vehicle and the surrounding vehicles.
6. The method according to claim 5, wherein: The determining of the state space based on the state information of the unmanned vehicle, the relative relationship between the unmanned vehicle and the road, the state information of vehicles surrounding the unmanned vehicle, and the relative relationship between the unmanned vehicle and the surrounding vehicles includes: Determining an initial state space of the unmanned vehicle based on state information of the unmanned vehicle, a relative relationship between the unmanned vehicle and a road, state information of vehicles surrounding the unmanned vehicle, and a relative relationship between the unmanned vehicle and the surrounding vehicles; Each parameter in the initial state space is normalized to obtain the state space.
7. A decision-making method for an unmanned vehicle, comprising: Obtaining a pre-trained decision model, wherein the decision model is obtained according to any one of claims 1 to 6; The operation of the unmanned vehicle is controlled by the decision model.
8. An operation decision-making device for an unmanned vehicle, comprising: A first determining unit is configured to determine a state space of the unmanned vehicle based on state information of the unmanned vehicle and environmental information of an environment in which the unmanned vehicle is located; A second determining unit is configured to determine an action space of the unmanned vehicle based on a parameter to be controlled when the unmanned vehicle is running; A training unit is configured to train an initial decision model for controlling the operation of the unmanned vehicle based on the state space, the action space, and a preset reward function to obtain a trained decision model, including: The following training operations are performed cyclically to obtain the decision model based on the obtained cumulative reward value: Starting from the initialization state of the preset training task corresponding to the training operation, updating the initial decision model after the previous training operation according to the state space, the action space and the preset reward function until the preset end condition is reached, ending the current training operation, and determining the task reward value corresponding to the current training operation; Determine the current cumulative reward value based on the task reward value corresponding to each training operation that has been executed; The step of starting from the initialization state of the preset training task corresponding to the training operation and updating the initial decision model trained in the previous training operation according to the state space, the action space, and the preset reward function includes: The following update operations are performed in a loop: Determining whether the preset end condition is met according to the current state space; In response to determining that the preset end condition has not been met, inputting the current state space into the preset reward function to determine a single-step reward value; updating the initial decision model according to the single-step reward value; The action space at the next moment is output according to the updated initial decision model to control the simulation operation of the unmanned vehicle and obtain the state space corresponding to the next moment.
9. A decision-making device for an unmanned vehicle, comprising: an acquisition unit configured to acquire a pre-trained decision model, wherein the decision model is obtained according to claim 8; A control unit is configured to control the operation of the unmanned vehicle through the decision model.
10. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic driving method for intersection scene and related equipment
CN113264064A