Automatic driving decision-making method and device and storage medium
By using a multi-armed slot machine model and reinforcement learning techniques, the probability of course selection is dynamically adjusted to train the decision-making strategy of autonomous vehicles. This solves the problem of uncertainty in the driving intentions of surrounding vehicles and improves the success rate and robustness of decision-making in complex scenarios such as unsignaled intersections.
Patent Information
- Application Number
- CN202511103027.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-18
AI Technical Summary
In high-traffic, high-interaction traffic scenarios, autonomous vehicles struggle to accurately assess the driving intentions of surrounding vehicles, leading to safety issues in motion planning, especially at unsignalized intersections where effective decision-making and response are difficult.
A multi-armed slot machine model is used for policy reinforcement learning. By constructing an increasing sequence of surrounding vehicle numbers, the probability of course selection is dynamically adjusted. Combined with a proximal policy optimization algorithm and an adaptive reward mechanism, the decision-making strategy of autonomous vehicles is trained to gradually adapt to the dynamic uncertainty of complex road environments.
It significantly improves the decision success rate and generalization performance of autonomous driving systems in complex scenarios, enhances robustness and training efficiency, and can better cope with uncertainties in complex scenarios such as unsignaled intersections.
Smart Images

Figure CN120972673A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present application relates to, but is not limited to, the technical field of artificial intelligence, and particularly relates to an automatic driving decision method, an electronic device and a computer readable storage medium. BACKGROUND
[0002] In recent years, significant progress has been made in the academic and industrial fields of automatic driving technology. However, in high traffic density and frequent interaction traffic scenarios, it is still a major challenge to achieve reliable automatic driving, and the main difficulty lies in the uncertainty of surrounding vehicle (SV) driving intention. If the evaluation of SV driving intention is not accurate enough, it will seriously affect the safety of motion planning, and even may cause traffic accidents. This problem is particularly prominent when an automatic driving vehicle drives into a signal-free intersection. In such scenarios, the automatic driving vehicle needs to coordinate with multiple vehicles from different directions at the same time, and the traffic flow pattern at the intersection is often highly unpredictable. The uncertainty of SV driving behavior further exacerbates the safety risk during interaction, making it difficult for the automatic driving system to accurately predict potential dangers and make appropriate responses. Therefore, achieving efficient decision-making and accurate environmental perception has become a core challenge for automatic driving. SUMMARY
[0003] The embodiment of the present application provides an automatic driving decision method, an electronic device and a computer readable storage medium, which aims to enhance the decision success rate of the automatic driving system in complex scenarios such as signal-free intersections.
[0004] In a first aspect, the embodiment of the present application provides an automatic driving decision method, which comprises: constructing a curriculum set based on an increasing sequence of surrounding vehicle numbers; regarding each curriculum in the curriculum set as an arm of a multi-armed bandit respectively; performing strategy reinforcement learning through multiple rounds of iterative training until a training end condition is met to obtain a target strategy, wherein each round of iterative training comprises the following steps: sampling a curriculum for a current iterative training period using the multi-armed bandit, setting an automatic driving vehicle and a corresponding number of surrounding vehicles in the target road scene according to a preset rule to obtain a reset environment, an agent interacting with the reset environment and deciding an action of the automatic driving vehicle through the strategy, calculating a current reward based on a result of the automatic driving vehicle executing the action, and updating parameters of the strategy and a probability distribution of the multi-armed bandit according to the current reward; deciding an action of an automatic driving vehicle located in the target road scene using the target strategy.
[0005] The embodiment of the present application has at least the following beneficial effects: The embodiment of the present application models the course selection as a multi-armed bandit problem, adjusts the course selection probability in real time by using a dynamic reward feedback mechanism, realizes performance-based adaptive course selection, enables the training process to preferentially screen high-value course combinations, and significantly improves sample utilization efficiency and strategy convergence speed. Meanwhile, by constructing a hierarchical course set with an increasing number of surrounding vehicles, the strategy gradually adapts to complex road environments, effectively deals with the dynamic uncertainty of complex road scenes, and in the training mode provided by the embodiment of the present application, the decision success rate of the strategy in complex scenes is significantly better than that in the fixed course training mode, which not only improves the training efficiency, but also enhances the generalization performance and robustness of the automatic driving system in complex road scenes such as signal-free crossroads.
[0006] In a possible implementation manner of the present application, the calculating the current reward based on the result of the automatic driving vehicle performing the action comprises: obtaining a reward score and a penalty score obtained after the automatic driving vehicle performs the action, the reward score being determined based on at least one of the following reward items: successfully completing a task, surviving in a task, and the penalty score being determined based on at least one of the following penalty items: collision, timeout, exceeding a road boundary, and lane changing; determining the current reward based on the reward score and the penalty score.
[0007] In a possible implementation manner of the present application, the updating process of the probability distribution of the multi-armed bandit comprises the following steps: determining an original reward of the selected arm in the current iteration training period according to the current reward, the maximum reward up to the current iteration training period, and the minimum reward up to the current iteration training period; dividing the original reward of the selected arm by the current probability of the selected arm to obtain a re-adjusted reward of the selected arm; updating the weight vector of the selected arm according to the re-adjusted reward of the selected arm; updating the probability distribution of the multi-armed bandit according to the updated weight vector of the selected arm.
[0008] In a possible implementation manner of the present application, the updating the probability distribution of the multi-armed bandit according to the updated weight vector of the selected arm comprises: updating the probability distribution of the target multi-armed bandit according to the updated weight vector of the selected arm; after a preset number of iteration training periods, synchronizing the probability distribution of the target multi-armed bandit to the multi-armed bandit.
[0009] In a possible implementation manner of the present application, the updating process of the parameters of the strategy comprises: updating parameters of the policy with the goal of maximizing a cumulative objective function associated with the curriculum set, expressed as follows:
[0010] wherein represents an objective function of the policy with . represents parameters of the policy the value of the policy at the th iteration training period, represents the curriculum set.
[0011] In a possible implementation manner of the present application, a proximal policy optimization algorithm is adopted in the process of updating the parameters of the policy to clip the objective function, and the expression of the objective function is as follows:
[0012] wherein represents the similarity between the new policy and the old policy; is an estimated advantage function; is a clipping parameter.
[0013] In a possible implementation manner of the present application, the target scene includes a no-signal intersection scene, and the automatic driving vehicle and the surrounding vehicles corresponding to the curriculum in number are set in the target road scene according to the preset rule, including: the automatic driving vehicle is set at a random position of any lane in the lower area of the no-signal intersection; on the premise of conforming to the traffic rules, the surrounding vehicles corresponding to the sampling curriculum in number are randomly generated on the random lanes in the left, upper and right areas of the no-signal intersection.
[0014] In a second aspect, the embodiments of the present application further provide a device for automatically designing a field model, and the device comprises: a curriculum construction module configured to construct a curriculum set based on an incremental surrounding vehicle number sequence; a curriculum selection module configured to take each curriculum in the curriculum set as an arm of a multi-armed bandit respectively; a training module configured to perform policy reinforcement learning through multiple rounds of iterative training until a training end condition is met to obtain a target policy, wherein each round of iterative training comprises the following steps: sampling a lesson for a current iteration training period using the multi-armed bandit, setting the autonomous vehicle and a corresponding number of surrounding vehicles according to a preset rule to obtain a reset environment in a target road scene, the agent interacting with the reset environment and deciding an action of the autonomous vehicle according to the policy, calculating a current reward based on a result of the autonomous vehicle performing the action, and updating parameters of the policy and probability distribution of the multi-armed bandit according to the current reward; a decision module configured to decide an action of the autonomous vehicle in the target road scene according to the target policy.
[0015] In a third aspect, an electronic device is provided, which includes at least one processor, at least one memory configured to store at least one program, and at least one of the programs is executed by the at least one processor to implement the autonomous driving decision method of the first aspect.
[0016] In a fourth aspect, a computer readable storage medium is provided, which stores computer executable instructions for implementing the autonomous driving decision method of the first aspect.
[0017] The beneficial effects of the second aspect to the fourth aspect can be referred to the beneficial effects of the first aspect, and will not be repeated here.
[0018] It should be understood that the foregoing general description and the following detailed description are only exemplary and do not limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 An architecture schematic diagram of an autonomous driving decision system suitable for the embodiments of the present application; Figure 2 A flow schematic diagram of an autonomous driving decision method provided by the embodiments of the present application; Figure 3 An autonomous driving policy training architecture schematic diagram of a two-lane signal-free intersection scene provided by the embodiments of the present application; Figure 4 A schematic diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0020] In order to make the purposes, technical methods and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application and not to limit the present application.
[0021] It should be noted that although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that in the flow chart. In the description of the specification and claims and the above description of the drawings, "at least one" means one or more, the meaning of multiple (or multiple items) is two or more, greater than, less than, more than, etc. is not included in the number, above, below, within, etc. is understood to include the number. If there is a description of "first", "second", etc. is only used to distinguish technical features for the purpose, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or implicitly indicating the relationship between the indicated technical features.
[0022] In the embodiments of the present application, the association relationship of the associated objects described by "and / or" represents that there can be three kinds of relationships, for example, A and / or B can represent the cases of A alone, A and B together, and B alone. Wherein A, B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" and the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b and c can represent: a alone, b alone, c alone, a and b together, a and c together, b and c together, or a and b and c together, wherein a, b, c can be single or multiple.
[0023] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Rather, the use of the words "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0024] In driving scenarios with high vehicle density and frequent interactions, the automatic driving task remains a challenging problem, with the core difficulty being the uncertainty of surrounding vehicle (SV) driving intentions. To address this challenge, the embodiments of the present application propose an automatic driving decision-making method, an electronic device, and a related computer-readable storage medium, aiming to achieve interactive perception of automatic driving capabilities, especially for dealing with uncertainties in complex environments such as signal-free intersections. These uncertainties mainly manifest as the unpredictability of surrounding vehicle driving intentions and the dynamic nature of SV numbers. The embodiments of the present application innovatively design an adaptive curriculum mechanism that can gradually adapt to increasing SV numbers. By introducing an automated curriculum selection strategy, this method can reasonably assign importance weights to different training courses, significantly improving sample utilization efficiency and model convergence effect. In addition, through a carefully designed reward function, the agent can more efficiently explore the optimal driving strategy.
[0025] In order to better understand and illustrate the scheme of the embodiments of the present application, some technical terms involved in the embodiments of the present application are briefly explained below.
[0026] Proximal Policy Optimization (PPO) is a widely used policy gradient method in the field of reinforcement learning, proposed by the OpenAI team, aiming to solve the problems of low sample efficiency and unstable training in traditional policy gradient algorithms. The core idea is to limit the magnitude of policy updates to ensure that the difference between the new policy and the old policy remains within a small range (i.e., "proximal" constraint), thereby improving training stability while maintaining learning efficiency. Specifically, PPO constructs a clipped objective function to avoid performance fluctuations caused by excessive policy parameter updates, and uses importance sampling techniques to fully reuse collected sample data, reducing the number of interactions with the environment.
[0027] Multi-armed Bandit is a classic model in the field of reinforcement learning that studies the trade-off between exploration and exploitation. Its name comes from a slot machine with multiple levers ("arms") — each lever corresponds to a different reward probability distribution, and the player obtains a random reward by pulling the lever. The goal is to maximize the total reward within a limited number of attempts. The core challenge here lies in balancing "exploration" (trying unknown levers to obtain their reward information) and "exploitation" (choosing the lever with the highest known reward). Over-exploration wastes opportunities, while over-exploitation may miss better choices. The multi-armed bandit model is widely used in scenarios requiring dynamic decision-making, such as recommendation systems, ad placement, and clinical trials. Common solutions include the ε-greedy algorithm, UCB (Upper Confidence Bound) algorithm, and Thompson sampling, which balance exploration and exploitation from different perspectives to maximize long-term cumulative rewards.
[0028] In order to better understand the scheme provided by the embodiments of the present application, the scheme will be described below in combination with a specific application scenario.
[0029] Please refer to Figure 1 , a schematic diagram of an automatic driving decision system architecture applicable to the embodiments of the present application can be understood that the automatic driving decision method provided by the embodiments of the present application can be applied to but not limited to the application scenarios such as Figure 1 as shown in the application scenario.
[0030] As shown in Figure 1 , the automatic driving decision system architecture in this example can include but is not limited to server 10, terminal 20, database 30 and network 40. The server 10, the terminal 20 and the database 30 can realize data interaction through the network 40. The terminal 20 is installed on an autonomous vehicle, and collects road scene data in real time by means of sensors such as vehicle-mounted laser radar and camera, wherein the data contains the number of surrounding vehicles, the position of each vehicle, the driving speed and direction, and the like, and the data is transmitted to the server 10 through the network 40. After receiving the data sent by the terminal 20, the server 10 starts to build a curriculum set, which is generated based on an increasing sequence of the number of surrounding vehicles (such as 0, 1, …, until a preset maximum number), and each curriculum corresponds to a specific number of surrounding vehicle scenes. At the same time, the server 10 takes each curriculum in the curriculum set as an arm of the multi-armed bandit, and initializes the probability distribution of the multi-armed bandit. Then, the server 10 performs multi-round iterative training on the strategy. In each round of iterative training, the server 10 uses the multi-armed bandit to sample the curriculum of the current iteration training period, sets the autonomous vehicle and the corresponding number of surrounding vehicles in the target road scene (such as a signal-free intersection) according to the preset rule, to obtain a reset environment. Then, the server 10 controls the policy agent to interact with the reset environment, and the policy decides the action of the autonomous vehicle. The server 10 calculates the current reward according to the result of the autonomous vehicle executing the action (such as whether a collision occurs or not), and then updates the parameters of the policy and the probability distribution of the multi-armed bandit according to the current reward. This iteration process is repeated until the training end condition is met (such as reaching a preset number of training rounds, or the reward is stable above a preset threshold for a plurality of consecutive rounds), the server 10 determines the policy obtained at this time as the target policy, and stores it in the database 30. When the autonomous vehicle drives to the target road scene (such as a signal-free intersection), the terminal 20 collects the current scene data in real time and sends it to the server 10 through the network 40. The server 10 retrieves the target policy from the database 30, interacts with the current environment data, decides the action of the autonomous vehicle by using the target policy, and sends the action instruction to the terminal 20 through the network 40. After receiving the instruction, the terminal 20 controls the autonomous vehicle to execute the corresponding action, ensuring the safety of the vehicle passing through the intersection.
[0031] It can be understood that the above is only an example, and the present embodiment is not limited thereto.
[0032] The terminal includes, but is not limited to, a smart phone (such as an Android phone, an iOS phone, etc.), a mobile phone simulator, a tablet computer, a notebook computer, a digital broadcast receiver, a MID (Mobile Internet Device), a PDA (Personal Digital Assistant), a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, and the like.
[0033] The server can be a standalone physical server, a server cluster composed of multiple physical servers, or a distributed system, and can also be a cloud server or a server cluster providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.
[0034] The network can include, but is not limited to, a wired network including a local area network, a metropolitan area network, and a wide area network, and a wireless network including Bluetooth, Wi-Fi, and other networks that enable wireless communication. The actual application scenario requirements can also be determined, and are not limited here.
[0035] See Figure 2 A flowchart of an automatic driving decision-making method provided by an embodiment of the present application is shown. The method can be executed by any electronic device, such as a server. As an optional embodiment, the method can be executed by a server in a system for automatic driving decision-making. For the sake of convenience, the server will be taken as an example to illustrate the method in the description of some optional embodiments below. As shown in Figure 2 The automatic driving decision-making method provided by the embodiment of the present application includes the following steps: Step S101, constructing a curriculum set based on an increasing surrounding vehicle number sequence.
[0036] For example, the increasing surrounding vehicle number sequence is: (Formula 1) Correspondingly, constructing a curriculum set based on an increasing surrounding vehicle number sequence can be expressed as:
[0037] wherein indicates the serial number of the curriculum, and its value is equal to the number of surrounding vehicles in the corresponding curriculum.
[0038] It can be understood that the training process of the reinforcement learning (RL) strategy is regarded as a curriculum learning problem of a target task setting in the embodiment of the present application, which contains an increasing surrounding vehicle number curriculum. The target task can include a left turn task at a signal-free intersection, etc.
[0039] Step S102, taking each curriculum in the curriculum set as an arm of a multi-armed bandit.
[0040] It can be understood that the automatic driving strategy training needs to cope with the uncertainty of the driving intention of the surrounding vehicles and the uncertainty of the number of surrounding vehicles, and the use of a multi-armed bandit (MAB) to select episodes with different numbers of surrounding vehicles for training can specifically solve these challenges: the multi-armed bandit sets scenarios with different numbers of surrounding vehicles as independent "arms", and optimizes the selection of the strategy by dynamically evaluating the training value of each scenario. For scenarios with few surrounding vehicles but ambiguous intentions (such as a single vehicle suddenly slowing down or changing direction), it will ensure that the agent masters the basic interaction logic through moderate exploration; for high-complexity scenarios with many surrounding vehicles and intertwined intentions (such as multiple vehicles simultaneously competing for a lane or lane weaving conflicts), it will significantly increase the selection weight, forcing the agent to focus on such key scenarios that are prone to decision confusion, quickly accumulating experience to cope with complex games, avoiding the problem of repeatedly consuming resources without being able to touch the core difficulties in traditional random training, enabling the agent to both stably handle intention guessing in sparse traffic and efficiently cope with coordinated decision-making in dense traffic, and ultimately forming a more robust passing strategy under uncertainty.
[0041] In step S103, policy reinforcement learning is performed through multiple rounds of iterative training until a training end condition is met, and a target policy is obtained, wherein each round of iterative training includes the following steps: a multi-armed bandit is used to sample a lesson for the current iterative training period, the autonomous driving vehicle and a number of surrounding vehicles corresponding to the sampled lesson are set in the target road scene according to a preset rule to obtain a reset environment, the agent interacts with the reset environment and makes a policy decision on the action of the autonomous driving vehicle, the current reward is calculated based on the result of the action performed by the autonomous driving vehicle, and the parameters of the policy and the probability distribution of the multi-armed bandit are updated according to the current reward.
[0042] For example, the training framework of policy reinforcement learning includes an agent, an environment, a reward, and a policy, wherein the agent learns the policy through interaction with the environment, and the core goal is to maximize the long-term reward; the environment provides the scene for the agent to interact with, and returns the state, reward, and termination information; the reward is used to guide the behavior of the agent learning; and the policy is used to determine the action selection of the agent in the state.
[0043] In the embodiments of the present application, each lesson in the lesson set is regarded as an arm in the multi-armed bandit problem, and when a lesson is selected to collect a reinforcement learning episode, the relevant state and reward can be obtained from the environment. This process is similar to a step of running the multi-armed bandit, so the lesson selection problem can be regarded as a sampling process of a multi-armed bandit with arms , and the multi-armed bandit agent will sample a series of arms times during the training process: (Formula 3) wherein, is the arm of the multi-armed bandit sampled in the i-th episode.
[0044] Specifically, in each round of training, the selected arm returns a reward, and the multi-armed bandit agent updates the importance weight according to the historical reward. The goal of embodiments of the present application is to design an adaptive mechanism for maximizing the reward of the sampled curriculum sequence, which is expressed as follows: (Formula 4) wherein, is the importance weight vector of the multi-armed bandit; denotes the re-adjusted reward.
[0045] For example, the importance weight vector of the multi-armed bandit and the probability distribution vector can be expressed as follows: (Formula 5) It can be understood that, due to the change of the expected reward related to the specific task, the optimal arm is different in different training stages. Inspired by the Exp3 algorithm, this problem can be solved by introducing a -greedy term in the probability update. This ensures that all arms have a probability of being selected throughout the training process. For a given i-th episode weight vector , the sampling probability of the arm can be calculated by the following exponential weight algorithm: (Formula 6) wherein, is a constant parameter for balancing the utilization of experience data and random exploration. Then, embodiments of the present application sample the i-th episode curriculum sample from the calculated distribution .
[0046] In a possible implementation of the present application, the updating process of the probability distribution of the multi-armed bandit includes the following steps: determining the original reward of the selected arm according to the current reward, the maximum reward up to the current iteration training period, and the minimum reward up to the current iteration training period; dividing the original reward of the selected arm by the current probability of the selected arm to obtain the re-adjusted reward of the selected arm; updating the weight vector of the selected arm according to the re-adjusted reward of the selected arm; and updating the probability distribution of the multi-armed bandit according to the updated weight vector of the selected arm.
[0047] Specifically, the multi-armed bandit agent will obtain a reward from a training episode and then re-scale the reward to evaluate the performance of the selected arm, which is calculated as follows: (Equation 7) wherein, is the original reward of the i-th arm obtained in the j-th episode, is the re-scaled reward of the i-th arm obtained in the j-th episode; and are the maximum reward and the minimum reward in the un-scaled reward history up to the j-th episode, respectively; and are two constant parameters that adjust the re-scaling process. Here, all obtained rewards are divided by their current probabilities, ensuring that arms with potential optimality but low current probabilities can be quickly discovered through policy updates. Then, the embodiments of the present application update the weight vector of the i-th arm as shown in the following expression: (Equation 8) wherein is a constant parameter for adjusting the growth rate of the importance weight of each arm. In a possible implementation of the present application, the probability distribution of the multi-armed bandit is updated according to the updated weight vector of the selected arm, including: updating the probability distribution of the target multi-armed bandit according to the updated weight vector of the selected arm; and after a preset number of iteration training cycles, synchronizing the probability distribution of the target multi-armed bandit to the multi-armed bandit. (Equation 8) wherein is a constant parameter for adjusting the growth rate of the importance weight of each arm.
[0048] In a possible implementation of the present application, the probability distribution of the multi-armed bandit is updated according to the updated weight vector of the selected arm, including: updating the probability distribution of the target multi-armed bandit according to the updated weight vector of the selected arm; and after a preset number of iteration training cycles, synchronizing the probability distribution of the target multi-armed bandit to the multi-armed bandit.
[0049] It can be understood that the purpose of introducing the target multi-armed bandit for the multi-armed bandit in the embodiments of the present application is to delay updating the probability distribution of the multi-armed bandit, to reduce the overestimation or underestimation of the arm return due to accidental sampling, and to alleviate the instability of training.
[0050] It can be understood that the embodiments of the present application model course selection as a multi-armed bandit problem, use a dynamic reward feedback mechanism to adjust course selection probabilities in real time, and realize performance-based adaptive course selection. Based on the above automatic course selection process, the embodiments of the present application can sample a course schedule sequence from a multi-armed bandit to generate episodes for training of a reinforcement learning policy. The corresponding information including observations, actions and rewards is stored in a replay buffer for policy training. The reinforcement learning policy is trained to maximize the following cumulative objective function related to the sampled course schedule sequence as follows: (Formula 9) wherein is the objective function of the policy network with In the embodiments of the present application, the clipping objective function of the proximal policy optimization (PPO) algorithm is used to train the reinforcement learning policy:
[0051] wherein denotes the similarity between the new policy and the old policy; is the estimated advantage function; is the clipping parameter.
[0052] In a possible implementation manner of the present application, the reward obtained by the autonomous vehicle performing the action is calculated, including: obtaining a reward score and a penalty score obtained after the autonomous vehicle performs the action, the reward score being determined based on at least one of the following reward items: successfully completing a task, surviving in a task, the penalty score being determined based on at least one of the following penalty items: collision, timeout, exceeding a road boundary and lane changing; determining the current reward based on the reward score and the penalty score.
[0053] In a possible implementation manner, the reward is calculated by the following reward function: (Formula 11) wherein and are the rewards for successfully completing a task and surviving in a task, respectively; , , and are the penalties for collision with surrounding vehicles, timeout, exceeding a road boundary and lane changing behavior, respectively. is an indicator function corresponding to different events, as follows: (Formula 12) It can be understood that, in order to obtain higher task efficiency, the reward item for successfully completing a task and the penalty item for collision with surrounding vehicles is set as follows:
[0054] wherein is a constant parameter. In addition, in order to avoid unnecessary lane changing behavior of the reinforcement learning agent, the lane changing penalty is set to be positively correlated with the number of lane changes . The remaining reward / punishment items are set to be constants.
[0055] The embodiment of the present application models the course selection as a multi-armed bandit problem, uses a dynamic reward feedback mechanism to adjust the course selection probability in real time, realizes performance-based adaptive course selection, enables the training process to preferentially select high-value course combinations, and significantly improves the sample utilization efficiency and policy convergence speed. At the same time, by constructing a hierarchical course set with an increasing number of surrounding vehicles, the strategy gradually adapts to complex road environments, effectively dealing with the dynamic uncertainty of complex road scenes. Under the training mode provided by the embodiment of the present application, the decision success rate of the strategy in complex scenes is significantly better than that of the fixed course training mode, not only improving the training efficiency, but also enhancing the generalization performance and robustness of the autonomous driving system in complex road scenes such as signal-free intersections.
[0056] In a possible implementation manner of the present application, the target road scene is a signal-free intersection scene. The autonomous driving vehicle and the surrounding vehicles corresponding to the courses are set in the target road scene according to a preset rule, including: setting the autonomous driving vehicle on a random position of any lane in the lower area of the signal-free intersection; under the premise of complying with traffic rules, randomly generating surrounding vehicles corresponding to the number of sampled courses on random lanes in the left, upper and right areas of the signal-free intersection.
[0057] Step S104, using the target strategy to decide the action of the autonomous driving vehicle located in the target road scene.
[0058] After the training is completed, the trained target strategy can provide accurate action decisions for the vehicle by processing sensor data (including cameras, LiDAR, radars, etc.) of the autonomous driving vehicle in real time.
[0059] The embodiment of the present application provides a training technical solution of an autonomous driving strategy based on interactive perception, which can effectively deal with the uncertainty caused by the driving intentions of surrounding vehicles and different traffic flow densities in complex road scenes such as signal-free intersections. Please refer to Figure 3 , and the following will take the autonomous driving strategy of a two-lane signal-free intersection scene as an example to describe the solution of the embodiment of the present application.
[0060] Figure 3The example shown is performed in a two-lane unsignalized intersection scenario in the simulator Highway_Env. Here, it is assumed that there are up to 6 surrounding vehicles (SVs) at the intersection at the same time, so the episode set contains 7 episodes in total. The actor-critic architecture is adopted in this example as the training framework for the policy. The actor network and critic network are set as fully connected neural networks with 128 units and 64 units in a single hidden layer based on PyTorch, and are trained with the Adam optimizer. The number of training rounds is set to 20. The learning rates of the actor network and critic network are set to 5x and 1x , respectively. The parameters of the multi-armed bandit (MAB) are initialized as . Specifically, the following steps are included: Step S201: Calculate the episode distribution according to the MAB parameters, and sample an episode from the distribution. Step S202: Reset the environment for RL agent interaction according to the settings of the sampled episode, including different numbers of surrounding vehicles (SVs). Specifically, the random position of the autonomous vehicle (AV) on any lane in the lower area of the intersection can be generated, and the corresponding number of SVs in the left, upper, and right areas of the intersection can be randomly generated on the random lane under the premise of complying with traffic rules.
[0061] Step S203: The RL agent interacts with the set environment, and the neural network policy calculates after receiving the observation from the environment. The decision output is input into the underlying control module to output the final control amount and apply it to the AV. The state-action-reward pair information of the RL agent during the interaction process is recorded and stored.
[0062] Step S204: After a certain number of episodes are completed, the parameters of the policy network are updated.
[0063] Step S205: Calculate the readjustment reward to update the target MAB parameters; if a certain number of episodes are completed, synchronize the target MAB parameters to the MAB parameters. Step S206: Repeat the above steps S201-S205 until the maximum number of episodes is reached, end the training, and save the trained neural network parameters.
[0064]
[0065] Based on the above automatic driving decision method, it can be applied to the decision module in the automatic driving system in the high-fidelity simulator CARLA or even the actual signal-free intersection scene. The state information of the surrounding vehicles is obtained in real time through the vehicle-mounted sensor, and is input into the reinforcement learning strategy together with the state information of the ego vehicle for reasoning, so as to generate the final decision result for reference by the downstream planning controller. Further, the method can also be extended to more automatic driving scenes to complete different automatic driving tasks.
[0066] The embodiment of the present application also provides a device for automatically designing a field model, which comprises: a course construction module, configured to construct a course set based on an incremental surrounding vehicle quantity sequence; a course selection module, configured to take each course in the course set as an arm of a multi-armed bandit respectively; a training module, configured to perform policy reinforcement learning through multiple rounds of iterative training until a training end condition is met, so as to obtain a target policy, wherein each round of iterative training comprises the following steps: sampling a course of a current iterative training period by using the multi-armed bandit, setting the ego vehicle and a corresponding number of surrounding vehicles corresponding to the sampled course in a target road scene according to a preset rule to obtain a reset environment, interacting the agent with the reset environment and automatically driving the ego vehicle through a policy decision, calculating a current reward based on the result of the action performed by the ego vehicle, and updating the parameters of the policy and the probability distribution of the multi-armed bandit according to the current reward; a decision module, configured to decide the action of the ego vehicle in the target road scene by using the target policy.
[0067] Please refer to Figure 4 The embodiment of the present application also provides an electronic device, which comprises: at least one processor 1101; at least one memory 1102, configured to store at least one program; The at least one program is executed by the at least one processor 1101 to implement the automatic driving decision method of any of the foregoing embodiments.
[0068] The embodiment of the present application also provides a computer program product, comprising a computer program or computer instructions, the computer program or computer instructions being stored in a computer readable storage medium, and the processor of the electronic device reads the computer program or computer instructions from the computer readable storage medium, and the processor executes the computer program or computer instructions, so that the electronic device executes the automatic driving decision method of any of the foregoing embodiments.
[0069] The embodiment of the present application further provides a computer readable storage medium storing computer executable instructions, and the computer executable instructions are used for executing the automatic driving decision method of any one of the foregoing embodiments.
[0070] It should be understood that, in the embodiments of the present application, the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0071] In addition, the processor can include one or a combination of central processing units (CPUs), baseband processors, digital signal processors (DSPs), microprocessor units (MPUs), microcontroller units (MCUs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), artificial intelligence processors (AI processors) or neural network processors (Neural Processing Units, NPUs).
[0072] It should also be understood that the memory in the embodiments of the present application can be volatile or nonvolatile memory, or can include both volatile and nonvolatile memory. The nonvolatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory, among others. The volatile memory can be random access memory (RAM), which acts as external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM), among others. It should be noted that the memory described herein is intended to include, among others, these and any other suitable types of memory.
[0073] The above-described embodiments can be implemented in part or in whole through software, hardware, firmware or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When loaded and executed by a computer, the computer instructions or computer programs cause the computer to perform all or part of the processes or functions described in the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center through a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing a set of one or more available media. The available media can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid state disk.
[0074] It should be understood that the size of the sequence number of each process described above in various embodiments of the present application does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0075] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0076] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. For example, there can be another division manner for the actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electric, mechanical or in other forms.
[0077] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments. In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can be a physically independent unit, or two or more units can be integrated in a unit. When the above functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk, and various media that can store program codes.
[0078] Finally, it should be noted that: the above embodiments are merely specific implementations of the present application, used to explain the technical solutions of the present application, rather than limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can make modifications or easily think of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed by the present application, or make equivalent replacements to some technical features; and these modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application.
Claims
1. An autonomous driving decision-making method, characterized in that, The method includes: A course set is constructed based on an increasing sequence of surrounding vehicle numbers; Each course in the course set is treated as one arm of a multi-armed slot machine; Policy reinforcement learning is performed through multiple rounds of iterative training until the training termination condition is met to obtain the target policy. Each round of iterative training includes the following steps: using the multi-armed slot machine to sample the course of the current iteration training cycle, setting the autonomous vehicle and the number of surrounding vehicles corresponding to the sampled course in the target road scene according to preset rules to obtain a reset environment, the agent interacts with the reset environment and makes decisions on the actions of the autonomous vehicle through the policy, calculates the current reward based on the result of the autonomous vehicle performing the action, and updates the parameters of the policy and the probability distribution of the multi-armed slot machine according to the current reward. The target strategy is used to determine the actions of autonomous vehicles located in the target road scenario.
2. The method according to claim 1, characterized in that, The calculation of the current reward based on the result of the autonomous vehicle performing the action includes: The autonomous vehicle obtains a reward score and a penalty score after performing the action. The reward score is determined based on at least one of the following reward items: successfully completing the task and surviving in the task. The penalty score is determined based on at least one of the following penalty items: colliding with surrounding vehicles, exceeding the time limit, exceeding the road boundary, and changing lanes. The current reward is determined based on the reward score and the penalty score.
3. The method according to claim 1, characterized in that, The process of updating the probability distribution of the multi-armed slot machine includes the following steps: Based on the current reward, the maximum reward up to the current iteration training period, and the minimum reward up to the current iteration training period, determine the original reward of the selected arm in the current iteration training period; Divide the original reward of the selected arm by the current probability of the selected arm to obtain the rebalancing reward of the selected arm. Update the weight vector of the selected arm based on the rebalancing reward of the selected arm; The probability distribution of the multi-armed slot machine is updated based on the weight vector updated for the selected arm.
4. The method according to claim 3, characterized in that, The step of updating the probability distribution of the multi-armed slot machine based on the updated weight vector of the selected arm includes: Update the probability distribution of the target multi-armed slot machine based on the updated weight vector of the selected arm; After a preset number of iterative training cycles, the probability distribution of the target multi-armed slot machine is synchronized to the multi-armed slot machine.
5. The method according to claim 1, characterized in that, The parameter update process of the strategy includes: The parameters of the policy are updated with the objective of maximizing the cumulative objective function related to the course set, as expressed below: in Indicates having The objective function of the strategy, Parameters representing the strategy In the The value of each iteration training cycle, This indicates a set of courses.
6. The method according to claim 5, characterized in that, During the process of updating the parameters of the strategy, the objective function is pruned using a near-end policy optimization algorithm. The expression of the objective function is as follows: in This indicates the similarity between the new and old strategies; It is the estimated advantage function; These are the clipping parameters.
7. The method according to claim 1, characterized in that, The target scenario includes a no-signal intersection scenario. The step of placing the autonomous vehicle and a corresponding number of surrounding vehicles in the target road scenario according to preset rules includes: The autonomous vehicle is positioned at a random location in any lane of the area below the unsignalized intersection. Subject to traffic rules, a number of surrounding vehicles corresponding to the sampling course are randomly generated in random lanes in the left, upper, and right areas of the unsignalized intersection.
8. An autonomous driving decision-making device, characterized in that, The device includes: The course construction module is used to build course sets based on an incremental sequence of the number of surrounding vehicles; The course selection module is used to assign each course in the course set as one arm of the multi-armed slot machine. The training module is used to perform policy reinforcement learning through multiple rounds of iterative training until the training termination condition is met and the target policy is obtained. Each round of iterative training includes the following steps: sampling the course of the current iteration training cycle using the multi-armed slot machine; setting the autonomous vehicle and the number of surrounding vehicles corresponding to the sampled course in the target road scene according to preset rules to obtain a reset environment; the agent interacts with the reset environment and makes decisions on the actions of the autonomous vehicle through the policy; calculating the current reward based on the result of the autonomous vehicle performing the actions; and updating the parameters of the policy and the probability distribution of the multi-armed slot machine according to the current reward. The decision-making module is used to make decisions about the actions of autonomous vehicles located in the target road scenario using the target strategy.
9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; At least one of the programs is executed by at least one of the processors to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are used to perform the method described in any one of claims 1 to 7.