Enhanced trajectory optimization method based on trusted reviewer and computer equipment
By screening out the target evaluator and the target driving movements, and using the target evaluator to train the agent, the problem that a single evaluator is difficult to improve the autonomous driving effect and achieve better autonomous driving performance.
Patent Information
- Application Number
- CN202411993979.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, using a single evaluator to evaluate the actions output by the agent, it is difficult to achieve a better training effect, resulting in poor autonomous driving effect.
By obtaining sample status data of the autonomous driving agent, the target evaluator and target driving actions are screened using two candidate evaluators, and the agent is trained based on the target evaluator.
It improves the effect of the intelligent body in autonomous driving control and improves the overall performance of autonomous driving.
Smart Images

Figure CN120106121A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a trusted critic-based enhanced trajectory optimization method and a computer device. Background Art
[0002] Research on reinforcement learning is becoming more and more extensive. A basic scenario is to control the interaction between the target and the environment. For example, reinforcement learning is applied in the scenario of autonomous driving. It can determine the driving action based on the collected environmental information (state), and control the car to perform the determined driving action to achieve the purpose of autonomous driving.
[0003] In related technologies, during the process of training an intelligent agent, an evaluator is usually used to evaluate the actions given by the intelligent agent, so that the intelligent agent can use the evaluation value to update itself, and the evaluator can also use the actions output by the intelligent agent in the next round to update itself, thereby realizing dynamic training of the intelligent agent and the evaluator.
[0004] However, using a single evaluator may not achieve good training results, that is, the actions output by the intelligent agent cannot meet the requirements, which will lead to poor autonomous driving effects in the autonomous driving scenario. Summary of the invention
[0005] The embodiment of the present application provides a trusted critic-based enhanced trajectory optimization method and a computer device, which can improve the effect of the trained intelligent agent when performing autonomous driving control, thereby improving the effect of autonomous driving. The technical solution is as follows:
[0006] In one aspect, a trusted critic-based enhanced trajectory optimization method is provided, the method comprising:
[0007] During any round of training of the autonomous driving agent, sample state data of the current round of training is obtained, wherein the sample state data includes sample environment data and sample vehicle state data, and the agent is used to output driving actions based on the input state data;
[0008] Based on the sample state data and the agent, determine a target evaluator used in this round of training from two candidate evaluators and determine a target driving action, wherein the target evaluator is a candidate evaluator that meets a preset condition, and the target driving action is a driving action used in this round of training;
[0009] Based on the target evaluator, the sample state data and the target driving action, the agent is trained in this round.
[0010] On the one hand, a trusted critic-based enhanced trajectory optimization device is provided, the device comprising:
[0011] A data acquisition module, used for acquiring sample state data of the current round of training during any round of training of the autonomous driving agent, wherein the sample state data includes sample environment data and sample vehicle state data, and the agent is used for outputting driving actions based on the input state data;
[0012] A determination module, configured to determine a target evaluator used in this round of training from two candidate evaluators based on the sample state data and the agent, and to determine a target driving action, wherein the target evaluator is a candidate evaluator that meets a preset condition, and the target driving action is a driving action used in this round of training;
[0013] A training module is used to perform this round of training on the intelligent agent based on the target evaluator, the sample state data and the target driving action.
[0014] In one possible implementation, the determination module is used to input the sample state data into the agent to obtain at least two candidate actions; based on the sample state data and the at least two candidate actions, determine the target evaluator used in this round of training from the two candidate evaluators; based on the target evaluator, determine the target driving action from the at least two candidate actions.
[0015] In a possible implementation, the at least two candidate actions include a first candidate action and a second candidate action, the two candidate evaluators include a first candidate evaluator and a second candidate evaluator, the determination module is used to input the sample status data and the first candidate action into the first candidate evaluator and the second candidate evaluator respectively, to obtain a first evaluation value and a second evaluation value, the first evaluation value being the evaluation value of the first candidate action by the first candidate evaluator, and the second evaluation value being the evaluation value of the first candidate action by the second candidate evaluator; input the sample status data and the second candidate action into the first candidate evaluator and the second candidate evaluator respectively, to obtain a third evaluation value and a fourth evaluation value, the third evaluation value being the evaluation value of the second candidate action by the first candidate evaluator, and the fourth evaluation value being the evaluation value of the second candidate action by the second candidate evaluator; based on the first evaluation value, the second evaluation value, the third evaluation value and the fourth evaluation value, determine the target evaluator from the first candidate evaluator and the second candidate evaluator.
[0016] In a possible implementation, the determination module is used to determine the candidate evaluator corresponding to the lower evaluation value between the first evaluation value and the second evaluation value as the first reference evaluator, and the first reference evaluator is the first candidate evaluator or the second candidate evaluator; determine the candidate evaluator corresponding to the lower evaluation value between the third evaluation value and the fourth evaluation value as the second reference evaluator, and the first reference evaluator is the first candidate evaluator or the second candidate evaluator; and determine the reference evaluator with the higher evaluation value between the first reference evaluator and the second reference evaluator as the target evaluator.
[0017] In a possible implementation, the determination module is used to determine a target evaluation value corresponding to the target evaluator, where the target evaluation value is the first evaluation value, the second evaluation value, the third evaluation value or the fourth evaluation value; and determine the target driving action by using the candidate action corresponding to the target evaluation value.
[0018] In one possible implementation, the determination module is used to input the sample state data into the agent when the agent is used to output a random strategy, to obtain an action strategy corresponding to the sample state data; to perform random sampling in the action strategy to obtain the at least two candidate actions; to input the sample state data into the agent when the agent is used to output a deterministic strategy, to obtain an initial action corresponding to the sample state data; to superimpose the initial action with the target noise to obtain an action strategy; and to perform random sampling in the action strategy to obtain the at least two candidate actions.
[0019] In one possible implementation, the training module is used to determine a first policy gradient corresponding to a random strategy when the agent is used to output a random strategy; perform a current round of training on the agent based on a target evaluation value of the target driving action output by the target evaluator, a first action value of the target driving action under the sample state data output by the agent, and the first policy gradient; determine a second policy gradient corresponding to the output deterministic strategy when the agent is used to output a deterministic strategy; perform a current round of training on the agent based on a target evaluation value of the target driving action output by the target evaluator, the target driving action, and the second policy gradient.
[0020] In a possible implementation, the training module is used to substitute the target evaluation value and the first action value into the first policy gradient to obtain a first loss function; and use the first loss function to perform a current round of training on the parameters of the agent;
[0021] In a possible implementation, the training module is used to substitute the target evaluation value and the target driving action into the second policy gradient to obtain a second loss function; and use the second loss function to perform this round of training on the parameters of the agent.
[0022] In a possible implementation, the device also includes a prediction module for obtaining target state data of the target vehicle, wherein the target state data includes environmental data of the environment in which the target vehicle is located and target vehicle state data of the target vehicle; the target state data is input into the trained intelligent agent to obtain the target driving action corresponding to the target state data.
[0023] On the one hand, a computer device is provided, comprising one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the trusted critic-based enhanced trajectory optimization method.
[0024] On the one hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the trusted critic-based enhanced trajectory optimization method.
[0025] On the one hand, a computer program product or a computer program is provided, which includes a program code, the program code is stored in a computer-readable storage medium, a processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device performs the above-mentioned trusted critic-based enhanced trajectory optimization method.
[0026] Through the technical solution provided by the embodiment of the present application, during any round of training of the agent, the sample state data of this round of training is obtained. Using the sample state data and the agent, the target evaluator used in this round of training is determined from two candidate evaluators and the target driving action is determined, thereby implementing the screening of candidate evaluators and the screening of driving actions. By using the screened target evaluator, sample state data and target driving action to train the agent, a better training effect can be achieved. Thus, in the process of using the trained agent for autonomous driving, a better effect can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0028] Figure 1 is a schematic diagram of the Q function in reinforcement learning provided in an embodiment of the present application;
[0029] Figure 2 is a schematic diagram of an implementation environment of a trusted critic-based enhanced trajectory optimization method provided in an embodiment of the present application;
[0030] Figure 3 is a hierarchical diagram for autonomous driving using an intelligent agent, provided in an embodiment of the present application;
[0031] Figure 4 is a flow chart of a trusted critic-based enhanced trajectory optimization method provided in an embodiment of the present application;
[0032] Figure 5 is a flow chart of another enhanced trajectory optimization method based on trusted critics provided in an embodiment of the present application;
[0033] Figure 6 is a schematic diagram of a Gaussian distribution transformation provided in an embodiment of the present application;
[0034] Figure 7 is a schematic diagram of a trusted evaluator selection provided in an embodiment of the present application;
[0035] Figure 8 This is a schematic diagram of the effect of using an optimistic navigator to make a decision and a conservative navigator to make a decision provided by an embodiment of the present application;
[0036] Fig. 9 This is a flow chart of an embodiment of the present application providing a method of autonomous driving using an intelligent agent;
[0037] Fig.10 It is a schematic diagram of the structure of a trusted critic-based enhanced trajectory optimization device provided in an embodiment of the present application;
[0038] Fig.11 It is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0040] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with basically the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on quantity and execution order.
[0041] Autonomous driving: also known as driverless technology, refers to intelligent car technology that uses computer equipment to achieve driverless driving. This technology relies on the coordinated cooperation of artificial intelligence, visual computing, radar, monitoring devices and global positioning systems, allowing computer equipment to automatically and safely operate motor vehicles without any active human operation.
[0042] Reinforcement Learning: Consider a discounted infinite-horizon Markov decision process (MDP) defined as the five-tuple The state space and action space Is continuous, state transition probability p: Represents the probability density of the next state. Given the state at time step t and actions We can get The environment will give a bounded reward r based on a specific state and action: where γ is a discount factor whose value is in the range [0,1) so that the infinite cumulative reward is mathematically finite, r min ,r max denote the next and previous rewards respectively. The standard reinforcement learning algorithm aims to maximize the expected reward and, that is, where ρ π (s t ,a t ) represents the strategy π(a t |s t ) is the state-action marginal distribution of the trajectory distribution induced by . In the actual algorithm, the discount factor is considered and the standard Q value function is
[0043] For a fixed policy, the Q-value can be updated iteratively from an arbitrary function Q: To begin, repeatedly apply the Bellman backtracking operator Defined as:
[0044]
[0045] Bellman backtracking operator: The Bellman backtracking operator is a core concept in dynamic programming theory and plays a key role in the fields of reinforcement learning and optimal control. This operator cleverly captures the recursive nature of the decision-making process and provides us with an elegant way to describe and solve complex optimization problems.
[0046] Agent: In reinforcement learning, an agent is a subject that performs actions and interacts with the environment. The goal of the agent is to learn a strategy to maximize the cumulative reward it obtains from the environment. In the process of reinforcement learning, the agent learns how to act by interacting with the environment. The agent can be a software program or a physical entity. In the embodiment of the present application, the agent is used to output driving actions during the autonomous driving process.
[0047] Evaluator (Critic): In reinforcement learning, the critic is an important component, especially in the actor-critic method. The main role of the critic is to evaluate the goodness of the actions chosen by the actor, that is, to predict the total reward in the future based on the current state of the environment. The critic achieves this goal by calculating the value function.
[0048] Conservative critic: In the current mainstream reinforcement learning methods, two Q functions are usually used and initialized differently to ensure diversity. Based on these two Q functions, Q 1 and Q 2 , for a given state s t , which conservative commentators define as follows:
[0049] Q csvt (s t ,·)=min{Q 1 (s t ,·),Q 2 (s t ,·)}
[0050] Among them, this function represents the overlapping area of multiple Q functions, see Figure 1 , which, to some extent, indicates the consensus among them and excludes uncertainty. Therefore, it is widely used to estimate the target Q value.
[0051] Credible critics: Credible critics are built on conservative critics and bold critics. Conservative critics have been defined in the previous article. The definition of bold critics is as follows:
[0052] Q bld (s t ,·)=max{Q 1 (s t ,·),Q 2 (s t ,·)}
[0053] It can be found that the credible critic is an intermediate state between the conservative critic and the adventurous critic. Its basic definition is as follows:
[0054] Q csvt (s t ,·) opt (s t ,·) bld (s t ,·)
[0055] Combined with the specific schematic diagram, it can be found that the critics who are credible to some extent represent a trade-off between conservatism and boldness. Figure 1 As shown in Figure 2, different confidence levels also reflect differences between Q functions. Exploring this difference in the context of policy optimization can help accelerate the closing of this gap, thereby improving sample efficiency.
[0056] Random strategy: refers to the selection of actions in a given state based on probability distribution, that is, the action selected in the same state may be different each time. This strategy is usually used in scenarios that require exploration and utilization. For example, in reinforcement learning, the agent explores the environment through random strategies and tries different actions to obtain more reward information. The advantage of random strategies is that they can explore more possibilities and avoid falling into local optimal solutions; the disadvantage is that more sample data is required to estimate the policy gradient, and the computational complexity is high.
[0057] Deterministic strategy: means that in a given state, the choice of action is uniquely determined, that is, the action selected in the same state each time is fixed. This strategy is usually used in environments that have been fully explored or in scenarios that require precise control. The advantage of the deterministic strategy is high computational efficiency because it does not require sampling integration in the action space; the disadvantage is that it may fall into a local optimal solution due to lack of exploration.
[0058] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain better results.
[0059] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge sub-models to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.
[0060] Normalization: Mapping sequences with different value ranges to the interval (0, 1) facilitates data processing. In some cases, the normalized values can be directly implemented as probabilities.
[0061] Gaussian Distribution: Also known as Normal Distribution, the curve of Gaussian distribution is bell-shaped, high in the middle and low at both ends. The expected value μ of Gaussian distribution determines the position of the Gaussian distribution curve, and the standard deviation σ determines the range of the curve. When μ=0, σ=1, the Gaussian distribution is a standard Gaussian distribution.
[0062] Learning Rate: It is used to control the learning progress of the model. The learning rate can guide the model in how to use the gradient of the loss function to adjust the network weights in the gradient descent method. If the learning rate is too large, the loss function may directly cross the global optimal point, which is manifested as excessive loss; if the learning rate is too small, the loss function changes very slowly, which will greatly increase the convergence complexity of the network and it is easy to be trapped in the local minimum or saddle point.
[0063] Among related technologies, research on reinforcement learning is becoming more and more extensive. A basic scenario is to control the interaction between the target and the environment to generate a trajectory, such as controlling the driving of a car, which can generate a driving trajectory. The decisions made in this interactive process will be rewarded, and the reward for the entire trajectory is the sum of the rewards given for each decision in the trajectory. The goal of reinforcement learning is to obtain a control strategy that can maximize the trajectory reward. For example, in the field of autonomous driving, in the context of reliability, the strategy that can safely drive the longest distance is the optimal strategy.
[0064] In related technologies, an evaluator is usually designed to conservatively evaluate the value of driving actions, often giving low-value evaluations to high-value driving actions, resulting in the real value of driving actions not matching the evaluation value, so that the evaluation value is lower than the real value. The mismatch between value and evaluation requires more learning and correction processes, and too low evaluations reduce the enthusiasm for exploring driving actions. Based on conservative behavioral strategies, the optimization of continuous motion control by reinforcement learning is hindered, and the effect of trained intelligent agents on autonomous driving control is reduced, resulting in poor autonomous driving results.
[0065] Figure 2 This is a schematic diagram of an implementation environment of a trusted critic-based enhanced trajectory optimization method provided in an embodiment of the present application, see Figure 2 , the implementation environment may include a terminal 210 and a server 240.
[0066] The terminal 210 is connected to the server 240 via a wireless network or a wired network. Optionally, the terminal 210 is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal 210 is installed and runs an application that supports enhanced trajectory optimization based on trusted critics.
[0067] Server 240 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, distribution networks (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.
[0068] Optionally, terminal 210 refers to one of multiple terminals, and the embodiment of the present application only takes terminal 210 as an example.
[0069] The following is an introduction to the application scenarios of the technical solutions provided in the embodiments of the present application. Figure 3 , providing a flow chart of autonomous driving based on an intelligent agent. The top-level state perception mainly processes the input of state data, including destination information, map information, and sensor information. The middle-level algorithm optimization involves path planning and specific action control, which are all processed by the reinforcement learning algorithm, that is, by the intelligent agent trained by the technical solution (trusted critic reinforcement learning) provided in the embodiment of the present application. The intelligent agent receives state information and gives specific action control information, that is, obtains driving actions. After the bottom-level physical drive obtains the action control information, it adjusts the steering wheel angle, controls the throttle and brakes.
[0070] After introducing the implementation environment of this application, the following describes an enhanced trajectory optimization method based on a trusted critic provided by an embodiment of this application. Figure 4 Taking the execution subject as a server as an example, the method includes the following steps.
[0071] 401. During any round of training of the autonomous driving intelligent agent, the server obtains sample state data of this round of training, where the sample state data includes sample environment data and sample vehicle state data, and the intelligent agent is used to output driving actions based on the input state data.
[0072] Among them, the agent needs to go through multiple rounds of training before being used for autonomous driving. For ease of understanding, the embodiment of the present application takes any round of training in the multiple rounds of training as an example for explanation. The sample state data is the training sample used to train the agent in this round. Generally speaking, the sample state data includes sample environment data and sample vehicle state data. The sample environment data is used to represent the environmental state of the environment in which the sample vehicle is located, and the sample vehicle state data is used to represent the vehicle state of the sample vehicle. The driving action includes at least one of the steering wheel control action, the throttle control action, and the brake control action. Executing the driving action will change the vehicle state of the sample vehicle.
[0073] 402. The server determines a target evaluator used in this round of training from two candidate evaluators based on the sample state data and the agent and determines a target driving action, wherein the target evaluator is a candidate evaluator that meets preset conditions, and the target driving action is the driving action used in this round of training.
[0074] Among them, the candidate evaluator is an evaluator used to evaluate the driving action output by the agent. Determining the target evaluator from the two candidate evaluators is to determine a credible critic (a credible critic is also called an optimistic critic) from the two candidate evaluators. Accordingly, the candidate evaluator that meets the preset conditions is also the credible critic. The target driving action is the driving action adopted from at least two driving actions output by the agent.
[0075] 403. The server performs this round of training on the intelligent agent based on the target evaluator, the sample state data, and the target driving action.
[0076] Through the technical solution provided by the embodiment of the present application, during any round of training of the agent, the sample state data of this round of training is obtained. Using the sample state data and the agent, the target evaluator used in this round of training is determined from two candidate evaluators and the target driving action is determined, thereby implementing the screening of candidate evaluators and the screening of driving actions. By using the screened target evaluator, sample state data and target driving action to train the agent, a better training effect can be achieved. Thus, in the process of using the trained agent for autonomous driving, a better effect can be achieved.
[0077] Figure 5 This is a flowchart of a trusted critic-based enhanced trajectory optimization method provided in an embodiment of the present application, see Figure 5 , methods include:
[0078] 501. During any round of training of the autonomous driving intelligent agent, the server obtains sample state data of this round of training, where the sample state data includes sample environment data and sample vehicle state data, and the intelligent agent is used to output driving actions based on the input state data.
[0079] Among them, the agent needs to go through multiple rounds of training before being used for autonomous driving. For ease of understanding, the embodiment of the present application takes any round of training in multiple rounds of training as an example for explanation. The sample state data is the training sample used to train the agent in this round. Generally speaking, the sample state data includes sample environment data and sample vehicle state data. The sample environment data is used to represent the environmental state of the environment in which the sample vehicle is located, and the sample vehicle state data is used to represent the vehicle state of the sample vehicle. The driving action includes at least one of the steering wheel control action, the throttle control action, and the brake control action. Executing the driving action will change the vehicle state of the sample vehicle. The driving action output by the agent is based on the learned strategy, and the process of training the agent is also called the process of strategy improvement.
[0080] In some embodiments, sample environmental data includes sensory data interacting with the environment (such as camera data, radar data, and ultrasonic sensor data), specific information related to the driving scene (including lane information, road type, traffic rules, and neighboring vehicle status), dynamic and static information related to the environment (such as weather conditions, lighting conditions, and road conditions), historical information reflecting dynamic changes in time series (such as trajectory information and lane and signal light change information in the past few seconds), and information provided by high-precision maps (including information such as geometry, geographic location, and static obstacles).
[0081] In some embodiments, the sample vehicle status data includes information reflecting the vehicle's own status, such as position and posture, speed, acceleration, current steering angle information, throttle status, brake status, acceleration level and deceleration level, etc.
[0082] In a possible implementation, during any round of training of the autonomous driving agent, the server obtains sample state data of the current round of training from the sample database.
[0083] Among them, the sample database stores multiple sample state data for intelligent agent training. The multiple sample state data can be state data collected during the actual vehicle driving process or state data obtained in a simulated driving scene. The embodiment of the present application does not limit this.
[0084] 502. The server inputs the sample state data into the agent to obtain at least two candidate actions.
[0085] Among them, at least two candidate actions are output by the agent based on the sample state data, and the at least two candidate actions are driving actions that the agent believes match the sample state data. In an embodiment of the present application, the agent may be used to output either a random strategy or a deterministic strategy. When the agent is used to output a random strategy, the agent outputs an action strategy, and the candidate action can be obtained by sampling the action strategy. Since the candidate action is sampled from the action strategy, when the agent remains unchanged, different candidate actions may be obtained by inputting the same sample state data into the agent; when the agent is used to output a deterministic strategy, it can directly output an action, i.e., a t =π φ (s t ), a t Indicates action, π φ () represents the policy function of the deterministic policy, s t Represents sample state data. How to use this action to determine at least two candidate actions will be described later. In some embodiments, in the case of complex scenarios such as driving in the city, a random strategy can be used to improve the safety of autonomous driving; in the case of a single scenario such as high-speed driving, a deterministic strategy can be used to reduce the consumption of computing resources.
[0086] The following will explain the situations where the agent is used to output random strategies and deterministic strategies respectively.
[0087] In a possible implementation, when the agent is used to output a random strategy, the server inputs the sample state data into the agent to obtain an action strategy corresponding to the sample state data. The server performs random sampling in the action strategy to obtain the at least two candidate actions.
[0088] Here, an action policy is a distribution from which candidate actions can be obtained by sampling.
[0089] For example, when the agent is used to output a random strategy, the server inputs the sample state data into the agent, and extracts features of the sample state data through the agent to obtain sample state data features of the sample state data. The server generates an action strategy corresponding to the sample state data based on the sample state features through the agent. The server performs random sampling in the action strategy to obtain the at least two candidate actions.
[0090] For example, the action strategy can be expressed by the following formula (1).
[0091] π Q (·|s t )=exp(Q(st ,)-log Z(s t )) (1)
[0092] in, Used to normalize the distribution, s t Indicates the current state, that is, the sample state data, Q(s t ,) represents the current state s t The return under Q (·|s t ) indicates that the state s t The corresponding π Q (·|s t ). In the autonomous driving scenario, Q(s t , a t ) indicates that in the current state s t Take action a t The expected reward may be positive. For example, if the action is to accelerate, you can reach the destination faster, thus saving time and getting more positive rewards. This reward may also be negative. For example, if the action is also to accelerate, but in the current environment there is a vehicle in front that slows down, and taking the acceleration action causes a collision, the estimated reward is negative. Z(s t ) is in s t The rewards of all possible actions in the state are summed up, reflecting the state s t The return of π Q (·|s t ) is to associate all actions with rewards, and actions with high rewards will have high probabilities.
[0093] In a possible implementation, when the agent is used to output a deterministic strategy, the server inputs the sample state data into the agent to obtain an initial action corresponding to the sample state data. The server superimposes the initial action with the target noise to obtain an action strategy. Random sampling is performed in the action strategy to obtain the at least two candidate actions.
[0094] The target noise is Gaussian noise.
[0095] For example, when the agent is used to output a deterministic strategy, the server inputs the sample state data into the agent, and extracts features of the sample state data through the agent to obtain sample state data features of the sample state data. The server performs full connection and normalization on the sample state features through the agent to obtain an initial action corresponding to the sample state data. The server superimposes the initial action with Gaussian noise to obtain an action strategy. Random sampling is performed in the action strategy to obtain the at least two candidate actions.
[0096] For example, referring to the following formula (2), candidate actions can be obtained.
[0097]
[0098] Among them, a t represents a candidate action, π φ () represents the policy function of the deterministic policy, π φ (s t ) is the initial action, k is the weight, which is set by the technician according to the actual situation, for example, set to 0.1. ∈ t is Gaussian noise, the Gaussian noise ∈ t It conforms to a Gaussian distribution with a mean of 0 and a variance of 1. The basic idea of the above approach is that for a deterministic policy, only a distribution needs to be introduced, which allows sampling of nearby candidate actions, thereby obtaining at least two candidate actions.
[0099] 503. The server determines a target evaluator to be used in this round of training from two candidate evaluators based on the sample state data and the at least two candidate actions. The target evaluator is a candidate evaluator that meets preset conditions.
[0100] Among them, the candidate evaluator is an evaluator for evaluating the driving action output by the intelligent agent, and the two candidate evaluators are trained with different sample data in different initialization states, and the parameters of the two candidate evaluators are different. Determining the target evaluator from the two candidate evaluators is to determine a credible critic from the two candidate evaluators, and accordingly, the candidate evaluator that meets the preset conditions is also the credible critic. In the embodiment of the present application, the target evaluator (credible critic) is determined by multiple sampling.
[0101] In a possible implementation, taking at least two candidate actions as two candidate actions as an example, the at least two candidate actions include a first candidate action and a second candidate action, and the two candidate evaluators include a first candidate evaluator and a second candidate evaluator, the server inputs the sample state data and the first candidate action into the first candidate evaluator and the second candidate evaluator respectively, and obtains a first evaluation value and a second evaluation value, wherein the first evaluation value is the evaluation value of the first candidate evaluator to the first candidate action, and the second evaluation value is the evaluation value of the second candidate evaluator to the first candidate action. The server inputs the sample state data and the second candidate action into the first candidate evaluator and the second candidate evaluator respectively, and obtains a third evaluation value and a fourth evaluation value, wherein the third evaluation value is the evaluation value of the first candidate evaluator to the second candidate action, and the fourth evaluation value is the evaluation value of the second candidate evaluator to the second candidate action. The server determines the target evaluator from the first candidate evaluator and the second candidate evaluator based on the first evaluation value, the second evaluation value, the third evaluation value, and the fourth evaluation value.
[0102] In order to explain the above implementation more clearly, the above implementation is explained in several parts below.
[0103] In the first part, the server inputs the sample state data and the first candidate action into the first candidate evaluator and the second candidate evaluator respectively to obtain a first evaluation value and a second evaluation value.
[0104] In a possible implementation, the server inputs the sample state data and the first candidate action into the first candidate evaluator, respectively, and extracts features from the sample state data and the first candidate action through the first candidate evaluator to obtain a first evaluation feature. The server fully connects and normalizes the first evaluation feature through the first candidate evaluator to obtain the first evaluation value. The server inputs the sample state data and the first candidate action into the second candidate evaluator, respectively, and extracts features from the sample state data and the first candidate action through the second candidate evaluator to obtain a second evaluation feature. The server fully connects and normalizes the first evaluation feature through the second candidate evaluator to obtain the second evaluation value.
[0105] Since the parameters of the first candidate evaluator and the second candidate evaluator are different, even if the same sample state data and the first candidate action are input, the obtained first evaluation value and the second evaluation value are different.
[0106] In the second part, the server inputs the sample state data and the second candidate action into the first candidate evaluator and the second candidate evaluator respectively to obtain a third evaluation value and a fourth evaluation value.
[0107] In a possible implementation, the server inputs the sample state data and the second candidate action into the first candidate evaluator, respectively, and extracts features from the sample state data and the second candidate action through the first candidate evaluator to obtain a third evaluation feature. The server fully connects and normalizes the third evaluation feature through the first candidate evaluator to obtain the third evaluation value. The server inputs the sample state data and the second candidate action into the second candidate evaluator, respectively, and extracts features from the sample state data and the second candidate action through the second candidate evaluator to obtain a fourth evaluation feature. The server fully connects and normalizes the third evaluation feature through the second candidate evaluator to obtain the fourth evaluation value.
[0108] Since the parameters of the first candidate evaluator and the second candidate evaluator are different, even if the same sample state data and the second candidate action are input, the obtained third evaluation value and the fourth evaluation value are different.
[0109] In the third part, the server determines the target evaluator from the first candidate evaluator and the second candidate evaluator based on the first evaluation value, the second evaluation value, the third evaluation value and the fourth evaluation value.
[0110] In a possible implementation, the server determines the candidate evaluator corresponding to the lower evaluation value of the first evaluation value and the second evaluation value as the first reference evaluator, and the first reference evaluator is the first candidate evaluator or the second candidate evaluator. The server determines the candidate evaluator corresponding to the lower evaluation value of the third evaluation value and the fourth evaluation value as the second reference evaluator, and the first reference evaluator is the first candidate evaluator or the second candidate evaluator. The server determines the reference evaluator with the higher evaluation value of the first reference evaluator and the second reference evaluator as the target evaluator.
[0111] Among them, the reference evaluator is associated with the reference action. For example, the combination of the first candidate evaluator and the first candidate action and the combination of the first candidate evaluator and the second candidate action will be regarded as different reference evaluators in the above steps. Correspondingly, the combination of the second candidate evaluator and the first candidate action and the combination of the second candidate evaluator and the second candidate action will also be regarded as different reference evaluators in the above steps. Of course, the combination of the first candidate evaluator and the first candidate action and the combination of the second candidate evaluator and the first candidate action will also be regarded as different reference evaluators. Therefore, the determined reference evaluator is actually a combination of the candidate evaluation and the candidate action, and the target evaluator finally determined is also the combination of the candidate evaluation and the candidate action.
[0112] For example, the target evaluator can be determined by the following formula (3).
[0113] Q t (s t ,a t )=max(min(Q 1 (s t ,a 1 ),Q 2 (s t ,a 1 )),min(Q 1 (s t ,a 2 ),Q 2 (s t ,a 2 ))) (3)
[0114] Among them, Q t () is the target evaluator, Q 1 () is the first candidate evaluator, Q 2 () is the second candidate evaluator, a 1 represents the first candidate action, a 2 represents the second candidate action, min(Q 1 (s t , a 1 ), Q 2 (s t , a 1 )) represents the first reference evaluator, min(Q 1 (s t , a 2 ), Q 2 (s t , a 2 )) represents the second reference evaluator, Q t () represents the target evaluator.
[0115] It should be noted that the target evaluator determined by the above method is also a credible critic, and is therefore also called a credible evaluator. Figure 6 ,The above concept can be described as selecting a trusted evaluator to guide the policy optimization in a manner similar to the distributed Bellman backtracking operator. Figure 6 The data in this paper are generated entirely through numerical simulation. (Left) Gaussian distribution is an ideal model for continuous action space. (Right) The transformed distribution is obtained by taking the maximum of two sampled values from the Gaussian distribution. This transformation can be viewed as a compression and translation of the Gaussian distribution, which forms the core of the uncertainty-aware reinforcement learning method. This transformation reduces uncertainty and embodies an optimistic exploration approach.
[0116] The target evaluator selection scheme provided in this application is as follows Figure 7As shown. If at least two candidate actions are sampled, four evaluators are actually obtained that can be used for policy optimization. Among these evaluators, candidate actions "a" and "c" are more optimistic, while "d" and "b" represent the maximum and minimum values, respectively. Optimism here refers to a positive estimate of potential rewards. Specifically, the sampled candidate actions "a" and "c" are considered to be more credible choices because they show higher Q values during the simulation. Using these credible evaluators to guide policy optimization can promote the policy to explore high-return areas more actively. However, the values of candidate actions "d" and "b" are also retained to form a reference range, so that potential benefits and risks can be more comprehensively evaluated. This approach reflects a trade-off, by selecting relatively credible evaluators to guide policy optimization, while not ignoring risks while exploring.
[0117] Safety is of vital importance in autonomous driving. Using a conservative strategy (min) can better avoid overly aggressive operations (such as sudden acceleration or risky overtaking), and using an aggressive exploration operation (max) ensures that even under conservative estimates, the potential best action in the current state can be found. The core idea of the above formula (3) is to balance aggressive and conservative action selection strategies through dual Q networks and conservative estimates in complex environments (such as autonomous driving). In specific implementations, the state data and action pairs of autonomous driving can be combined for training to achieve safer and more efficient driving behavior.
[0118] 504. The server determines the target driving action from the at least two candidate actions based on the target evaluator, where the target driving action is the driving action used in this round of training.
[0119] In a possible implementation, the server determines a target evaluation value corresponding to the target evaluator, where the target evaluation value is the first evaluation value, the second evaluation value, the third evaluation value, or the fourth evaluation value. The server determines the target driving action as the candidate action corresponding to the target evaluation value.
[0120] For example, referring to the above formula (3), the target evaluator Q t () Corresponding candidate action a t Identify the target driving action.
[0121] 505. The server performs this round of training on the intelligent agent based on the target evaluator, the sample state data, and the target driving action.
[0122] In a possible implementation, when the agent is used to output a random strategy, the server determines a first policy gradient corresponding to the random strategy. The server performs this round of training on the agent based on the target evaluation value of the target driving action output by the target evaluator, the first action value of the target driving action under the sample state data output by the agent, and the first policy gradient.
[0123] Among them, the first policy gradient is used to optimize the strategy. The strategy refers to the strategy of the intelligent agent to output driving actions. The first policy gradient can indicate the direction of policy optimization.
[0124] For example, when the agent is used to output a random policy, the server determines a first policy gradient corresponding to the random policy. The server substitutes the target evaluation value and the first action value into the first policy gradient to obtain a first loss function. The server uses the first loss function to perform this round of training on the parameters of the agent.
[0125] For example, when the agent is used to output a random policy, the first policy gradient is expressed by the following formula (4).
[0126] J(π′(·|s t ))=D KL (π′(·|s t )||π Q (·|s t )) (4)
[0127] Among them, J(π′(·|s t )) is the first policy gradient, D KL (q||p)=∑ x q(x)(log q(x)-logp(x)) is the KL divergence, where x is a random variable, and p(x) and q(x) represent two distributions of the random variable x. Intuitively, when minimizing the first policy gradient, the optimal solution is π Q (·|s t ). If the Q value converges to the optimal value, the strategy will also converge to the optimal solution. Therefore, the actual role of KL divergence is to perform strategy extraction.
[0128] During the policy improvement process, the server solves the following formula (5).
[0129]
[0130] Among them, KL divergence can be expressed in the desired form, such as: In theory, the optimal solution of KL divergence can be directly obtained, but the cost of explicit solution is too high. Instead, the optimal solution can be gradually approached by iteratively minimizing KL divergence. In order to make the strategy trainable, the reparameterization technique is applied to reparameterize the target action into the following formula (6).
[0131]
[0132] Among them, a t represents the re-parameterized target action, f(s t ,∈ t )=μ(s t )+σ(s t )*∈ t , μ and σ represent the mean and variance of the Gaussian strategy respectively. Consider the parameterized strategy π φ and Q function Q θ , which can minimize the policy objective and derive the policy gradient relative to φ as shown in the following formula (7)
[0133]
[0134] Among them, f φ (s t ,∈ t )=μ φ (s t )+σ φ (s t )*∈ t .
[0135] Combining the above formula, we can get the unbiased gradient estimate of the first policy gradient, as shown in the following formula (8).
[0136]
[0137] It is important to note that the trusted evaluator and target action are implicitly derived by sampling at least two candidate actions and comparing their Q-values (values of the Q-function). Since the policy and Q-function (function of the candidate evaluator) are parameterized using a neural network, the two terms in formula (8) can be forward-computed using a deep learning framework. The automatic gradient mechanism inherent in the framework will automatically handle back-propagation.
[0138] In a possible implementation, when the agent is used to output a deterministic policy, the server determines a second policy gradient corresponding to the output deterministic policy. The server performs this round of training on the agent based on the target evaluation value of the target driving action output by the target evaluator, the target driving action, and the second policy gradient.
[0139] Among them, the second policy gradient is used to optimize the strategy. The strategy refers to the strategy of the intelligent agent to output driving actions. The second policy gradient can indicate the direction of policy optimization.
[0140] For example, when the agent is used to output a deterministic policy, the server determines a second policy gradient corresponding to the output deterministic policy. The server substitutes the target evaluation value and the target driving action into the second policy gradient to obtain a second loss function. The server uses the second loss function to perform this round of training on the parameters of the agent.
[0141] For example, when the agent is used to output a deterministic policy, the second policy gradient is expressed by the following formula (9).
[0142]
[0143] See also Figure 8 Taking passing two consecutive right-angle turns as an example, a conservative navigator can guide the driver to slow down and then pass the two right-angle turns in sequence, as shown by the solid line. An optimistic navigator usually chooses to enter the turn at an angle and pass two right-angle turns continuously at high speed, as shown by the dotted line. Generally speaking, high-speed cornering can pass through consecutive turns quickly. This shows that optimistic navigators are more efficient, and this optimism is based on rich experience. The core idea of the technical solution provided in the embodiment of the present application is to find an optimistic navigator, that is, a trusted evaluator, and use this trusted evaluator to reach an optimistic estimate of the decision, thereby better guiding the learning of the strategy.
[0144] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described one by one here.
[0145] Through the technical solution provided by the embodiment of the present application, during any round of training of the agent, the sample state data of this round of training is obtained. Using the sample state data and the agent, the target evaluator used in this round of training is determined from two candidate evaluators and the target driving action is determined, thereby implementing the screening of candidate evaluators and the screening of driving actions. By using the screened target evaluator, sample state data and target driving action to train the agent, a better training effect can be achieved. Thus, in the process of using the trained agent for autonomous driving, a better effect can be achieved.
[0146] In addition, the trusted evaluator does not need to explicitly model the distributed reward value function. The value estimation method replaces the modeling method, which greatly reduces the time and space complexity of the traditional method. The policy gradient reinforcement learning algorithm guided by the trusted evaluator improves the sample efficiency of the current optimal algorithm and reduces the learning cost and computational cost of the strategy.
[0147] In addition to the above steps 501-505, the embodiment of the present application also provides a technical solution for using the trained intelligent agent to perform autonomous driving, taking the execution subject as the terminal as an example, see Fig. 9 , the method comprises the following steps.
[0148] 901. The terminal obtains target state data of a target vehicle, where the target state data includes environmental data of an environment in which the target vehicle is located and target vehicle state data of the target vehicle.
[0149] The target state data is data collected by sensors during the driving process of the target vehicle.
[0150] 902. The terminal inputs the target state data into the trained intelligent agent to obtain the target driving action corresponding to the target state data.
[0151] The trained intelligent agent is the intelligent agent trained by the above steps 501-505.
[0152] Fig.10 This is a schematic diagram of the structure of a trusted critic-based enhanced trajectory optimization device provided in an embodiment of the present application, see Fig.10 The device includes: a data acquisition module 1001, a determination module 1002 and a training module 1003.
[0153] The data acquisition module 1001 is used to obtain the sample state data of the current round of training during any round of training of the autonomous driving intelligent agent, and the sample state data includes sample environment data and sample vehicle state data. The intelligent agent is used to output driving actions based on the input state data.
[0154] The determination module 1002 is used to determine the target evaluator used in this round of training from two candidate evaluators and determine the target driving action based on the sample state data and the intelligent agent. The target evaluator is a candidate evaluator that meets preset conditions, and the target driving action is the driving action used in this round of training.
[0155] The training module 1003 is used to perform this round of training on the intelligent agent based on the target evaluator, the sample state data and the target driving action.
[0156] In a possible implementation, the determination module 1002 is configured to input the sample state data into the agent to obtain at least two candidate actions. Based on the sample state data and the at least two candidate actions, determine a target evaluator used in this round of training from the two candidate evaluators. Based on the target evaluator, determine the target driving action from the at least two candidate actions.
[0157] In a possible implementation, the at least two candidate actions include a first candidate action and a second candidate action, the two candidate evaluators include a first candidate evaluator and a second candidate evaluator, and the determination module 1002 is used to input the sample state data and the first candidate action into the first candidate evaluator and the second candidate evaluator respectively, to obtain a first evaluation value and a second evaluation value, the first evaluation value is the evaluation value of the first candidate evaluator to the first candidate action, and the second evaluation value is the evaluation value of the second candidate evaluator to the first candidate action. The sample state data and the second candidate action are input into the first candidate evaluator and the second candidate evaluator respectively, to obtain a third evaluation value and a fourth evaluation value, the third evaluation value is the evaluation value of the first candidate evaluator to the second candidate action, and the fourth evaluation value is the evaluation value of the second candidate evaluator to the second candidate action. Based on the first evaluation value, the second evaluation value, the third evaluation value and the fourth evaluation value, the target evaluator is determined from the first candidate evaluator and the second candidate evaluator.
[0158] In a possible implementation, the determination module 1002 is used to determine the candidate evaluator corresponding to the lower evaluation value of the first evaluation value and the second evaluation value as the first reference evaluator, and the first reference evaluator is the first candidate evaluator or the second candidate evaluator. The candidate evaluator corresponding to the lower evaluation value of the third evaluation value and the fourth evaluation value is determined as the second reference evaluator, and the first reference evaluator is the first candidate evaluator or the second candidate evaluator. The reference evaluator with the higher evaluation value between the first reference evaluator and the second reference evaluator is determined as the target evaluator.
[0159] In a possible implementation, the determination module 1002 is used to determine a target evaluation value corresponding to the target evaluator, where the target evaluation value is the first evaluation value, the second evaluation value, the third evaluation value, or the fourth evaluation value, and to determine the target driving action from a candidate action corresponding to the target evaluation value.
[0160] In a possible implementation, the determination module 1002 is used to input the sample state data into the agent when the agent is used to output a random strategy, and obtain an action strategy corresponding to the sample state data. Perform random sampling in the action strategy to obtain the at least two candidate actions. When the agent is used to output a deterministic strategy, input the sample state data into the agent, and obtain an initial action corresponding to the sample state data. Superimpose the initial action with the target noise to obtain an action strategy. Perform random sampling in the action strategy to obtain the at least two candidate actions.
[0161] In a possible implementation, the training module 1003 is used to determine the first policy gradient corresponding to the random policy when the agent is used to output a random policy. Based on the target evaluation value of the target driving action output by the target evaluator, the first action value of the target driving action under the sample state data output by the agent, and the first policy gradient, the agent is trained in this round. When the agent is used to output a deterministic policy, the second policy gradient corresponding to the output deterministic policy is determined. Based on the target evaluation value of the target driving action output by the target evaluator, the target driving action and the second policy gradient, the agent is trained in this round.
[0162] In a possible implementation, the training module 1003 is used to substitute the target evaluation value and the first action value into the first policy gradient to obtain a first loss function, and use the first loss function to perform this round of training on the parameters of the agent.
[0163] In a possible implementation, the training module 1003 is used to substitute the target evaluation value and the target driving action into the second policy gradient to obtain a second loss function, and use the second loss function to perform this round of training on the parameters of the agent.
[0164] In a possible implementation, the device further includes a prediction module for obtaining target state data of a target vehicle, the target state data including environmental data of an environment in which the target vehicle is located and target vehicle state data of the target vehicle. The target state data is input into a trained agent to obtain a target driving action corresponding to the target state data.
[0165] It should be noted that: the enhanced trajectory optimization device based on trusted critics provided in the above embodiment only uses the division of the above functional modules as an example when training an intelligent agent. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the enhanced trajectory optimization device based on trusted critics provided in the above embodiment and the enhanced trajectory optimization method embodiment based on trusted critics belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0166] Through the technical solution provided by the embodiment of the present application, during any round of training of the agent, the sample state data of this round of training is obtained. Using the sample state data and the agent, the target evaluator used in this round of training is determined from two candidate evaluators and the target driving action is determined, thereby implementing the screening of candidate evaluators and the screening of driving actions. By using the screened target evaluator, sample state data and target driving action to train the agent, a better training effect can be achieved. Thus, in the process of using the trained agent for autonomous driving, a better effect can be achieved.
[0167] The above-mentioned computer device can be implemented as a server. The structure of the server is introduced below:
[0168] Fig.11 It is a structural diagram of a server provided in an embodiment of the present application. The server 1100 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1101 and one or more memories 1102, wherein the one or more memories 1102 store at least one computer program, and the at least one computer program is loaded and executed by the one or more processors 1101 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server 1100 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 1100 may also include other components for implementing device functions, which will not be described in detail here.
[0169] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program, and the computer program can be executed by a processor to implement the trusted critic-based enhanced trajectory optimization method in the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device.
[0170] In an exemplary embodiment, a computer program product or a computer program is also provided, which includes a program code, the program code is stored in a computer-readable storage medium, a processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device performs the above-mentioned trusted critic-based enhanced trajectory optimization method.
[0171] In some embodiments, the computer program involved in the embodiments of the present application may be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected by a communication network. Multiple computer devices distributed at multiple locations and interconnected by a communication network may constitute a blockchain system.
[0172] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0173] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A trusted critic-based enhanced trajectory optimization method, characterized in that: The method comprises: During any round of training of the autonomous driving agent, sample state data of the current round of training is obtained, wherein the sample state data includes sample environment data and sample vehicle state data, and the agent is used to output driving actions based on the input state data; Based on the sample state data and the agent, determine a target evaluator used in this round of training from two candidate evaluators and determine a target driving action, wherein the target evaluator is a candidate evaluator that meets a preset condition, and the target driving action is a driving action used in this round of training; Based on the target evaluator, the sample state data and the target driving action, the agent is trained in this round.
2. The method according to claim 1, characterized in that The step of determining a target evaluator used in this round of training from two candidate evaluators based on the sample state data and the intelligent agent and determining a target driving action includes: Inputting the sample state data into the agent to obtain at least two candidate actions; Based on the sample state data and the at least two candidate actions, determining a target evaluator used in this round of training from the two candidate evaluators; The target driving action is determined from the at least two candidate actions based on the target evaluator.
3. The method according to claim 2, characterized in that The at least two candidate actions include a first candidate action and a second candidate action, the two candidate evaluators include a first candidate evaluator and a second candidate evaluator, and determining a target evaluator used in this round of training from the two candidate evaluators based on the sample state data and the at least two candidate actions includes: Input the sample state data and the first candidate action into the first candidate evaluator and the second candidate evaluator respectively, to obtain a first evaluation value and a second evaluation value, wherein the first evaluation value is an evaluation value of the first candidate evaluator on the first candidate action, and the second evaluation value is an evaluation value of the second candidate evaluator on the first candidate action; Input the sample state data and the second candidate action into the first candidate evaluator and the second candidate evaluator respectively, to obtain a third evaluation value and a fourth evaluation value, wherein the third evaluation value is an evaluation value of the first candidate evaluator on the second candidate action, and the fourth evaluation value is an evaluation value of the second candidate evaluator on the second candidate action; The target evaluator is determined from among the first candidate evaluators and the second candidate evaluators based on the first evaluation value, the second evaluation value, the third evaluation value, and the fourth evaluation value.
4. The method according to claim 3, characterized in that The step of determining the target evaluator from the first candidate evaluator and the second candidate evaluator based on the first evaluation value, the second evaluation value, the third evaluation value, and the fourth evaluation value comprises: Determine a candidate evaluator corresponding to a lower evaluation value between the first evaluation value and the second evaluation value as a first reference evaluator, wherein the first reference evaluator is the first candidate evaluator or the second candidate evaluator; Determine a candidate evaluator corresponding to a lower evaluation value between the third evaluation value and the fourth evaluation value as a second reference evaluator, wherein the first reference evaluator is the first candidate evaluator or the second candidate evaluator; The reference evaluator with a higher corresponding evaluation value between the first reference evaluator and the second reference evaluator is determined as the target evaluator.
5. The method according to claim 3 or 4, characterized in that: The determining the target driving action from the at least two candidate actions based on the target evaluator includes: Determine a target evaluation value corresponding to the target evaluator, wherein the target evaluation value is the first evaluation value, the second evaluation value, the third evaluation value or the fourth evaluation value; The candidate actions corresponding to the target evaluation value are determined as the target driving action.
6. The method according to claim 2, characterized in that The step of inputting the sample state data into the agent to obtain at least two candidate actions includes: In the case where the agent is used to output a random strategy, the sample state data is input into the agent to obtain an action strategy corresponding to the sample state data; random sampling is performed in the action strategy to obtain the at least two candidate actions; In the case where the agent is used to output a deterministic strategy, the sample state data is input into the agent to obtain an initial action corresponding to the sample state data; the initial action is superimposed with the target noise to obtain an action strategy; and random sampling is performed in the action strategy to obtain the at least two candidate actions.
7. The method according to claim 1, characterized in that The performing a current round of training on the agent based on the target evaluator, the sample state data and the target driving action includes: In the case where the agent is used to output a random strategy, determining a first policy gradient corresponding to the random strategy; performing a current round of training on the agent based on a target evaluation value of the target driving action output by the target evaluator, a first action value of the target driving action under the sample state data output by the agent, and the first policy gradient; When the agent is used to output a deterministic strategy, a second policy gradient corresponding to the output deterministic strategy is determined; and based on the target evaluation value of the target driving action output by the target evaluator, the target driving action and the second policy gradient, the agent is trained in this round.
8. The method according to claim 7, characterized in that The present round of training of the agent based on the target evaluation value of the target driving action output by the target evaluator, the first action value of the target driving action under the sample state data output by the agent, and the first policy gradient comprises: Substituting the target evaluation value and the first action value into the first policy gradient to obtain a first loss function; Using the first loss function to perform this round of training on the parameters of the agent; The step of performing a current round of training on the agent based on the target evaluation value of the target driving action output by the target evaluator, the target driving action, and the second policy gradient includes: Substituting the target evaluation value and the target driving action into the second policy gradient to obtain a second loss function; The second loss function is used to perform this round of training on the parameters of the agent.
9. The method according to claim 7, characterized in that: The method comprises: Acquire target state data of a target vehicle, the target state data including environmental data of an environment in which the target vehicle is located and target vehicle state data of the target vehicle; The target state data is input into the trained intelligent agent to obtain the target driving action corresponding to the target state data.
10. A computer device, characterized in that: The computer device includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the trusted critic-based enhanced trajectory optimization method according to any one of claims 1 to 9.